Remote Lead Site Reliability Engineer
$145k - $200kMattermost
Mattermost is the leading collaborative workflow platform for defense, intelligence, security, and critical infrastructure. Trusted by the U.S. Department of War and Fortune 500s, our platform runs on-premises and in private clouds, delivering secure messaging, file sharing, workflow automation, audio/screenshare, and project management—all with full data and operational control. Mattermost powers high-stakes workflows across mission planning, real-time, real-world operations, DevSecOps, incident response, and cyber defense—enabling secure collaboration from tactical edge and DDIL environments to enterprise HQ. Teams operate across web, desktop, and mobile, with embedded interoperability for Microsoft Teams, Outlook, and Microsoft 365.
To learn more, visit
Mattermost is seeking an experienced and visionary Lead Site Reliability Engineer (SRE) to guide the architecture, reliability, and operational excellence of the infrastructure powering our secure, mission-critical collaboration platform.
In this role, you will provide technical leadership across our SRE function, driving strategic initiatives for scalability, observability, performance, and automation across cloud and hybrid environments. You will mentor engineers, establish best practices, and collaborate closely with development, security, and operations teams to ensure our customers in defense, government, and critical infrastructure sectors experience exceptional reliability and performance.
Responsibilities Include:
- Define the strategy, architecture, and roadmap for Mattermost’s site reliability engineering function, aligning infrastructure initiatives with product and business goals.
- Lead the design, deployment, and optimization of production-grade containerized workloads, infrastructure-as-code, and compliant cloud environments for regulated domains (e.g., FedRAMP, DoD).
- Establish and evolve observability, monitoring, and alerting frameworks to ensure performance, reliability, and capacity planning at scale.
- Drive incident management processes, including on-call rotations, root cause analysis, and systemic reliability improvements.
- Partner with security and compliance teams to meet data sovereignty, security, and regulatory requirements.
- Champion automation and operational excellence to improve efficiency, reduce risk, and scale operations.
- Oversee cloud cost management and capacity planning to optimize infrastructure spending while meeting performance targets.
- Build and maintain a developer platform that enables fast, secure software delivery and improves application stability in production.
- Mentor and coach SRE team members, fostering a culture of learning, collaboration, and technical excellence.
Requirements:
- BS in Computer Science, Cybersecurity, Software Engineering, or a related technical field, or equivalent experience, with 5+ years of relevant experience in site reliability engineering, DevOps, or cloud infrastructure roles.
- Proven expertise in container orchestration platforms, ideally Kubernetes.
- Extensive experience with infrastructure-as-code, ideally Terraform.
- Strong background in cloud platforms, ideally AWS.
- Demonstrated experience designing and implementing monitoring, alerting, and performance optimization strategies.
- Exceptional troubleshooting and incident management skills for distributed systems.
- Proficiency in at least one scripting or programming language for automation.
- Excellent communication skills with a track record of influencing cross-functional teams.
- Experience leading globally distributed teams in a remote-first environment.
Preferences:
- Familiarity with observability stacks such as Grafana and Prometheus.
- Experience designing high-availability, disaster recovery, and scaling architectures.
- Exposure to GCP and Azure cloud environments.
- Leadership experience in highly regulated industries such as defense, finance, or critical infrastructure.
- Experience with U.S. federal compliance frameworks and authorization processes, including FedRAMP, DoD ATO, NIST 800-53, and related government standards.
- Experience preparing, delivering, and maintaining software offerings through AWS Marketplace and other cloud provider marketplaces (e.g., Azure Marketplace, Google Cloud Marketplace), including packaging, compliance validation, and ongoing operational support.
- Open-source contributions in reliability, DevOps, or infrastructure tooling.
- Certifications in cloud infrastructure, reliability, or DevOps engineering (e.g., CKA, CKAD, AWS Certified Solutions Architect).
Compensation
Salary range: $145,000 – $200,000
Mattermost takes a market-based approach to pay. Compensation is determined based on skills, experience, qualifications, and work location. Ranges may be updated as market conditions evolve. .
U.S. Eligibility & Compliance
This role may require obtaining and maintaining a U.S. government security clearance. Candidates must meet federal eligibility requirements to be considered. For more information visit Security Clearances — United States Department of State
Applicants must meet eligibility requirements for access to export-controlled information as defined by U.S. export control laws, including EAR and ITAR. For more information visit the Bureau of Industry and Security and the Directorate of Defense Trade Controls .
Mattermost is an EEO Employer, we are a remote-first, open-source company.
We are continually working to expand our hiring in more countries and regions, ensuring compliance with local laws and regulations, which takes time.
Mattermost values your unique perspective—we welcome all applicants. We encourage individuals from all backgrounds to apply and are committed to assessing candidates based on their skills and qualifications. We do not tolerate discrimination against staff or applicants based on race, religion, national origin, age, disability, pregnancy status, veteran status, or other personal characteristics.
If you require accommodations during the interview process, please let us know—we’re happy to assist.
Jobicy JobID: 153852$99k - $225k
Site Reliability Engineer, LeadThe Opportunity: As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance... ...expected to have their cameras on during meetings.Remote: If this position is listed as remote, there may...Remote workFull timeContract workPart timeWork at officeLocal area$100.1k - $180.2k
...brands, offering comprehensive engineering, supply chain, and... ...a vast network of over 100 sites worldwide, Jabil combines global... ...the globe.Jabil is seeking a Lead Site Reliability Infrastructure and Security... ...: Remote - USA; Austin, TXType: Full...Remote workTemporary workWork at officeLocal areaWorldwide- ...Site Reliability Engineer Company: GitLab Work Type: Remote Employment: Full Time Location: CA, US Seniority: Senior Level Technologies: Terraform, Ansible, Kubernetes, Go, Ruby, Jsonnet, Prometheus, ELK, Grafana Requirements: Senior-level SRE with strong Terraform/IaC...Remote workFull time
- Site Reliability Engineer Company: ContainIQ Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Requirements: Remote, full-time role; prior SRE experience preferred. Role: Individual Contributor Category: Information TechnologyRemote workFull time
- ...Senior Site Reliability Engineer Remote – Home Based Job Summary We’re partnering with a company in the SaaS space to find a Senior Site Reliability Engineer . In this role, you’ll be part of the IT Operations group responsible for maintaining all environments...Remote workTemporary workWork from homeFlexible hours
$140k - $180k
...grows, we are investing in the reliability and platform systems that... ...fast, resilient, and easy for engineering teams to operate. Position... ...Overview We’re hiring a Senior Site Reliability Engineer to... ...urgency. Endear is a lean, remote team where individuals have...Remote workWork from homeHome officeFlexible hours- ...Site Reliability Engineer Company: Quzara Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: Azure, Terraform, Bicep, Ansible, Azure Monitor, Azure Automation, Azure Policy, Azure Site Recovery, TLS/SSL Requirements: 4+ years in SRE...Remote workFull time
- ...Site Reliability Engineer Company: Milestone Systems Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Golang, Python, Linux, Shell scripting, Kubernetes, Docker, Terraform, CI/CD, GitOps, ArgoCD, Spinnaker, Prometheus, Datadog...Remote workFull time
- ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS, Kubernetes, Terraform... ...SRE with 5+ years AWS infra, 3+ years in senior/lead roles; strong automation with Terraform, Ansible,...Remote workFull time
- ...Site Reliability Engineer Company: Crunchafi Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Azure, AKS, Azure Kubernetes Service, Terraform, Bicep, ARM templates, GitHub Actions, Azure DevOps, Kubernetes, Docker, App Insights...Remote workFull time
$104.9k - $174.7k
...link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability... ...you may work a hybrid schedule. If not, this role is fully remote. We do not restrict applicants based on job site or posting...Remote workFull timeWork at officeLocal areaWork from home- ...Site Reliability Engineer II This role sits at the intersection of software engineering, infrastructure, and network operations, helping scale... ..., and continuous-improvement skills. Benefits Remote work opportunity in India, with flexibility to work from...Remote workWork at officeWork from homeFlexible hours
$95k - $171k
..., Kubernetes, and ensuring reliability for AI workloads within Akamai... ...platform. As an Site Reliability Engineer II, you will be responsible... ...and protects life online. Leading companies worldwide choose... ...energize and inspire you! #LI-Remote Compensation Akamai is...Remote workPermanent employmentWork experience placementWork at officeWork from homeWorldwideFlexible hours- ...Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Azure, Terraform, ARM, Kubernetes, SonarCloud, CheckPoint, NewRelic, Hadoop, Kafka, Presto, SQL, Redis, WebApps, CI/CD, SaaS, Azure DevOps...Remote workFull time
- ...Site Reliability Engineer We are looking for a Site Reliability Engineer to join our IT Operations group. You will develop automated solutions... ...,IIS,LinuxBackground Check :YesDrug Screen :YesNotes :Remote? Yes, 100% remote anywhere in the US is ok but prefers candidate...Remote workWork experience placementWork at officeWork from home2 days per week1 day per week
- ...Senior Site Reliability Engineer Company: Filevine Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Python, Bash, PowerShell, AWS, Kubernetes, EKS, CloudWatch, Lambda, S3, IAM, CI/CD, Monitoring Requirements: 8+ years in software...Remote workFull time
- ...Site Reliability Engineer Company: Akamai Work Type: Remote Employment: Full Time Location: US, PL Seniority: Mid Level Technologies: C++, Linux, UNIX, Networking, Cloud, Distributed systems Requirements: 3-5 years of relevant experience; Bachelor's degree or equivalent...Remote workFull time
- ...Senior Site Reliability Engineer India What We Do At GoGuardian, we're helping build a future... ...Participate in on-call rotations and lead incident response, ensuring... ...time. Working with us means joining a remote team of diverse, committed, mission-driven...Remote workWork from home
- ...About the Role We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability... ...; ensure we meet or exceed targets consistently ~ Lead observability strategy by designing comprehensive...Remote work
$135k - $170k
...Symmetrio is recruiting a Site Reliability Engineer for its customer, a rapidly growing international... ...technical issues. This is a full-time, remote position. Candidates must reside in... ...customer-facing connectivity incidents: lead the call, keep the customer updated,...Remote workFull time- ...Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: Ansible, Puppet, Chef, Python, Ruby, Bash, PowerShell, Terraform, CloudFormation, AWS, Azure, GCP, Linux, Windows, Networking, Security Requirements...Remote workFull time
- ...Senior Site Reliability Engineer Company: Regrello Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Go, AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, GitHub Actions, GitLab CI, CircleCI, Helm, Helmfile, LaunchDarkly, Otel, Prometheus...Remote workFull time
$165k - $195k
...of ways to work, ranging from a fully remote experience to working full-time in one... ...Your Role We're looking for a Senior Site Reliability Engineer II to help us scale our infrastructure... ...Participate in on-call rotation and lead incident response/root-cause analysis...Remote workFull timeWork at officeLocal areaWork from homeFlexible hours$152k - $205k
...device intelligence. We lead our industry with... ...globally dispersed, 100% remote company. We were named... ...Are you a systems-minded engineer who is happiest when production... ...? Do you want to own reliability for a platform that... ...looking for a Senior Site Reliability Engineer to...Remote workLocal areaWork from homeVisa sponsorship- .... Our partner is looking for a Senior Site Reliability Engineer based in United States. As a Senior... ...detected quickly enough. Take a leading role in high-severity incident response... ...market and relevant experience. Fully remote working environment. Opportunity to...Remote workWork from homeFree visa
$7.5k
...manager, and we have ambitious goals for the future. As a Site Reliability Engineer (SRE), you will work at the intersection of production... ...and trading systems Diagnose and fix bugs in code Lead complex deployments Automate manual workflows...Remote workLocal area- ...Job Title: Site Reliability Engineer (Azure Government & Infrastructure) Pay Type : SALARIED EXEMPT Location: Remote Citizenship Requirement: U.S. Citizen (Required) Summary of Position Role/Responsibilities The Site Reliability Engineer (SRE) for...Remote workFull timeMonday to Friday
- ...Senior Site Reliability Engineer We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they...Remote work
$152k - $195k
...and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and... ...observability — define SLOs, alerts, and dashboards. Lead incident response and postmortems, focusing on root cause and...Remote work- ...Job Title Location Remote - United States Job Category Information Technology, Platform Engineering, Site Reliability Engineering Industry Computer Software, SaaS, National Security Employee Type FT Exempt Manage Others No Minimum Experience 5 Years...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote Lead Site Reliability Engineer. Be the first to apply!
- lead infrastructure engineer Remote
- lead engineer Remote
- lead operating engineer Remote
- lead algorithm engineer Remote
- lead web developer Remote
- lead network engineer Remote
- site reliability engineer sre Remote
- site reliability engineer Remote
- site reliability engineer remote Remote
- procurement specialist remote Remote

