Lead Site Reliability Engineer
$145k - $200kMattermost
Lead Site Reliability Engineer
Mattermost is the leading collaborative workflow platform for defense, intelligence, security, and critical infrastructure. Trusted by the U.S. Department of War and Fortune 500s, our platform runs on-premises and in private clouds, delivering secure messaging, file sharing, workflow automation, audio/screenshare, and project management—all with full data and operational control. Mattermost powers high-stakes workflows across mission planning, real-time, real-world operations, DevSecOps, incident response, and cyber defense—enabling secure collaboration from tactical edge and DDIL environments to enterprise HQ. Teams operate across web, desktop, and mobile, with embedded interoperability for Microsoft Teams, Outlook, and Microsoft 365. To learn more, visit
Mattermost is seeking an experienced and visionary Lead Site Reliability Engineer (SRE) to guide the architecture, reliability, and operational excellence of the infrastructure powering our secure, mission-critical collaboration platform.
In this role, you will provide technical leadership across our SRE function, driving strategic initiatives for scalability, observability, performance, and automation across cloud and hybrid environments. You will mentor engineers, establish best practices, and collaborate closely with development, security, and operations teams to ensure our customers in defense, government, and critical infrastructure sectors experience exceptional reliability and performance.
Responsibilities Include:
- Define the strategy, architecture, and roadmap for Mattermost's site reliability engineering function, aligning infrastructure initiatives with product and business goals.
- Lead the design, deployment, and optimization of production-grade containerized workloads, infrastructure-as-code, and compliant cloud environments for regulated domains (e.g., FedRAMP, DoD).
- Establish and evolve observability, monitoring, and alerting frameworks to ensure performance, reliability, and capacity planning at scale.
- Drive incident management processes, including on-call rotations, root cause analysis, and systemic reliability improvements.
- Partner with security and compliance teams to meet data sovereignty, security, and regulatory requirements.
- Champion automation and operational excellence to improve efficiency, reduce risk, and scale operations.
- Oversee cloud cost management and capacity planning to optimize infrastructure spending while meeting performance targets.
- Build and maintain a developer platform that enables fast, secure software delivery and improves application stability in production.
- Mentor and coach SRE team members, fostering a culture of learning, collaboration, and technical excellence.
Requirements:
- BS in Computer Science, Cybersecurity, Software Engineering, or a related technical field, or equivalent experience, with 5+ years of relevant experience in site reliability engineering, DevOps, or cloud infrastructure roles.
- Proven expertise in container orchestration platforms, ideally Kubernetes.
- Extensive experience with infrastructure-as-code, ideally Terraform.
- Strong background in cloud platforms, ideally AWS.
- Demonstrated experience designing and implementing monitoring, alerting, and performance optimization strategies.
- Exceptional troubleshooting and incident management skills for distributed systems.
- Proficiency in at least one scripting or programming language for automation.
- Excellent communication skills with a track record of influencing cross-functional teams.
- Experience leading globally distributed teams in a remote-first environment.
Preferences:
- Familiarity with observability stacks such as Grafana and Prometheus.
- Experience designing high-availability, disaster recovery, and scaling architectures.
- Exposure to GCP and Azure cloud environments.
- Leadership experience in highly regulated industries such as defense, finance, or critical infrastructure.
- Experience with U.S. federal compliance frameworks and authorization processes, including FedRAMP, DoD ATO, NIST 800-53, and related government standards.
- Experience preparing, delivering, and maintaining software offerings through AWS Marketplace and other cloud provider marketplaces (e.g., Azure Marketplace, Google Cloud Marketplace), including packaging, compliance validation, and ongoing operational support.
- Open-source contributions in reliability, DevOps, or infrastructure tooling.
- Certifications in cloud infrastructure, reliability, or DevOps engineering (e.g., CKA, CKAD, AWS Certified Solutions Architect).
Compensation
Salary range: $145,000 – $200,000
Mattermost takes a market-based approach to pay. Compensation is determined based on skills, experience, qualifications, and work location. Ranges may be updated as market conditions evolve.
U.S. Eligibility & Compliance
This role may require obtaining and maintaining a U.S. government security clearance. Candidates must meet federal eligibility requirements to be considered. For more information visit Security Clearances — United States Department of State
Applicants must meet eligibility requirements for access to export-controlled information as defined by U.S. export control laws, including EAR and ITAR. For more information visit the Bureau of Industry and Security and the Directorate of Defense Trade Controls.
Mattermost is an EEO Employer, we are a remote-first, open-source company. We are continually working to expand our hiring in more countries and regions, ensuring compliance with local laws and regulations, which takes time. Mattermost values your unique perspective—we welcome all applicants. We encourage individuals from all backgrounds to apply and are committed to assessing candidates based on their skills and qualifications. We do not tolerate discrimination against staff or applicants based on race, religion, national origin, age, disability, pregnancy status, veteran status, or other personal characteristics. If you require accommodations during the interview process, please let us know—we're happy to assist.
- ...This job is responsible for building and leading a team to deliver technology products... ...standards, promoting design, engineering, and organizational practices, and advocating... ...stakeholders.Overview:Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the...SuggestedFull timeWork at officeDay shift
$99k - $225k
Site Reliability Engineer, LeadThe Opportunity: As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and...SuggestedFull timeContract workPart timeWork at officeLocal areaRemote work$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...SuggestedWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week- Job Title:Lead Information Security Site Reliability EngineerWells Fargo is back in the office collaborating for fabulous outcomes!This role is in the... ...About this Role: We are seeking a Lead Site Reliability Engineer (SRE) to lead a team of SRE engineers to help mature...SuggestedFull timeWork experience placementWork at officeWork from homeVisa sponsorship2 days per week3 days per week
- Job Summary Job Summary The Support Lead (SRE) is responsible for overseeing the support operations and site reliability engineering tasks, ensuring the effective functioning of systems and applications. The primary goal is to enhance system performance, availability,...Suggested
$175k - $250k
...developed by our expert team of lawyers, engineers and research scientists. We’ve found... ...Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the... ...networking) across 50+ global regions Lead incident management processes, including...Full timeRelocation package$140k - $230k
...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development... ...deployment processes, and drive automation initiatives. Lead incident resolution: You will conduct thorough root cause...Full time- ...only provider of enterprise-scale context engines capable of analyzing trillions of real-... ...seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team.... ...issues before they impact end-users. Lead troubleshooting efforts for complex production...Full time
- ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely... ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale...Permanent employmentWork experience placementWork at officeLocal area
$104.9k - $174.7k
...you passionate about improving reliability, scalability, and resilience... ...-based container platforms, leading vulnerability remediation efforts... ...of reliability engineering tasks within team backlogs.Lead... ...(IaaS).Background in DevOps, site reliability engineering practices...Full timeLocal area$176k - $282k
Responsibilities Lead and support the migration of on-premise corporate hardware... ...operational excellence, security, reliability, performance efficiency, cost optimization, and sustainability.Perform Site Reliability Engineering (SRE) functions, including automation...Contract workShift work$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...Full timeTemporary workWork experience placementFlexible hours$105.6k - $145.2k
Architect the Future as our Site Reliability Engineer!Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site... ..., and efficiency, ensuring best practices are followed.Lead incident response efforts and conduct deep-dive root cause...Ongoing contractFull timeWork at officeLocal areaWorldwide$138.1k - $198.2k
...technology that simply works. The SRE Engineering Enablement Team supports our CI Platforms... ...engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter... ...experiment on and build great products Lead the design and platform evolution of critical...Permanent employmentFull timeTemporary workWork experience placementLocal areaRemote workFlexible hours$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...Full timeWorldwide$133k - $190k
Site Reliability Engineer needed for a full time opportunity with SOC's direct client based in Herndon, VA. Direct Hire Role **Due to federal requirements, candidates must hold and possess an Active DOW TS/SCI security clearance to be considered for this role.** SOC is...Full time$112k - $179k
...About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in... ...extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise...Contract workWorldwideShift work$123k - $165k
Job Posting Title:Site Reliability Engineer IIReq ID:10143234Job Description:Department/Group OverviewOur engineering fleet is a horizontal set of teams providing engineering services across the organization. Our specific team provides reliability engineering and operational...Full time- ...developers on how to make things better. We collectively strive to build and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.ResponsibilitiesEKS & Karpenter Management: Manage, upgrade, and autoscale...For contractors
$112k - $137k
...Financial Group (MUFG), one of the world’s leading financial groups. Across the globe, we’... ...will work at an MUFG office or client sites four days per week and work remotely... ...highly motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable, and...Full timeWork at officeLocal areaRemote work- ...billion and backed by world-leading investors including T. Rowe Price... ...’s next.About the teamThe Engineering team at Airwallex is a diverse... ...together to build scalable, reliable, and secure products that empower... ....What you’ll doAs a Senior Site Reliability Engineer, you’ll...Temporary workLocal areaWorldwide
$152.6k - $191.5k
...for partnering with leaders across engineering and technology to define objective reliability goals for services. Key... ....Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced... ...benefits eligible. We provide industry-leading benefits, access to paid time off...Full timeWork at officeDay shift- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and...
- ...platforms, and vendors to assess resiliency, reliability, and operational risk.Design and... ...enterprise resiliency and reliability standards.Lead blameless post‑incident reviews for high... ....Actively participate in reliability engineering and resilience communities of practice,...Full time
$174.92k - $209.91k
...access to data as simple and reliable as electricity. With Fivetran... ...and ready to query, with no engineering or maintenance required. We’re... ...bringing together two industry-leading companies with a shared... ...our teams, systems, and career sites.About the RoleFivetran is building...Full timeWork at officeRemote work- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the IP, you will solve complex and broad business problems with simple and straightforward solutions...
$165k - $265k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink...Permanent employmentTemporary workWorldwideWeekend work$179.2k - $268.8k
...Latitude team, you’ll work alongside leading experts across machine learning and robotics... ..., test operations, systems and safety engineering - all dedicated to redefining the... ...and Palo Alto, Calif.Meet the team:As a Site Reliability Engineer on the team, you will be...Permanent employmentFull timeWork at officeImmediate startVisa sponsorship$102.1k - $202.2k
...yearEmployment type: Full-TimeWork site: 0 days / week in-office -... ...EngineeringDiscipline: Site Reliability EngineeringCompany:... ...workloads. As a Site Reliability Engineer II, you will take ownership of... ...SLOs and operational metrics. Lead post-incident reviews for owned...Ongoing contractWork at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead c# developer United States
- lead web developer United States
- lead industrial engineer United States
- lead gameplay engineer United States
- lead product engineer United States
- lead operating engineer United States
- lead algorithm engineer United States
- lead quality engineer United States
- lead telecom engineer United States
- lead sharepoint developer United States

