Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$145k - $200k

Mattermost

Lead Site Reliability Engineer

Mattermost is the leading collaborative workflow platform for defense, intelligence, security, and critical infrastructure. Trusted by the U.S. Department of War and Fortune 500s, our platform runs on-premises and in private clouds, delivering secure messaging, file sharing, workflow automation, audio/screenshare, and project management—all with full data and operational control. Mattermost powers high-stakes workflows across mission planning, real-time, real-world operations, DevSecOps, incident response, and cyber defense—enabling secure collaboration from tactical edge and DDIL environments to enterprise HQ. Teams operate across web, desktop, and mobile, with embedded interoperability for Microsoft Teams, Outlook, and Microsoft 365. To learn more, visit

Mattermost is seeking an experienced and visionary Lead Site Reliability Engineer (SRE) to guide the architecture, reliability, and operational excellence of the infrastructure powering our secure, mission-critical collaboration platform.

In this role, you will provide technical leadership across our SRE function, driving strategic initiatives for scalability, observability, performance, and automation across cloud and hybrid environments. You will mentor engineers, establish best practices, and collaborate closely with development, security, and operations teams to ensure our customers in defense, government, and critical infrastructure sectors experience exceptional reliability and performance.

Responsibilities Include:
  • Define the strategy, architecture, and roadmap for Mattermost's site reliability engineering function, aligning infrastructure initiatives with product and business goals.
  • Lead the design, deployment, and optimization of production-grade containerized workloads, infrastructure-as-code, and compliant cloud environments for regulated domains (e.g., FedRAMP, DoD).
  • Establish and evolve observability, monitoring, and alerting frameworks to ensure performance, reliability, and capacity planning at scale.
  • Drive incident management processes, including on-call rotations, root cause analysis, and systemic reliability improvements.
  • Partner with security and compliance teams to meet data sovereignty, security, and regulatory requirements.
  • Champion automation and operational excellence to improve efficiency, reduce risk, and scale operations.
  • Oversee cloud cost management and capacity planning to optimize infrastructure spending while meeting performance targets.
  • Build and maintain a developer platform that enables fast, secure software delivery and improves application stability in production.
  • Mentor and coach SRE team members, fostering a culture of learning, collaboration, and technical excellence.
Requirements:
  • BS in Computer Science, Cybersecurity, Software Engineering, or a related technical field, or equivalent experience, with 5+ years of relevant experience in site reliability engineering, DevOps, or cloud infrastructure roles.
  • Proven expertise in container orchestration platforms, ideally Kubernetes.
  • Extensive experience with infrastructure-as-code, ideally Terraform.
  • Strong background in cloud platforms, ideally AWS.
  • Demonstrated experience designing and implementing monitoring, alerting, and performance optimization strategies.
  • Exceptional troubleshooting and incident management skills for distributed systems.
  • Proficiency in at least one scripting or programming language for automation.
  • Excellent communication skills with a track record of influencing cross-functional teams.
  • Experience leading globally distributed teams in a remote-first environment.
Preferences:
  • Familiarity with observability stacks such as Grafana and Prometheus.
  • Experience designing high-availability, disaster recovery, and scaling architectures.
  • Exposure to GCP and Azure cloud environments.
  • Leadership experience in highly regulated industries such as defense, finance, or critical infrastructure.
  • Experience with U.S. federal compliance frameworks and authorization processes, including FedRAMP, DoD ATO, NIST 800-53, and related government standards.
  • Experience preparing, delivering, and maintaining software offerings through AWS Marketplace and other cloud provider marketplaces (e.g., Azure Marketplace, Google Cloud Marketplace), including packaging, compliance validation, and ongoing operational support.
  • Open-source contributions in reliability, DevOps, or infrastructure tooling.
  • Certifications in cloud infrastructure, reliability, or DevOps engineering (e.g., CKA, CKAD, AWS Certified Solutions Architect).

Compensation

Salary range: $145,000 – $200,000

Mattermost takes a market-based approach to pay. Compensation is determined based on skills, experience, qualifications, and work location. Ranges may be updated as market conditions evolve.

U.S. Eligibility & Compliance

This role may require obtaining and maintaining a U.S. government security clearance. Candidates must meet federal eligibility requirements to be considered. For more information visit Security Clearances — United States Department of State

Applicants must meet eligibility requirements for access to export-controlled information as defined by U.S. export control laws, including EAR and ITAR. For more information visit the Bureau of Industry and Security and the Directorate of Defense Trade Controls.

Mattermost is an EEO Employer, we are a remote-first, open-source company. We are continually working to expand our hiring in more countries and regions, ensuring compliance with local laws and regulations, which takes time. Mattermost values your unique perspective—we welcome all applicants. We encourage individuals from all backgrounds to apply and are committed to assessing candidates based on their skills and qualifications. We do not tolerate discrimination against staff or applicants based on race, religion, national origin, age, disability, pregnancy status, veteran status, or other personal characteristics. If you require accommodations during the interview process, please let us know—we're happy to assist.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in United States vacancy
  •  ...This job is responsible for building and leading a team to deliver technology products...  ...standards, promoting design, engineering, and organizational practices, and advocating...  ...stakeholders.Overview:Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the... 
    Suggested
    Full time
    Work at office
    Day shift

    Bank of America

    Chandler, AZ
    8 hours ago
  • $99k - $225k

    Site Reliability Engineer, LeadThe Opportunity:  As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and... 
    Suggested
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Chantilly, Loudoun County, VA
    2 days ago
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity... 
    Suggested
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    Jersey City, NJ
    8 hours ago
  • Job Title:Lead Information Security Site Reliability EngineerWells Fargo is back in the office collaborating for fabulous outcomes!This role is in the...  ...About this Role: We are seeking a Lead Site Reliability Engineer (SRE) to lead a team of SRE engineers to help mature... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Work from home
    Visa sponsorship
    2 days per week
    3 days per week

    Wells Fargo

    Charlotte, NC
    8 hours ago
  • Job Summary Job Summary The Support Lead (SRE) is responsible for overseeing the support operations and site reliability engineering tasks, ensuring the effective functioning of systems and applications. The primary goal is to enhance system performance, availability,... 
    Suggested

    TechDigital Group

    Fairfax, VA
    1 day ago
  • $175k - $250k

     ...developed by our expert team of lawyers, engineers and research scientists. We’ve found...  ...Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the...  ...networking) across 50+ global regions Lead incident management processes, including... 
    Full time
    Relocation package

    Harvey

    Remote
    1 day ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...deployment processes, and drive automation initiatives. Lead incident resolution: You will conduct thorough root cause... 
    Full time

    Zoox

    Remote
    1 day ago
  •  ...only provider of enterprise-scale context engines capable of analyzing trillions of real-...  ...seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team....  ...issues before they impact end-users. Lead troubleshooting efforts for complex production... 
    Full time

    Lovelace Ai

    Pittsburgh, PA
    1 day ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    4 days ago
  • $104.9k - $174.7k

     ...you passionate about improving reliability, scalability, and resilience...  ...-based container platforms, leading vulnerability remediation efforts...  ...of reliability engineering tasks within team backlogs.Lead...  ...(IaaS).Background in DevOps, site reliability engineering practices... 
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Texas
    8 hours ago
  • $176k - $282k

    Responsibilities Lead and support the migration of on-premise corporate hardware...  ...operational excellence, security, reliability, performance efficiency, cost optimization, and sustainability.Perform Site Reliability Engineering (SRE) functions, including automation... 
    Contract work
    Shift work

    Peraton Corporation

    Annapolis Junction, MD
    8 hours ago
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Full time
    Work at office

    Restaurant 365

    Austin, TX
    4 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying... 
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    4 days ago
  • $105.6k - $145.2k

    Architect the Future as our Site Reliability Engineer!Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site...  ..., and efficiency, ensuring best practices are followed.Lead incident response efforts and conduct deep-dive root cause... 
    Ongoing contract
    Full time
    Work at office
    Local area
    Worldwide

    Trimble Navigation

    Westminster, CO
    1 day ago
  • $138.1k - $198.2k

     ...technology that simply works.  The SRE Engineering Enablement Team supports our CI Platforms...  ...engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter...  ...experiment on and build great products Lead the design and platform evolution of critical... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Chicago, IL
    4 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    2 days ago
  • $133k - $190k

    Site Reliability Engineer needed for a full time opportunity with SOC's direct client based in Herndon, VA. Direct Hire Role **Due to federal requirements, candidates must hold and possess an Active DOW TS/SCI security clearance to be considered for this role.** SOC is... 
    Full time

    SOC Support Services

    McLean, VA
    3 days ago
  • $112k - $179k

     ...About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in...  ...extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington, VA
    8 hours ago
  • $123k - $165k

    Job Posting Title:Site Reliability Engineer IIReq ID:10143234Job Description:Department/Group OverviewOur engineering fleet is a horizontal set of teams providing engineering services across the organization. Our specific team provides reliability engineering and operational... 
    Full time

    Hulu

    New York, NY
    8 hours ago
  •  ...developers on how to make things better. We collectively strive to build and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.ResponsibilitiesEKS & Karpenter Management: Manage, upgrade, and autoscale... 
    For contractors

    Varo Money

    San Francisco, CA
    2 days ago
  • $112k - $137k

     ...Financial Group (MUFG), one of the world’s leading financial groups. Across the globe, we’...  ...will work at an MUFG office or client sites four days per week and work remotely...  ...highly motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable, and... 
    Full time
    Work at office
    Local area
    Remote work

    MUFG

    Tempe, AZ
    8 hours ago
  •  ...billion and backed by world-leading investors including T. Rowe Price...  ...’s next.About the teamThe Engineering team at Airwallex is a diverse...  ...together to build scalable, reliable, and secure products that empower...  ....What you’ll doAs a Senior Site Reliability Engineer, you’ll... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    8 hours ago
  • $152.6k - $191.5k

     ...for partnering with leaders across engineering and technology to define objective reliability goals for services. Key...  ....Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced...  ...benefits eligible. We provide industry-leading benefits, access to paid time off... 
    Full time
    Work at office
    Day shift

    Bank of America

    Plano, TX
    3 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and... 

    JP Morgan Chase

    Chicago, IL
    1 day ago
  •  ...platforms, and vendors to assess resiliency, reliability, and operational risk.Design and...  ...enterprise resiliency and reliability standards.Lead blameless post‑incident reviews for high...  ....Actively participate in reliability engineering and resilience communities of practice,... 
    Full time

    Vanguard

    Wayne, PA
    3 days ago
  • $174.92k - $209.91k

     ...access to data as simple and reliable as electricity. With Fivetran...  ...and ready to query, with no engineering or maintenance required. We’re...  ...bringing together two industry-leading companies with a shared...  ...our teams, systems, and career sites.About the RoleFivetran is building... 
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    2 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the IP, you will solve complex and broad business problems with simple and straightforward solutions... 

    JP Morgan Chase

    Plano, TX
    4 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  • $179.2k - $268.8k

     ...Latitude team, you’ll work alongside leading experts across machine learning and robotics...  ..., test operations, systems and safety engineering - all dedicated to redefining the...  ...and Palo Alto, Calif.Meet the team:As a Site Reliability Engineer on the team, you will be... 
    Permanent employment
    Full time
    Work at office
    Immediate start
    Visa sponsorship

    Latitude AI

    Pittsburgh, PA
    4 days ago
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 0 days / week in-office -...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...workloads. As a Site Reliability Engineer II, you will take ownership of...  ...SLOs and operational metrics. Lead post-incident reviews for owned... 
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!