Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$145k - $200k
Full-time

Mattermost

Mattermost is the leading collaborative workflow platform for defense, intelligence, security, and critical infrastructure. Trusted by the U.S. Department of War and Fortune 500s, our platform runs on-premises and in private clouds, delivering secure messaging, file sharing, workflow automation, audio/screenshare, and project management—all with full data and operational control. Mattermost powers high-stakes workflows across mission planning, real-time, real-world operations, DevSecOps, incident response, and cyber defense—enabling secure collaboration from tactical edge and DDIL environments to enterprise HQ. Teams operate across web, desktop, and mobile, with embedded interoperability for Microsoft Teams, Outlook, and Microsoft 365. To learn more, visit Mattermost is seeking an experienced and visionary Lead Site Reliability Engineer (SRE) to guide the architecture, reliability, and operational excellence of the infrastructure powering our secure, mission-critical collaboration platform. In this role, you will provide technical leadership across our SRE function, driving strategic initiatives for scalability, observability, performance, and automation across cloud and hybrid environments. You will mentor engineers, establish best practices, and collaborate closely with development, security, and operations teams to ensure our customers in defense, government, and critical infrastructure sectors experience exceptional reliability and performance. Responsibilities Include:

  • Define the strategy, architecture, and roadmap for Mattermost’s site reliability engineering function, aligning infrastructure initiatives with product and business goals.
  • Lead the design, deployment, and optimization of production-grade containerized workloads, infrastructure-as-code, and compliant cloud environments for regulated domains (e.g., FedRAMP, DoD).
  • Establish and evolve observability, monitoring, and alerting frameworks to ensure performance, reliability, and capacity planning at scale.
  • Drive incident management processes, including on-call rotations, root cause analysis, and systemic reliability improvements.
  • Partner with security and compliance teams to meet data sovereignty, security, and regulatory requirements.
  • Champion automation and operational excellence to improve efficiency, reduce risk, and scale operations.
  • Oversee cloud cost management and capacity planning to optimize infrastructure spending while meeting performance targets.
  • Build and maintain a developer platform that enables fast, secure software delivery and improves application stability in production.
  • Mentor and coach SRE team members, fostering a culture of learning, collaboration, and technical excellence.
Requirements:
  • BS in Computer Science, Cybersecurity, Software Engineering, or a related technical field, or equivalent experience, with 5+ years of relevant experience in site reliability engineering, DevOps, or cloud infrastructure roles.
  • Proven expertise in container orchestration platforms, ideally Kubernetes.
  • Extensive experience with infrastructure-as-code, ideally Terraform.
  • Strong background in cloud platforms, ideally AWS.
  • Demonstrated experience designing and implementing monitoring, alerting, and performance optimization strategies.
  • Exceptional troubleshooting and incident management skills for distributed systems.
  • Proficiency in at least one scripting or programming language for automation.
  • Excellent communication skills with a track record of influencing cross-functional teams.
  • Experience leading globally distributed teams in a remote-first environment.
Preferences:
  • Familiarity with observability stacks such as Grafana and Prometheus.
  • Experience designing high-availability, disaster recovery, and scaling architectures.
  • Exposure to GCP and Azure cloud environments.
  • Leadership experience in highly regulated industries such as defense, finance, or critical infrastructure.
  • Experience with U.S. federal compliance frameworks and authorization processes, including FedRAMP, DoD ATO, NIST 800-53, and related government standards.
  • Experience preparing, delivering, and maintaining software offerings through AWS Marketplace and other cloud provider marketplaces (e.g., Azure Marketplace, Google Cloud Marketplace), including packaging, compliance validation, and ongoing operational support.
  • Open-source contributions in reliability, DevOps, or infrastructure tooling.
  • Certifications in cloud infrastructure, reliability, or DevOps engineering (e.g., CKA, CKAD, AWS Certified Solutions Architect).
Compensation Salary range: $145,000 – $200,000 Mattermost takes a market-based approach to pay. Compensation is determined based on skills, experience, qualifications, and work location. Ranges may be updated as market conditions evolve.. U.S. Eligibility & Compliance This role may require obtaining and maintaining a U.S. government security clearance. Candidates must meet federal eligibility requirements to be considered. For more information visitSecurity Clearances — United States Department of State Applicants must meet eligibility requirements for access to export-controlled information as defined by U.S. export control laws, including EAR and ITAR. For more information visit theBureau of Industry and Security and theDirectorate of Defense Trade Controls. Mattermost is an EEO Employer, we are a remote-first, open-source company. We are continually working to expand our hiring in more countries and regions, ensuring compliance with local laws and regulations, which takes time. Mattermost values your unique perspective—we welcome all applicants. We encourage individuals from all backgrounds to apply and are committed to assessing candidates based on their skills and qualifications. We do not tolerate discrimination against staff or applicants based on race, religion, national origin, age, disability, pregnancy status, veteran status, or other personal characteristics. If you require accommodations during the interview process, please let us know—we’re happy to assist.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in United States vacancy
  •  ...This job is responsible for building and leading a team to deliver technology products...  ...standards, promoting design, engineering, and organizational practices, and advocating...  ...stakeholders.Overview:Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the... 
    Suggested
    Full time
    Work at office
    Day shift

    Bank of America

    Chandler, AZ
    5 hours ago
  • $99k - $225k

    Site Reliability Engineer, LeadThe Opportunity:  As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and... 
    Suggested
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Chantilly, Loudoun County, VA
    2 days ago
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity... 
    Suggested
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    New York, NY
    4 days ago
  • Job title: Site Reliability Engineer (SRE) Bill rate: $52/hr W2 Client address: 2900 W Plano Pkwy Plano, TX 75075 - Role is hybrid (3 days/wk) Years of experience required: 11+ Mandatory skills: Azure DevOps (ADO), GitHub & GitHub Actions, JFrog Artifactory Site... 
    Suggested
    Full time

    IPolarity LLC

    Hanover, PA
    3 days ago
  • $350k

     ...and novel use-cases. We’re hiring to grow the platform alongside the Tinker community. About the Role We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research... 
    Suggested
    Full time
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    3 days ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud...  ...concrete engineering and prioritization decisions. Lead incident response and resolution for production issues, acting... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    3 days ago
  • $192k

     ...here. Role Overview: LeoLabs is seeking a skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will...  ...improves deployment reliability. Within 12 months, you’ll: Lead cross-functional initiatives to improve availability,... 
    Full time
    Work experience placement
    Remote work
    Flexible hours

    LeoLabs, Inc.

    United States
    1 day ago
  • $101.97k - $203.94k

     ...and one community at a time. Position Summary As a Senior Site Reliability Engineer, you will be responsible for ensuring the stability,...  ...Responsibilities Application Performance Monitoring and Observability: Lead the design and implementation of end‑to‑end observability... 
    Hourly pay
    Full time
    Temporary work
    Work experience placement
    Local area

    CVS Health

    Massachusetts
    1 day ago
  • $139k - $257.55k

     ...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,...  ...productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express,... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe

    United States
    1 day ago
  • $146.4k

     ...Communications group. A cross-functional engineering team that develops the...  ...system supports fast and reliable configuration of Akamai's...  ...metadata systems. As a Senior Site Reliability Engineer, you...  ...ESPP). Akamai provides industry-leading benefits including healthcare... 
    Full time
    Work experience placement
    Work at office

    Akamai

    United States
    8 hours ago
  • $180k - $200k

    Company Name: tastytrade Role: Senior Site Reliability Engineer Location: Chicago, IL (Hybrid, 3 days/week in office) Role Summary Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options, futures... 
    Full time
    Work at office
    3 days per week

    tastylive

    Chicago, IL
    1 day ago
  • About the Role We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability,...  ...critical systems; ensure we meet or exceed targets consistently Lead observability strategy by designing comprehensive... 
    Full time

    MeridianLink

    United States
    3 days ago
  •  ...for current or future sponsorship. Maintain and enhance the reliability, availability, and performance of Navy Federal’s systems and...  ...benefits, review the Benefits page [ of the Navy Federal Career Site. Protect Yourself from Job Scams: Navy Federal Credit Union... 
    Full time
    Internship

    Navy Federal Credit Union

    Vienna, VA
    3 days ago
  • $108k - $180k

     ...beauty of fashion accessible to all, promoting its industry-leading, on-demand production methodology, for a smarter, future-ready industry. Position Summary We are seeking a Staff Site Reliability Engineer (Official Title: Staff Site Reliability Engineer I) with... 
    Full time
    Temporary work
    Work at office
    Worldwide
    Flexible hours

    SHEIN

    San Diego, CA
    1 hour ago
  • $135k - $155k

     ...businesses and our colleagues achieve their financial goals. As a leading commercial bank, we remain passionate about serving our...  ...opportunities, and enjoy meaningful work! The Director of Site Reliability Engineer is a pivotal technical leader within the Software... 
    Full time
    Work experience placement
    Shift work
    Afternoon shift

    Webster Bank

    Southington, CT
    1 day ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    4 days ago
  •  ...looking for a Senior SRE to join our Platform Engineering team as the operations owner of our...  .... You’ll be responsible for the reliability, scalability, and continued evolution of...  ...for teams using observability platforms Lead or contribute to platform modernization... 
    Full time

    Dimensional Fund Advisors

    Austin, TX
    8 hours ago
  • $165k - $270k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Redmond, WA
    5 days ago
  •  ...education. Client is currently seeking a talented Software Engineer who is able to work into the Site Reliability Engineer role. This candidate is expected to work...  ...and how to build/utilize (panel of 3-hm, lead, and arch); 2nd round w/ director (panel of 2) Top Must... 
    Remote work

    Intelliswift

    Durham, NC
    5 days ago
  • $130k - $200k

    IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Full time
    Work at office
    Immediate start

    IXL Learning

    San Mateo, CA
    3 days ago
  • $130k - $153k

     ...our customers, and in our growing commitment to land stewardship and recreational access.WHAT YOU WILL DOonX is seeking a Site Reliability Engineer to build and maintain the infrastructure that enables our developers to ship reliably at scale. You'll manage onX's infrastructure... 
    Full time
    Part time
    Work at office

    onXmaps

    Bozeman, MT
    2 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    LexisNexis Risk Solutions Group

    Atlanta, GA
    4 days ago
  • $125k - $185k

     ...HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions and operations. By bringing...  ..., and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    5 days ago
  • $138.4k - $173k

     ...infrastructure as well as help improve the reliability, quality of services and overall...  ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability...  ...about our locations by visiting our site.Compensation & BenefitsThe base salary that... 
    Full time
    Flexible hours

    AppFolio

    Santa Barbara, CA
    5 days ago
  • $152k - $241.5k

     ...technology—and amazing people. NVIDIA is leading the way in groundbreaking developments...  ...automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-...  ..., Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through... 
    Full time

    Nvidia

    Durham, NC
    8 hours ago
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has...  ...guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence...  ...a diverse team of experts as you use leading-edge tech to empower everyone to meet a... 
    Work at office
    Local area

    Realtor.com

    Austin, TX
    1 day ago
  • $145k - $175k

     ...funding options. Our engaging and rewarding environment is designed to help you gain your full potential. Job OverviewThe Site Reliability Engineer supports deployments, cloud infrastructure, and monitoring systems that power Rewards Network's applications and services.... 
    Full time
    Work at office
    Local area
    Flexible hours
    3 days per week

    Rewards Network

    Chicago, IL
    5 days ago
  • $80k - $133k

     ...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and... 
    Permanent employment
    Full time
    Contract work
    Remote work
    Flexible hours

    Guidehouse

    San Antonio, TX
    2 days ago
  • $165k - $190k

    Obsidian Security is the leading SaaS security platform, trusted by global enterprises...  ...DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable,...  ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    2 days ago
  • $158.5k - $172k

     ...velocity energy of a powerhouse startup.As a leading U.S. ordering and delivery marketplace,...  ....About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will...  ...high-impact position driving continuous reliability, deep system optimization, and automation... 
    Full time
    Work at office
    3 days per week

    GrubHub

    Chicago, IL
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!