Lead Site Reliability Engineer
$145k - $200kMattermost
Mattermost is the leading collaborative workflow platform for defense, intelligence, security, and critical infrastructure. Trusted by the U.S. Department of War and Fortune 500s, our platform runs on-premises and in private clouds, delivering secure messaging, file sharing, workflow automation, audio/screenshare, and project management—all with full data and operational control. Mattermost powers high-stakes workflows across mission planning, real-time, real-world operations, DevSecOps, incident response, and cyber defense—enabling secure collaboration from tactical edge and DDIL environments to enterprise HQ. Teams operate across web, desktop, and mobile, with embedded interoperability for Microsoft Teams, Outlook, and Microsoft 365. To learn more, visit Mattermost is seeking an experienced and visionary Lead Site Reliability Engineer (SRE) to guide the architecture, reliability, and operational excellence of the infrastructure powering our secure, mission-critical collaboration platform. In this role, you will provide technical leadership across our SRE function, driving strategic initiatives for scalability, observability, performance, and automation across cloud and hybrid environments. You will mentor engineers, establish best practices, and collaborate closely with development, security, and operations teams to ensure our customers in defense, government, and critical infrastructure sectors experience exceptional reliability and performance. Responsibilities Include:
- Define the strategy, architecture, and roadmap for Mattermost’s site reliability engineering function, aligning infrastructure initiatives with product and business goals.
- Lead the design, deployment, and optimization of production-grade containerized workloads, infrastructure-as-code, and compliant cloud environments for regulated domains (e.g., FedRAMP, DoD).
- Establish and evolve observability, monitoring, and alerting frameworks to ensure performance, reliability, and capacity planning at scale.
- Drive incident management processes, including on-call rotations, root cause analysis, and systemic reliability improvements.
- Partner with security and compliance teams to meet data sovereignty, security, and regulatory requirements.
- Champion automation and operational excellence to improve efficiency, reduce risk, and scale operations.
- Oversee cloud cost management and capacity planning to optimize infrastructure spending while meeting performance targets.
- Build and maintain a developer platform that enables fast, secure software delivery and improves application stability in production.
- Mentor and coach SRE team members, fostering a culture of learning, collaboration, and technical excellence.
- BS in Computer Science, Cybersecurity, Software Engineering, or a related technical field, or equivalent experience, with 5+ years of relevant experience in site reliability engineering, DevOps, or cloud infrastructure roles.
- Proven expertise in container orchestration platforms, ideally Kubernetes.
- Extensive experience with infrastructure-as-code, ideally Terraform.
- Strong background in cloud platforms, ideally AWS.
- Demonstrated experience designing and implementing monitoring, alerting, and performance optimization strategies.
- Exceptional troubleshooting and incident management skills for distributed systems.
- Proficiency in at least one scripting or programming language for automation.
- Excellent communication skills with a track record of influencing cross-functional teams.
- Experience leading globally distributed teams in a remote-first environment.
- Familiarity with observability stacks such as Grafana and Prometheus.
- Experience designing high-availability, disaster recovery, and scaling architectures.
- Exposure to GCP and Azure cloud environments.
- Leadership experience in highly regulated industries such as defense, finance, or critical infrastructure.
- Experience with U.S. federal compliance frameworks and authorization processes, including FedRAMP, DoD ATO, NIST 800-53, and related government standards.
- Experience preparing, delivering, and maintaining software offerings through AWS Marketplace and other cloud provider marketplaces (e.g., Azure Marketplace, Google Cloud Marketplace), including packaging, compliance validation, and ongoing operational support.
- Open-source contributions in reliability, DevOps, or infrastructure tooling.
- Certifications in cloud infrastructure, reliability, or DevOps engineering (e.g., CKA, CKAD, AWS Certified Solutions Architect).
- ...This job is responsible for building and leading a team to deliver technology products... ...standards, promoting design, engineering, and organizational practices, and advocating... ...stakeholders.Overview:Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the...SuggestedFull timeWork at officeDay shift
$99k - $225k
Site Reliability Engineer, LeadThe Opportunity: As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and...SuggestedFull timeContract workPart timeWork at officeLocal areaRemote work$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...SuggestedWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week- Job title: Site Reliability Engineer (SRE) Bill rate: $52/hr W2 Client address: 2900 W Plano Pkwy Plano, TX 75075 - Role is hybrid (3 days/wk) Years of experience required: 11+ Mandatory skills: Azure DevOps (ADO), GitHub & GitHub Actions, JFrog Artifactory Site...SuggestedFull time
$350k
...and novel use-cases. We’re hiring to grow the platform alongside the Tinker community. About the Role We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research...SuggestedFull timeVisa sponsorshipWork visaRelocation package$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud... ...concrete engineering and prioritization decisions. Lead incident response and resolution for production issues, acting...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$192k
...here. Role Overview: LeoLabs is seeking a skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will... ...improves deployment reliability. Within 12 months, you’ll: Lead cross-functional initiatives to improve availability,...Full timeWork experience placementRemote workFlexible hours$101.97k - $203.94k
...and one community at a time. Position Summary As a Senior Site Reliability Engineer, you will be responsible for ensuring the stability,... ...Responsibilities Application Performance Monitoring and Observability: Lead the design and implementation of end‑to‑end observability...Hourly payFull timeTemporary workWork experience placementLocal area$139k - $257.55k
...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,... ...productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express,...Full timeTemporary workLocal areaRemote workWorldwide$146.4k
...Communications group. A cross-functional engineering team that develops the... ...system supports fast and reliable configuration of Akamai's... ...metadata systems. As a Senior Site Reliability Engineer, you... ...ESPP). Akamai provides industry-leading benefits including healthcare...Full timeWork experience placementWork at office$180k - $200k
Company Name: tastytrade Role: Senior Site Reliability Engineer Location: Chicago, IL (Hybrid, 3 days/week in office) Role Summary Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options, futures...Full timeWork at office3 days per week- About the Role We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability,... ...critical systems; ensure we meet or exceed targets consistently Lead observability strategy by designing comprehensive...Full time
- ...for current or future sponsorship. Maintain and enhance the reliability, availability, and performance of Navy Federal’s systems and... ...benefits, review the Benefits page [ of the Navy Federal Career Site. Protect Yourself from Job Scams: Navy Federal Credit Union...Full timeInternship
$108k - $180k
...beauty of fashion accessible to all, promoting its industry-leading, on-demand production methodology, for a smarter, future-ready industry. Position Summary We are seeking a Staff Site Reliability Engineer (Official Title: Staff Site Reliability Engineer I) with...Full timeTemporary workWork at officeWorldwideFlexible hours$135k - $155k
...businesses and our colleagues achieve their financial goals. As a leading commercial bank, we remain passionate about serving our... ...opportunities, and enjoy meaningful work! The Director of Site Reliability Engineer is a pivotal technical leader within the Software...Full timeWork experience placementShift workAfternoon shift- ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely... ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale...Permanent employmentWork experience placementWork at officeLocal area
- ...looking for a Senior SRE to join our Platform Engineering team as the operations owner of our... .... You’ll be responsible for the reliability, scalability, and continued evolution of... ...for teams using observability platforms Lead or contribute to platform modernization...Full time
$165k - $270k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...Permanent employmentTemporary workWorldwideWeekend work- ...education. Client is currently seeking a talented Software Engineer who is able to work into the Site Reliability Engineer role. This candidate is expected to work... ...and how to build/utilize (panel of 3-hm, lead, and arch); 2nd round w/ director (panel of 2) Top Must...Remote work
$130k - $200k
IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...Full timeWork at officeImmediate start$130k - $153k
...our customers, and in our growing commitment to land stewardship and recreational access.WHAT YOU WILL DOonX is seeking a Site Reliability Engineer to build and maintain the infrastructure that enables our developers to ship reliably at scale. You'll manage onX's infrastructure...Full timePart timeWork at office$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...Full timeWork at officeLocal areaRemote workWork from home$125k - $185k
...HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions and operations. By bringing... ..., and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance...Full timeWork experience placementWork at officeRemote workWork from homeRelocation package$138.4k - $173k
...infrastructure as well as help improve the reliability, quality of services and overall... ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability... ...about our locations by visiting our site.Compensation & BenefitsThe base salary that...Full timeFlexible hours$152k - $241.5k
...technology—and amazing people. NVIDIA is leading the way in groundbreaking developments... ...automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-... ..., Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through...Full time- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has... ...guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence... ...a diverse team of experts as you use leading-edge tech to empower everyone to meet a...Work at officeLocal area
$145k - $175k
...funding options. Our engaging and rewarding environment is designed to help you gain your full potential. Job OverviewThe Site Reliability Engineer supports deployments, cloud infrastructure, and monitoring systems that power Rewards Network's applications and services....Full timeWork at officeLocal areaFlexible hours3 days per week$80k - $133k
...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and...Permanent employmentFull timeContract workRemote workFlexible hours$165k - $190k
Obsidian Security is the leading SaaS security platform, trusted by global enterprises... ...DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable,... ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate...Work from home$158.5k - $172k
...velocity energy of a powerhouse startup.As a leading U.S. ordering and delivery marketplace,... ....About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will... ...high-impact position driving continuous reliability, deep system optimization, and automation...Full timeWork at office3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead angular developer United States
- lead operating engineer United States
- lead android developer United States
- lead field engineer United States
- lead gameplay engineer United States
- lead project engineer United States
- lead network engineer United States
- lead structural engineer United States
- mainframe lead developer United States
- lead web developer United States
