Senior Software / Site Reliability Lead Engineer
$142.7k - $158.3kGeneral Dynamics Mission Systems
Role Description
What You Will Own:
- Cross-pod reliability standards: Set the reliability bar and ensure it is met consistently across applications. Collaborate with Functional SREs to connect technical reliability metrics to business-side outcomes. You own the engineering signal; together you tell the full reliability story.
- SLOs and reliability metrics: Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.
- Monitoring and observability: Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards. You will know when something is degrading before users do. Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong. Alerts you build catch real problems; they don't cry wolf.
- Incident response: Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog. When something breaks, you coordinate the response and ensure it doesn't break the same way again.
- Production Readiness: Define and enforce the criteria that determine whether an AI service is ready for production. You are the gate between "it works in dev" and "it's ready to ship."
- Toil elimination: Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.
What You Won't Own:
- Infrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works.
- Business process decisions or backlog prioritization.
- Business-side reliability metrics - you partner with the Functional SRE on those, but they own that domain.
What Makes This Role Different:
- AI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.
- You are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform.
- Your production readiness criteria directly determine whether AI services go live. You have real authority to say "not ready."
- You operate across projects simultaneously — embedded deeply enough to understand large-scale systems, while maintaining consistent standards across all projects.
- Your software engineering background means you can engage directly with development teams at the design level — catching reliability problems before they become operational ones.
Qualifications
- Bachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience.
- Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines.
- Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.
- Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar).
- Experience with containerized environments — Docker, Kubernetes, container orchestration at scale.
- Experience defining and managing SLOs, error budgets, and incident response procedures in production.
- U.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.
Requirements
- Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines.
- Software engineering fundamentals — you can read, write, and meaningfully review production-quality code. You understand how architectural and design decisions made early translate into operational problems later.
- Software design experience — you have participated in or led design reviews, defined service interfaces or APIs, and pushed back on design decisions using reliability and operability as criteria.
- Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that have caught real problems.
- Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar).
- Experience with containerized environments — Docker, Kubernetes, container orchestration at scale.
- Experience defining and managing SLOs, error budgets, and incident response procedures in production.
Benefits
- Remote — 100% telework.
- 9/80 schedule.
- Defense industry experience is not required.
Company Description
General Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!
$142.7k - $158.3k
...position involves owning the reliability of AI services and ensuring that... ...and use them to drive engineering decisions. ~Monitoring and... ...incident management end-to-end. Lead post-incident reviews and maintain... ...standards. ~Your software engineering background allows...SeniorSoftwareFull timeRemote work- Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and... ..., and error budgets.The role combines software engineering, cloud engineering, automation... ...failover, and disaster recovery.Lead incident response efforts, act as an escalation...SeniorSoftwareLocal areaRemote workFlexible hoursShift work
$139k - $257.55k
...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through... ...with the resources of a large software company.What you'll doThis is a role... ...customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio...SeniorSoftwareFull timeTemporary workLocal areaRemote workWorldwide$158.5k - $172k
...of a powerhouse startup.As a leading U.S. ordering and delivery... ...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation... ...position driving continuous reliability, deep system optimization,... ..., secure, and friction-free software delivery workflows.Secure and...SeniorSoftwareFull timeWork at office3 days per week$117k - $209.33k
...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build... ...-facing services.You will combine software engineering and production... ...operational automation at scaleExperience leading or participating in Gamedays,...SeniorSoftwareFull timeFor contractorsRemote work- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...plane services and dataplane software running on SmartNICsDevelop tooling... ...teams to improve service reliability and deployment... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production...SeniorSoftwareWork at officeLocal areaWork from homeFlexible hours
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... ...if youHave 6+ years of experience in software development and operating distributed systemsAre...SeniorSoftwareWork at officeLocal areaRemote workWorldwideFlexible hours$96k - $163k
...greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At... ...during the application build phase in software run principles that include operational... ...capacity planning, and monitoring that leads to fault-tolerant, scalable products....SeniorSoftwareFull timePart timeWorldwideFlexible hours$90k - $180k
...spans the spectrum of healthcare, with leading businesses and products in... ...than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar... ...eliminate performance bottlenecks in software and infrastructure, ensuring low-latency...SeniorSoftwareRemote work$127k - $249k
...zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support,... ...team works alongside the various Atlas software engineering teams to provide expertise... ...OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure...SeniorSoftwareLocal areaRemote workWorldwideFlexible hours$112.7k - $193.2k
.... Growing together.We are seeking an experienced Senior Manager to lead enterprise Site Reliability Engineering (SRE), DevOps, IT Service Management (ITSM), and... ...Technology, or related field10+ years of experience in Software Engineering, Site Reliability Engineering,...SeniorSoftwareMinimum wageFull timeWork experience placementLocal areaRemote work$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security... ...Design and Implementation: Help lead the design and deployment of security... ...transform, and disrupt industries with software. MongoDB’s unified data platform, the...SeniorSoftwareLocal areaRemote workWorldwideFlexible hours$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole... ...: Less than 25%Profession: Software EngineeringDiscipline: Site Reliability EngineeringCompany:... ...most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability improvements across...SeniorSoftwareOngoing contractLocal area3 days per week$15k
...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster... ...mechanisms when off-the-shelf ones won't doHelp software and research teams design policies around fair cluster...SeniorSoftwareWork at officeLocal areaRemote work$150k - $180k
...seeking an experienced SeniorSite Reliability Engineer to help design, build,... ...organization.This position is based on-site in either our Arlington, VA... ...the team's capabilities.Lead by example in fostering a... ...in infrastructure and software architecture, capable of designing...SeniorSoftwarePermanent employmentFull timeWork at officeLocal areaRemote workWorldwide$262k - $364k
...design consulting, developing software platforms and frameworks,... ...pushing for changes that improve reliability and velocity.Practice... ...languages.4 years of experience leading projects, and providing... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE...SeniorSoftware$149.4k - $202k
...Noctua Technology is seeking a Senior Software Engineer specializing in Site Reliability Engineering to join their team. This role focuses on the reliability and performance of cloud-native applications, emphasizing Infrastructure as Code and automation. The ideal candidate...SeniorSoftwareRemote work$118.6k - $195.68k
...Hat IT OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and... ...and development of software like Kubernetes operators, webhooks,... ...Engineering teamsDesign software tests and lead peer reviews to increase the quality...SeniorSoftwarePermanent employmentFull timeContract workWork experience placementWork at officeRemote workFlexible hours$130k - $180k
...collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer... ...process. You’ll engage in and often lead architectural discussions, reduce... ...cloud product. We’re seeking senior software and systems engineers specializing in...SeniorSoftwareWork at officeLocal areaRemote workWorldwideMonday to FridayFlexible hours$96k - $163k
...potential. Title and Summary Senior Site Reliability Engineer Overview The BizOps team is... ...application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead Mastercard in DevOps automation and...SeniorSoftwareFull timePart timeWorldwideFlexible hoursShift work$139k - $257.55k
...Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,... ...we used to run our own operations. ~Partner across software, ML, and platform teams to bake reliability in from the start...SeniorSoftwareFull timeRemote work- ...Essential Functions: Partner with software developers, platform engineers, and IT staff to improve system design... ...requirements, service quality, reliability, security, and compliance needs. Drive... ...: 8+ years of experience in Site Reliability Engineering, DevOps, Platform...SeniorSoftwareWork at officeRemote work
$96k - $163k
...governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer, Performance Engineering Senior Site Reliability... ...from LoadRunner to Gatling or BlazeMeter. • Strong software performance analysis, bottleneck identification, scaling,...SeniorSoftwareFull timePart timeWorldwideFlexible hours- ...physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability,... ...radius and prevent cascading failures.Lead production incident response, postmortems... ..., distributed systems, or production software engineering.Have deep experience...SeniorSoftwareWork at officeLocal areaWork from homeFlexible hours
$134.25k - $214.8k
...safety and justice issues with our ecosystem of devices and cloud software. Like our products, we work better together. We connect with... ...that holds up in court. That's us.Axon's Platform team is the engine behind what hundreds of thousands of officers rely on every day...SeniorSoftwareWork experience placementWork at officeRemote work$267k - $356k
...currently Tuesday.Lambda's Storage Engineering team is the backbone behind... ...the industry, which means reliability and performance aren't just... ...behind Lambda's own software-defined data plane.Build and... ...storage across new and existing sites using tools such as Ansible,...SeniorSoftwareWork experience placementWork at officeLocal areaWork from homeFlexible hours$137.9k - $221.4k
...We’re looking for someone to lead development aspects of the Infrastructure engineering team at ServiceTitan. You must... ...architectural thought process. Our Site Reliability and Infrastructure Engineering... ...issues or configuring software ~Reducing the total cost of...SeniorSoftwareFull timeImmediate startFlexible hours$141k
...year over year, ClickHouse leads the market in real-time analytics... ...our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible... ...You will be leveraging your software engineering expertise to...SeniorSoftwareLocal areaRemote workHome officeFlexible hours- ...secured identity, improving engineering velocity while maintaining security... ...building our production and software as a service infrastructure... ...to trust us for secure and reliable access to their infrastructure... ...with the Hiring Manager or a Lead Engineer. They will walk you...SeniorSoftwareWork at officeLocal areaRemote workSleeping nights
- Edmond, OKYouVersion - YouVersion Engineering /Full-Time/ Salary /On-siteThe YouVersion Senior Site Reliability Engineer is responsible for... ...with the development of software, the performance of regular maintenance... ....Church, our mission is to lead people to become fully...SeniorSoftwareFull timeContract workTemporary workWork experience placementCasual workInternshipLocal areaWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software / Site Reliability Lead Engineer. Be the first to apply!
- software team lead Remote
- software lead Remote
- agile software developer Remote
- software developer internship no experience Remote
- intermediate software engineer Remote
- software engineer staff Remote
- experienced software developer Remote
- work from home software developer Remote
- software developer no experience Remote
- software developer fintech Remote





