Senior Software / Site Reliability Lead Engineer
General Dynamics Mission Systems
Role Description
What You Will Own:
- Cross-pod reliability standards: Set the reliability bar and ensure it is met consistently across applications. Collaborate with Functional SREs to connect technical reliability metrics to business-side outcomes. You own the engineering signal; together you tell the full reliability story.
- SLOs and reliability metrics: Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.
- Monitoring and observability: Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards. You will know when something is degrading before users do. Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong. Alerts you build catch real problems; they don't cry wolf.
- Incident response: Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog. When something breaks, you coordinate the response and ensure it doesn't break the same way again.
- Production Readiness: Define and enforce the criteria that determine whether an AI service is ready for production. You are the gate between "it works in dev" and "it's ready to ship."
- Toil elimination: Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.
What You Won't Own:
- Infrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works.
- Business process decisions or backlog prioritization.
- Business-side reliability metrics - you partner with the Functional SRE on those, but they own that domain.
What Makes This Role Different:
- AI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.
- You are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform.
- Your production readiness criteria directly determine whether AI services go live. You have real authority to say "not ready."
- You operate across projects simultaneously — embedded deeply enough to understand large-scale systems, while maintaining consistent standards across all projects.
- Your software engineering background means you can engage directly with development teams at the design level — catching reliability problems before they become operational ones.
Qualifications
- Bachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience.
- Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines.
- Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.
- Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar).
- Experience with containerized environments — Docker, Kubernetes, container orchestration at scale.
- Experience defining and managing SLOs, error budgets, and incident response procedures in production.
- U.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.
Requirements
- Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines.
- Software engineering fundamentals — you can read, write, and meaningfully review production-quality code. You understand how architectural and design decisions made early translate into operational problems later.
- Software design experience — you have participated in or led design reviews, defined service interfaces or APIs, and pushed back on design decisions using reliability and operability as criteria.
- Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that have caught real problems.
- Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar).
- Experience with containerized environments — Docker, Kubernetes, container orchestration at scale.
- Experience defining and managing SLOs, error budgets, and incident response procedures in production.
Benefits
- Remote — 100% telework.
- 9/80 schedule.
- Defense industry experience is not required.
Company Description
General Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!
$142.7k - $158.3k
...position involves owning the reliability of AI services and ensuring that... ...and use them to drive engineering decisions. ~Monitoring and... ...incident management end-to-end. Lead post-incident reviews and maintain... ...standards. ~Your software engineering background allows...SeniorSoftwareFull timeRemote work$139k - $257.55k
...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through... ...with the resources of a large software company.What you'll doThis is a role... ...customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio...SeniorSoftwareFull timeTemporary workLocal areaRemote workWorldwide$158.5k - $172k
...of a powerhouse startup.As a leading U.S. ordering and delivery... ...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation... ...position driving continuous reliability, deep system optimization,... ..., secure, and friction-free software delivery workflows.Secure and...SeniorSoftwareFull timeTemporary workWork at officeFlexible hours3 days per week$117k - $209.33k
...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build... ...-facing services.You will combine software engineering and production... ...operational automation at scaleExperience leading or participating in Gamedays,...SeniorSoftwareFull timeFor contractorsRemote work- ...Grafana Labs is seeking a Staff Software Engineer - SRE to scale Grafana Cloud databases (Mimir, Loki, Tempo, Pyroscope) across AWS, GCP, and Azure. You will own production reliability for high-SLA environments and partner with product engineering squads to deliver reliable...SeniorSoftwareRemote work
- ...communicator. Expected to actively lead and triage proactively... ..., My SQL and Mongo DB Seniority level Seniority level Mid-Senior... ...set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,0... .../San Diego, CA) Senior Software Engineer - Optical Network...SeniorSoftwareContract workRemote work
- ...Senior Site Reliability Engineer United Kingdom - Remote At NiCE, we don't limit our challenges. We challenge our limits. Always. We're ambitious... ...and taking a holistic view of system health Build software and systems to manage platform infrastructure and applications...SeniorSoftwareWork experience placementWork at officeRemote workFlexible hours
- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...plane services and dataplane software running on SmartNICsDevelop tooling... ...teams to improve service reliability and deployment... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production...SeniorSoftwareWork at officeLocal areaWork from homeFlexible hours
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... ...if youHave 6+ years of experience in software development and operating distributed systemsAre...SeniorSoftwareWork at officeLocal areaRemote workWorldwideFlexible hours$127k - $249k
...zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support,... ...team works alongside the various Atlas software engineering teams to provide expertise... ...OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure...SeniorSoftwareLocal areaRemote workWorldwideFlexible hours$90k - $180k
...spans the spectrum of healthcare, with leading businesses and products in... ...than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar... ...eliminate performance bottlenecks in software and infrastructure, ensuring low-latency...SeniorSoftwareRemote work$112.7k - $193.2k
.... Growing together.We are seeking an experienced Senior Manager to lead enterprise Site Reliability Engineering (SRE), DevOps, IT Service Management (ITSM), and... ...Technology, or related field10+ years of experience in Software Engineering, Site Reliability Engineering,...SeniorSoftwareMinimum wageFull timeWork experience placementLocal areaRemote work$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security... ...Design and Implementation: Help lead the design and deployment of security... ...transform, and disrupt industries with software. MongoDB’s unified data platform, the...SeniorSoftwareLocal areaRemote workWorldwideFlexible hours$118k - $177k
...Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely . Everforth ECS is seeking talented professionals to join our... ...suite of multiple Commercial Off the Shelf (COTS) products, software configuration packages, and custom code which work...SeniorSoftwareRemote work- ...Senior Site Reliability Engineer AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank... ...within observability platforms. Lead maintenance efforts and platform improvements...SeniorSoftwareLocal areaRemote workFlexible hours
- ...Job title: Senior Site Reliability Engineer Location: Urbandale IA Duration: 1+ year of contract Job Description: Person will work a split schedule... ...as Java or Go. (3 - 6 years) Experience in programming/software development (Java (70%), Go (10%), Scala (10%), Python (5...SeniorSoftwareContract workRemote work
- ...Senior Site Reliability Engineer We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be... ...engineering teams to embed reliability and performance into the software delivery lifecycle. Design, implement, and evolve...SeniorSoftwareRemote work
$15k
...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster... ...mechanisms when off-the-shelf ones won't doHelp software and research teams design policies around fair cluster...SeniorSoftwareWork at officeLocal areaRemote work$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole... ...: Less than 25%Profession: Software EngineeringDiscipline: Site Reliability EngineeringCompany:... ...most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability improvements across...SeniorSoftwareOngoing contractLocal area3 days per week$185k - $200k
Back to All JobsSenior Site Reliability Engineer (SRE) Dayton, OH (Remote) full time Top Secret (TS... ...Overview Metronome is seeking a Senior Site Reliability Engineer (SRE) to support... ..., cloud/platform engineering, DevOps, software engineering, or a related discipline....SeniorSoftwareFull timeRemote work- ...Senior Site Reliability Engineer (SRE) Salt Lake City, UT Are you passionate about building highly... ...You'll work at the intersection of software engineering and infrastructure, partnering... ...platforms and integrations, leading proof-of-concepts and defining adoption...SeniorSoftwareWork at officeRemote work1 day per week
- Reliability Engineering Design, implement, and operate scalable, resilient, and... ...secure, repeatable, and reliable software delivery.Observability and... ...service restoration, and lead incident response when appropriate... ...more years of experience in Site Reliability Engineering,...SeniorSoftwareRemote work
$150k - $180k
...seeking an experienced SeniorSite Reliability Engineer to help design, build,... ...organization.This position is based on-site in either our Arlington, VA... ...the team's capabilities.Lead by example in fostering a... ...in infrastructure and software architecture, capable of designing...SeniorSoftwarePermanent employmentFull timeWork at officeLocal areaRemote workWorldwide$174k - $252k
...consulting, developing software platforms and... ...changes that improve reliability and velocity.Practice... ...in Computer Science, Engineering, a related field, or equivalent... ...2 years of experience leading projects and providing... ...Science or Engineering.Site Reliability Engineering...SeniorSoftware- ...Senior Site Reliability Engineer Deimos is a cloud-native developer and security operations technology... ...foundational design and the craft of software engineering. As such our engineers enjoy... ...or similar). Reliability-as-Code: Lead the drive to manage our entire...SeniorSoftwareCurrently hiringRemote workWork from home
- ...Site Reliability Engineer Teikametrics is revolutionizing retail through our patented Artificial Retail Intelligence platform. Our proprietary... ...DevOps tools and best practices required for efficient software development and deployment. This highly visible role will...SeniorSoftwareRemote workWork from homeFlexible hours
- ...Senior Site Reliability Engineer (SRE) Founded in 2010, Semios Group is a leading agricultural technology company helping growers, agronomists, and ag retailers manage over... ...an on-call roster. Work with product and software development colleagues to improve the...SeniorSoftwareRemote workWork from home
$118.6k - $195.68k
...Hat IT OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and... ...and development of software like Kubernetes operators, webhooks,... ...Engineering teamsDesign software tests and lead peer reviews to increase the quality...SeniorSoftwarePermanent employmentFull timeContract workWork experience placementWork at officeRemote workFlexible hours$139k - $257.55k
...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through... ...with the resources of a large software company. What you'll do This... ...customer experiences. Adobe's industry-leading offerings including Adobe Acrobat Studio...SeniorSoftwareTemporary workLocal areaRemote workWorldwide$134.25k - $214.8k
...with our ecosystem of devices and cloud software. Like our products, we work better... ...where you matter.Your ImpactAre you an engineer who gets excited about the challenge of... ...of the Observability team within Axon's Site Reliability organization — a focused team responsible...SeniorSoftwareWork experience placementWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software / Site Reliability Lead Engineer. Be the first to apply!
- software team lead Remote
- software lead Remote
- cybersecurity software engineer Remote
- graduate software engineer Remote
- software developer fintech Remote
- new graduate software engineer Remote
- senior robotics software engineer Remote
- software engineer visa sponsorship Remote
- software engineer unity Remote
- software qa engineer Remote


