Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software / Site Reliability Lead Engineer

Full-time

General Dynamics Mission Systems

Role Description

What You Will Own:

  • Cross-pod reliability standards: Set the reliability bar and ensure it is met consistently across applications. Collaborate with Functional SREs to connect technical reliability metrics to business-side outcomes. You own the engineering signal; together you tell the full reliability story.
  • SLOs and reliability metrics: Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.
  • Monitoring and observability: Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards. You will know when something is degrading before users do. Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong. Alerts you build catch real problems; they don't cry wolf.
  • Incident response: Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog. When something breaks, you coordinate the response and ensure it doesn't break the same way again.
  • Production Readiness: Define and enforce the criteria that determine whether an AI service is ready for production. You are the gate between "it works in dev" and "it's ready to ship."
  • Toil elimination: Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.

What You Won't Own:

  • Infrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works.
  • Business process decisions or backlog prioritization.
  • Business-side reliability metrics - you partner with the Functional SRE on those, but they own that domain.

What Makes This Role Different:

  • AI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.
  • You are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform.
  • Your production readiness criteria directly determine whether AI services go live. You have real authority to say "not ready."
  • You operate across projects simultaneously — embedded deeply enough to understand large-scale systems, while maintaining consistent standards across all projects.
  • Your software engineering background means you can engage directly with development teams at the design level — catching reliability problems before they become operational ones.

Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience.
  • Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines.
  • Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.
  • Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar).
  • Experience with containerized environments — Docker, Kubernetes, container orchestration at scale.
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production.
  • U.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.

Requirements

  • Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines.
  • Software engineering fundamentals — you can read, write, and meaningfully review production-quality code. You understand how architectural and design decisions made early translate into operational problems later.
  • Software design experience — you have participated in or led design reviews, defined service interfaces or APIs, and pushed back on design decisions using reliability and operability as criteria.
  • Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that have caught real problems.
  • Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar).
  • Experience with containerized environments — Docker, Kubernetes, container orchestration at scale.
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production.

Benefits

  • Remote — 100% telework.
  • 9/80 schedule.
  • Defense industry experience is not required.

Company Description

General Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!

Vacancy posted 16 days ago
Similar jobs that could be interesting for youBased on the Senior Software / Site Reliability Lead Engineer in Remote vacancy
  • $142.7k - $158.3k

     ...position involves owning the reliability of AI services and ensuring that...  ...and use them to drive engineering decisions. ~Monitoring and...  ...incident management end-to-end. Lead post-incident reviews and maintain...  ...standards. ~Your software engineering background allows... 
    Senior
    Software
    Full time
    Remote work

    General Dynamics Mission Systems, Inc

    Remote
    18 days ago
  • $139k - $257.55k

     ...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through...  ...with the resources of a large software company.What you'll doThis is a role...  ...customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio... 
    Senior
    Software
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    1 day ago
  • $158.5k - $172k

     ...of a powerhouse startup.As a leading U.S. ordering and delivery...  ...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation...  ...position driving continuous reliability, deep system optimization,...  ..., secure, and friction-free software delivery workflows.Secure and... 
    Senior
    Software
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    Chicago, IL
    4 days ago
  • $117k - $209.33k

     ...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build...  ...-facing services.You will combine software engineering and production...  ...operational automation at scaleExperience leading or participating in Gamedays,... 
    Senior
    Software
    Full time
    For contractors
    Remote work

    Autodesk

    Plano, TX
    15 hours ago
  •  ...Grafana Labs is seeking a Staff Software Engineer - SRE to scale Grafana Cloud databases (Mimir, Loki, Tempo, Pyroscope) across AWS, GCP, and Azure. You will own production reliability for high-SLA environments and partner with product engineering squads to deliver reliable... 
    Senior
    Software
    Remote work

    United States Digital Space LLC

    United States
    2 days ago
  •  ...communicator. Expected to actively lead and triage proactively...  ..., My SQL and Mongo DB Seniority level Seniority level Mid-Senior...  ...set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,0...  .../San Diego, CA) Senior Software Engineer - Optical Network... 
    Senior
    Software
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    3 days ago
  •  ...Senior Site Reliability Engineer United Kingdom - Remote At NiCE, we don't limit our challenges. We challenge our limits. Always. We're ambitious...  ...and taking a holistic view of system health Build software and systems to manage platform infrastructure and applications... 
    Senior
    Software
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    NICE

    United States
    1 day ago
  •  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...plane services and dataplane software running on SmartNICsDevelop tooling...  ...teams to improve service reliability and deployment...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production... 
    Senior
    Software
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range...  ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...  ...if youHave 6+ years of experience in software development and operating distributed systemsAre... 
    Senior
    Software
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Chicago, IL
    1 day ago
  • $127k - $249k

     ...zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support,...  ...team works alongside the various Atlas software engineering teams to provide expertise...  ...OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure... 
    Senior
    Software
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    2 days ago
  • $90k - $180k

     ...spans the spectrum of healthcare, with leading businesses and products in...  ...than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar...  ...eliminate performance bottlenecks in software and infrastructure, ensuring low-latency... 
    Senior
    Software
    Remote work

    Abbott

    Sunnyvale, CA
    1 day ago
  • $112.7k - $193.2k

     .... Growing together.We are seeking an experienced Senior Manager to lead enterprise Site Reliability Engineering (SRE), DevOps, IT Service Management (ITSM), and...  ...Technology, or related field10+ years of experience in Software Engineering, Site Reliability Engineering,... 
    Senior
    Software
    Minimum wage
    Full time
    Work experience placement
    Local area
    Remote work

    UnitedHealth Group

    Basking Ridge, NJ
    6 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security...  ...Design and Implementation: Help lead the design and deployment of security...  ...transform, and disrupt industries with software. MongoDB’s unified data platform, the... 
    Senior
    Software
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    4 days ago
  • $118k - $177k

     ...Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely . Everforth ECS is seeking talented professionals to join our...  ...suite of multiple Commercial Off the Shelf (COTS) products, software configuration packages, and custom code which work... 
    Senior
    Software
    Remote work

    ECS

    Fairfax, VA
    2 days ago
  •  ...Senior Site Reliability Engineer AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank...  ...within observability platforms. Lead maintenance efforts and platform improvements... 
    Senior
    Software
    Local area
    Remote work
    Flexible hours

    AgileEngine

    United States
    15 hours ago
  •  ...Job title: Senior Site Reliability Engineer Location: Urbandale IA Duration: 1+ year of contract Job Description: Person will work a split schedule...  ...as Java or Go. (3 - 6 years) Experience in programming/software development (Java (70%), Go (10%), Scala (10%), Python (5... 
    Senior
    Software
    Contract work
    Remote work

    Artech

    Urbandale, IA
    3 days ago
  •  ...Senior Site Reliability Engineer We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be...  ...engineering teams to embed reliability and performance into the software delivery lifecycle. Design, implement, and evolve... 
    Senior
    Software
    Remote work

    Omilia - Conversational Intelligence

    United States
    15 hours ago
  • $15k

     ...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster...  ...mechanisms when off-the-shelf ones won't doHelp software and research teams design policies around fair cluster... 
    Senior
    Software
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    2 days ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...: Less than 25%Profession: Software EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability improvements across... 
    Senior
    Software
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    15 hours ago
  • $185k - $200k

    Back to All JobsSenior Site Reliability Engineer (SRE) Dayton, OH (Remote) full time Top Secret (TS...  ...Overview Metronome is seeking a Senior Site Reliability Engineer (SRE) to support...  ..., cloud/platform engineering, DevOps, software engineering, or a related discipline.... 
    Senior
    Software
    Full time
    Remote work

    Metronome LLC

    Dayton, OH
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Salt Lake City, UT Are you passionate about building highly...  ...You'll work at the intersection of software engineering and infrastructure, partnering...  ...platforms and integrations, leading proof-of-concepts and defining adoption... 
    Senior
    Software
    Work at office
    Remote work
    1 day per week

    PrincePerelson & Associates

    Salt Lake City, UT
    4 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and...  ...secure, repeatable, and reliable software delivery.Observability and...  ...service restoration, and lead incident response when appropriate...  ...more years of experience in Site Reliability Engineering,... 
    Senior
    Software
    Remote work

    Patterson-UTI

    Houston, TX
    2 days ago
  • $150k - $180k

     ...seeking an experienced SeniorSite Reliability Engineer to help design, build,...  ...organization.This position is based on-site in either our Arlington, VA...  ...the team's capabilities.Lead by example in fostering a...  ...in infrastructure and software architecture, capable of designing... 
    Senior
    Software
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    4 days ago
  • $174k - $252k

     ...consulting, developing software platforms and...  ...changes that improve reliability and velocity.Practice...  ...in Computer Science, Engineering, a related field, or equivalent...  ...2 years of experience leading projects and providing...  ...Science or Engineering.Site Reliability Engineering... 
    Senior
    Software

    Google

    Sunnyvale, TX
    4 days ago
  •  ...Senior Site Reliability Engineer Deimos is a cloud-native developer and security operations technology...  ...foundational design and the craft of software engineering. As such our engineers enjoy...  ...or similar). Reliability-as-Code: Lead the drive to manage our entire... 
    Senior
    Software
    Currently hiring
    Remote work
    Work from home

    Deimos

    United States
    15 hours ago
  •  ...Site Reliability Engineer Teikametrics is revolutionizing retail through our patented Artificial Retail Intelligence platform. Our proprietary...  ...DevOps tools and best practices required for efficient software development and deployment. This highly visible role will... 
    Senior
    Software
    Remote work
    Work from home
    Flexible hours

    Teikametrics

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Founded in 2010, Semios Group is a leading agricultural technology company helping growers, agronomists, and ag retailers manage over...  ...an on-call roster. Work with product and software development colleagues to improve the... 
    Senior
    Software
    Remote work
    Work from home

    Semios

    United States
    3 days ago
  • $118.6k - $195.68k

     ...Hat IT OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and...  ...and development of software like Kubernetes operators, webhooks,...  ...Engineering teamsDesign software tests and lead peer reviews to increase the quality... 
    Senior
    Software
    Permanent employment
    Full time
    Contract work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Red Hat

    Raleigh, NC
    2 days ago
  • $139k - $257.55k

     ...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through...  ...with the resources of a large software company. What you'll do This...  ...customer experiences. Adobe's industry-leading offerings including Adobe Acrobat Studio... 
    Senior
    Software
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe

    New York, NY
    1 day ago
  • $134.25k - $214.8k

     ...with our ecosystem of devices and cloud software. Like our products, we work better...  ...where you matter.Your ImpactAre you an engineer who gets excited about the challenge of...  ...of the Observability team within Axon's Site Reliability organization — a focused team responsible... 
    Senior
    Software
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    15 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software / Site Reliability Lead Engineer. Be the first to apply!