Staff Site Reliability Engineer
$191k - $253kAI Chopping Block
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting‑edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.
ABOUT THE TEAM:
CorpTech Platform is the internal engineering force multiplier behind Anduril’s corporate systems. It drives CorpTech engineering forward through strategic investment in data platforms, software platforms, QA and release excellence, ERP engineering, and AI infrastructure.
We build the foundations that power both CorpOS, enabling Finance and Growth to operate with speed and precision, and ArsenalOS, the digital backbone of Anduril’s hardware enterprise. By unifying systems, accelerating development velocity, and embedding intelligence into every layer, CorpTech Platform transforms how the company operates end to end.
We are also driving Anduril toward becoming an autonomous enterprise. Through initiatives like the Autonomous Software Factory, we are rethinking how software is built, tested, deployed, and evolved by integrating AI directly into the engineering lifecycle. CorpTech Platform is how Anduril scales its business systems with the same rigor, speed, and adaptability as its mission critical products.
ABOUT THE JOB:
This Staff SRE role sets the reliability architecture for the systems that run Anduril’s business and manufacturing operations. You are not responding to tickets. You are the person who decides how production works: how services are observed, how deployments are safe, how incidents are managed, and how reliability scales as the platform grows. You own the infrastructure and mechanisms that make reliability a property of the platform rather than a function of individual heroics.
The role carries broad organizational reach. You will influence how software engineering teams across CorpTech Platform build and operate their systems, set the standards that govern production readiness, and make infrastructure decisions that affect every service in the portfolio. You bring both deep systems expertise and the judgment to know which reliability investments matter most at each stage of the platform’s growth.
WHAT YOU"LL DO:
- Set the reliability architecture for CorpTech Platform’s production environment, including observability infrastructure, deployment systems, incident management, capacity planning, and failure‑domain isolation.
- Design and operate the observability platform, including metrics, distributed tracing, structured logging, alerting, and dashboarding infrastructure that gives engineering teams real production visibility at scale.
- Own deployment infrastructure and release‑safety mechanisms, including progressive rollout systems, canary analysis, automated rollback, and deployment gates that let the organization ship fast without gambling on production stability.
- Define and govern SLO frameworks that make reliability measurable and actionable, creating shared language between SRE, product engineering, and leadership for making trade‑off decisions.
- Identify systemic reliability risks across the platform and drive infrastructure or automation investments that eliminate entire failure classes rather than patching individual symptoms.
- Establish production‑readiness standards and review processes that engineering teams adopt during design and pre‑launch, embedding reliability into the development lifecycle rather than bolting it on after release.
- Lead incident response for complex, multi‑system failures and drive post‑incident processes that produce durable systemic improvements rather than shallow fixes.
- Develop reliability patterns for AI‑enabled systems, including monitoring for model behavior drift, non‑deterministic output degradation, and graceful fallback under novel failure modes.
- Provide technical direction to SRE and infrastructure engineers, facilitate cross‑team architecture decisions, and represent reliability concerns in platform‑level planning and prioritization.
- Drive capacity planning, cost optimization, and performance engineering for the platform’s critical paths and shared infrastructure.
REQUIRED QUALIFICATIONS:
- 10+ years of experience in site reliability engineering, production engineering, infrastructure engineering, or a closely related discipline, including experience operating at architecture or platform‑wide scope.
- Demonstrated experience designing and owning reliability infrastructure that multiple engineering teams depend on, including observability platforms, deployment systems, or incident management tooling at meaningful scale.
- Deep technical fluency across distributed systems, container orchestration (Kubernetes), cloud platforms (AWS, GCP, or Azure), networking, storage, and the failure modes specific to each layer.
- Proficiency in systems programming languages (Go, Python, Rust, or equivalent) used for building production infrastructure, tooling, and automation.
- Demonstrated experience defining SRE standards, production‑readiness frameworks, or operational maturity models and influencing adoption across engineering teams without formal authority.
- Track record of leading incident response for complex, multi‑system failures and converting post‑incident findings into infrastructure investments that prevent recurrence.
- Proven ability to work independently through ambiguity, define reliability strategy without top‑down direction, and sequence investments against competing organizational demands.
- Excellent written and verbal communication skills, with the ability to influence senior engineers, engineering leadership, and cross‑functional partners on infrastructure and reliability decisions.
- Degree in Computer Science, Information Systems, Engineering, or related technical field, or equivalent practical experience.
- U.S. Person status is required as this position needs to access export controlled data.
PREFERRED QUALIFICATIONS:
- Experience designing or operating observability platforms (Datadog, Grafana, Prometheus, OpenTelemetry) at organizational scale, including decisions about instrumentation standards, data pipelines, retention, and cost management.
- Experience building deployment and release‑safety infrastructure, including canary systems, progressive delivery, feature flagging, and automated rollback mechanisms.
- Experience with reliability engineering for AI‑enabled or ML‑serving systems, including monitoring non‑deterministic behavior, building evaluation‑driven alerts, and designing graceful degradation for model failures.
- Familiarity with frontier AI tooling, AI coding assistants, and AI‑enabled software development workflows.
- Experience in hyper growth startup‑like environments, with demonstrated success scaling reliability practices alongside rapid engineering growth.
- Familiarity with enterprise systems and business process domains such as ERP, MES, WMS, CRM, finance systems, or manufacturing systems.
- Eligible to obtain and maintain a U.S. Secret security clearance.
US Salary Range $191,000 — $253,000 USD
The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top‑tier benefits for full‑time employees, including:
Benefits
At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.
Protecting Yourself from Recruitment Scams
Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information.
To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind:
- No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates.
- Please always verify communications:
- Direct from Anduril: If you receive an email from one of our recruiters, it will only come from an @anduril.com address.
- Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to View email address on click.appcast.io.
- Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain is @anduril.com before providing any personal information or clicking on links.
- What to Do If You Suspect Fraud: Should you encounter any questionable or fraudulent outreach claiming to be from Anduril, please report it immediately to View email address on click.appcast.io. Your proactive caution is invaluable in protecting your personal information and upholding the security and trustworthiness of our recruitment efforts.
Data Privacy
To view Anduril's candidate data privacy policy, please visit
By submitting your application, you consent to Anduril Industries using a third‑party service provider to conduct pre‑employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third‑party service provider provides risk‑intelligence services that may include analysis of sanctions and watchlists, adverse media, public‑record information, and other lawful open‑source or commercial data sources. This third‑party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.
#J-18808-Ljbffr$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...such as Python, Go, Perl, or Ruby. ~ Mentored other engineers and influenced technical direction through design reviews, architecture...Suggested- ...Early Warning Services LLC in Scottsdale seeks a Principal Site Reliability Engineer to design high-availability systems and scale microservice architectures. You will partner with development teams to implement observability, automation, and resilient deployment patterns...Suggested
$182.8k - $247.3k
...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...SuggestedWork experience placement$180k - $270k
...AWS DevOps / Agentic SRE Engineer (Active TS/SCI Clearance) ~ Chantilly, Virginia... ...Release Engineer to own the deployment, reliability, observability, and operation of secure... ...engineering position combining DevOps, Site Reliability Engineering, platform engineering...SuggestedContract workTemporary workFor contractorsLocal areaRemote workFlexible hours$194k - $267k
...Staff Site Reliability Engineer - Kubernetes Important: if an employer asks you to log into their system via iCloud or Google, send a code, an SMS or Telegram password, run some code, or install software — refuse. These are signs of fraud. Secure Every Identity,...SuggestedPermanent employmentWork at officeLocal areaWorldwideFlexible hours$194k - $237k
## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on: Posted 5 Days Agojob requisition id: REQ2026... ...sponsorship.**Overall Purpose**The Principal Site Reliability Engineer partners with development teams by designing availability and...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$181k - $265k
...we’re excited to hire a high performing Staff SRE to join our growing Platform team. In... .... You will collaborate closely with engineering leadership, product managers, and cross-... ...Understand the importance of performant and reliable systems Education - Ideally looking...Work at officeImmediate start3 days per week- ...NVIDIA Corporation is seeking a Sr. Systems Software Engineer to build platform software around open-source container runtimes and Kubernetes. You will contribute to GPU-accelerated applications, work on Kubernetes integration, and collaborate across NVIDIA teams to...
- ...NVIDIA Corporation is seeking a Senior System Software Engineer for the Autonomous Vehicles Platform in Santa Clara, CA. You will develop high-performance, scalable system software for data collection and fleets of autonomous vehicles, interfacing with sensors and heterogeneous...
$136k - $218.5k
## Systems Quality and Reliability Engineer - LPUApplylocations: US, CA, Santa Claratime type: Full timeposted on: Posted Todayjob requisition id: JR2018660We are seeking Systems Quality and Reliability Engineer to join our LPU team!NVIDIA has continuously reinvented itself...$185k - $260k
...LinkedIn Top Startups. Your Role As a Senior Software Engineer on the Developer Platform team, you will be responsible for designing... ...initiatives to ensure that new features are performant and reliable, creating innovative frameworks to address our unique testing...Immediate startRemote workHome officeFlexible hours- ...Angell West is seeking a full time or part time Staff Cardiologist Veterinarian – for our busy and growing hospital located in Waltham... ...with cages for continuous oxygen supply, and we maintain an on-site blood bank. Labwork and samples are sent twice daily to the fully...Full timePart timeInternshipRelocation packageFlexible hours
- ...Sprig is hiring for a Platform Build Engineer to lead the next era of fast, reliable CI/CD and hermetic builds for a massive TypeScript/React/Go monorepo. You will reduce build times, make tests reliable, and create ephemeral environments per task to accelerate developer...
- ...Notion is seeking a Developer Platform engineer to build tools, APIs, and platform experiences that connect Notion to the world — bringing... ...work will enable developers, admins, and builders to create reliable integrations and automations on top of Notion. You’ll contribute...
- ...Vercel is seeking a Staff level Engineer to lead the Finance Data Platform, powering revenue reporting and compliance across the business. Define the technical vision and build auditable pipelines with Kafka, ClickHouse, Tinybird, and Snowflake. You will mentor engineers...
- ...Trustly, Inc. is looking for a Staff AI Enablement Engineer to shape how teams across the company adopt AI at scale. You will design internal platforms... ...ability, and a passion for turning prototypes into reliable production systems. You’ll partner with product and engineering...
$228.5k
...mission-driven Senior Manager, Platform Engineering who believes that great products emerge... ..., data-driven systems that need to work reliably during critical moments, scale as our partners... ...a culture of care to ensure that our staff are best equipped to lead happy, healthy...Full timeTemporary workRemote workHome officeFlexible hoursShift work- ...EvenUp Inc is hiring a Senior Platform Foundation Engineer to lead the architecture of core platform capabilities powering the whole system. You will own the stack from design through deployment, enabling other teams with APIs, data exports, and streaming capabilities...
$200k - $250k
...companies, bringing together early-stage engineers, product builders, and business athletes... ...can hold your own in a conversation with a staff engineer about API architecture, webhooks... ...Gigs Republic, our bi-annual company off-site . Our offices are designed to feel like...Work at officeRemote workWork from homeRelocationHome office- ...Yahoo Inc. is seeking a Partner Success Engineer to own end-to-end partner integrations for Yahoo Mail’s data platform, ensuring API... ...collaboration with enterprise teams and executives. You will drive reliability, specify data contracts, and mentor teams while shaping...
$200k - $240k
...is today’s application monitoring standard and our team is building its AI-native future. About the role: The Solutions Engineering team at Sentry is responsible for helping our largest customers successfully embed Sentry into their applications and workflows,...Hourly payFlexible hours$200k - $250k
...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to... ...it out on their hardest use case, and get agents running reliably at scale. This is a hands-on, highly technical team. Solutions...Work at officeFlexible hours- ...Databricks is hiring an AI Forward Deployed Engineer (FDE) to design, build, and productionize GenAI applications for the U.S. federal sector... ...Washington, D.C./Maryland/Virginia metro area with periodic on-site work. Remote work is possible. You will own end-to-end GenAI...Remote work
- ...Cencora seeks a Forward Deployed Engineer to embed with business units and deliver production-grade Generative AI solutions in days to weeks... ...engineering, solution architecture, and consulting, with on-site engagement as needed and a strong emphasis on delivering value quickly...
- ...Voyager Technologies in El Segundo, CA is seeking a Senior Software Engineer to join a multidisciplinary team developing geospatial analytics software focused on SAR imagery. You will architect and lead production-ready data systems handling large geospatial datasets....
- ...TikTok is hiring a Software Engineer Graduate (Recommendation Infrastructure) for a 2027 start in San Jose. You will join the Recommendation Architecture Team to design and optimize the recommendation system, building high-performance online services and data pipelines...
$186.07k - $218.9k
...surges.” learn more about working at Coinbase. Senior Software Engineer (EAA) The EAA Compliance CXAE team, part of Coinbase's... ...capabilities that benefit multiple teams across the organization. Own reliability for Tier‑1 compliance systems by anticipating potential issues...Local area$195k - $255k
...later without any hidden fees or compounding interest. The Money Movement & Card Ledger team is looking for a passionate Software Engineer to help build the tools and systems that we use to manage money movement, bank data integration & merchant data. Affi m Card is...Work at officeRemote workFlexible hours- ...enterprise technologies. You will collaborate with architects, systems engineers, and mission stakeholders while helping influence technical... ...issues Optimize applications for performance, scalability, reliability, and maintainability Support automated testing and software...
$180k - $270k
...enabling fast, natural, and truly conversational interaction with diverse and creative characters. We’re looking for a Software Engineer to help improve the speech, audio, and media systems at the heart of the Cantina experience. This team’s responsibilities...Work at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- construction site safety Eastern, KY
- on-site clinical research associate (traveling/remote) Eastern, KY
- historic site Eastern, KY
- official site Eastern, KY
- site safety Eastern, KY
- staff automation engineer
- project engineer assistant project manager
- engineering aide
- staff chemical engineer
- staff design engineer

