Staff Site Reliability Engineer
$163.43k - $213.97kIonQ Inc.
About IonQ: IonQ, Inc. [NYSE: IONQ] is the world's leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ's newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting-edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two-qubit gate fidelity, setting a world record in quantum computing performance. Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Canada, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space. IonQ is making quantum platforms more accessible and impactful than ever before.
Location: Santa Clara, CA
Travel: Up to 25%
Job ID: 1739 The Role: We are seeking a Staff Site Reliability Engineer. As Staff SRE Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing. Responsibilities :
At IonQ, we believe in fair treatment, access, opportunity, and advancement for all while striving to identify and eliminate barriers. We empower employees to thrive by fostering a culture of autonomy, productivity, and respect. We are dedicated to creating an environment where individuals can feel welcomed, respected, supported, and valued.
We are committed to equity and justice. We welcome different voices and viewpoints and do not discriminate on the basis of race, religion, ancestry, physical and/or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, transgender status, age, sexual orientation, military or veteran status, or any other basis protected by law. We are proud to be an Equal Employment Opportunity employer. US Technical Jobs. The position you are applying for will require access to technology that is subject to U.S. export control and government contract restrictions. Employment with IonQ is contingent on either verifying "U.S. Person" (e.g., U.S. citizen, U.S. national, U.S. permanent resident, or lawfully admitted into the U.S. as a refugee or granted asylum) status for export controls and government contracts work, obtaining any necessary license, and/or confirming the availability of a license exception under U.S. export controls. Please note that in the absence of confirming you are a U.S. Person for export control and government contracts work purposes, IonQ may choose not to apply for a license or decline to use a license exception (if available) for you to access export-controlled technology that may require authorization, and similarly, you may not qualify for government contracts work that requires U.S. Persons, and IonQ may decline to proceed with your application on those bases alone. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law.
US Non-Technical Jobs. Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law. If you are interested in being a part of our team and mission, we encourage you to apply!
Location: Santa Clara, CA
Travel: Up to 25%
Job ID: 1739 The Role: We are seeking a Staff Site Reliability Engineer. As Staff SRE Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing. Responsibilities :
- Production reliability: own service-level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions.
- Engineer observability: design and operate the observability stack so production services are fully instrumented and define the standards platform and application teams follow.
- Govern SLOs and error budgets: define and manage service-level objectives, run regular reviews with service owners, and drive corrective action when services consume error budgets unsafely.
- Drive resilience: design and execute chaos experiments and validate that failure modes are covered by tested safeguards.
- Lead incident response: define the incident process and serve as incident commander for the highest-severity incidents, including security incidents within the coverage window.
- Run on-call and escalation: establish and manage rotations and escalation paths that provide continuous coverage with clean follow-the-sun handoffs.
- Disaster recovery: own disaster-recovery testing and failover validation against defined recovery objectives and turn exercise findings into architectural and operational improvements.
- Cloud security posture: co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring with DevSecOps.
- Data, streaming, and AI Ops: own reliability of stateful and streaming services, capacity planning and rightsizing, and autonomous agents for triage, predictive alerting, remediation, and self-healing.
- Scale the team and broaden impact: mentor engineers at different seniority levels, set standards adopted across teams, and align Architecture, DevSecOps, Cloud Operations, and Product Development behind a shared reliability roadmap.
- 7+ years of production engineering experience with recent hands-on reliability work.
- Hands-on, recent experience operating large-scale, fault-tolerant production systems on AWS or GCP.
- Observability ownership: have instrumented production systems and governed service-level objectives and error budgets, not only installed dashboards.
- Resilience practice: have designed and executed failure experiments or disaster-recovery exercises with real failover validation.
- Incident command: have personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix.
- Demonstrated ownership of reliability outcomes with measurable results, such as availability, mean time to recovery, and error-budget adherence.
- Evidence of multi-team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team.
- Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments.
- Strong experience prioritizing risk using identity, workload, and exposure-path context to focus remediation on issues that materially increase attack likelihood and operational impact.
- Experience with autonomous remediation and self-healing workflows powered by AIOps, including Amazon Bedrock Agent Core or equivalent agentic automation frameworks.
- Hands-on experience in capacity management, resource rightsizing, efficiency engineering, and practical cost optimization based on FinOps principles.
- Experience with load-balancing design and operations, including health-based failover, global traffic management, and performance optimization for highly available services.
- Experience with AI traffic management via an LLM gateway, including request routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and observability for multi-model or multi-provider environments.
- Ability to connect networking, security, and reliability considerations into cohesive platform design decisions that improve resilience, performance, and operability.
At IonQ, we believe in fair treatment, access, opportunity, and advancement for all while striving to identify and eliminate barriers. We empower employees to thrive by fostering a culture of autonomy, productivity, and respect. We are dedicated to creating an environment where individuals can feel welcomed, respected, supported, and valued.
We are committed to equity and justice. We welcome different voices and viewpoints and do not discriminate on the basis of race, religion, ancestry, physical and/or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, transgender status, age, sexual orientation, military or veteran status, or any other basis protected by law. We are proud to be an Equal Employment Opportunity employer. US Technical Jobs. The position you are applying for will require access to technology that is subject to U.S. export control and government contract restrictions. Employment with IonQ is contingent on either verifying "U.S. Person" (e.g., U.S. citizen, U.S. national, U.S. permanent resident, or lawfully admitted into the U.S. as a refugee or granted asylum) status for export controls and government contracts work, obtaining any necessary license, and/or confirming the availability of a license exception under U.S. export controls. Please note that in the absence of confirming you are a U.S. Person for export control and government contracts work purposes, IonQ may choose not to apply for a license or decline to use a license exception (if available) for you to access export-controlled technology that may require authorization, and similarly, you may not qualify for government contracts work that requires U.S. Persons, and IonQ may decline to proceed with your application on those bases alone. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law.
US Non-Technical Jobs. Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law. If you are interested in being a part of our team and mission, we encourage you to apply!
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Santa Clara, CA vacancy
$101k - $161k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-... ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-...Suggested$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...SuggestedNight shift$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...SuggestedFull timeWorldwide$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...SuggestedFull time- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...SuggestedFlexible hours
$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...Full time- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...Full timeWork at office2 days per week
- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...Work at officeLocal areaWork from homeFlexible hours
$90k - $180k
...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are...Remote work$186.9k - $267.7k
...requiring approximately 2 days per week on-site at Cisco offices in either San... ...behave as intended, improving reliability and reducing risks. This unified approach... ...enhanced observability and control.As a Staff Site Reliability Engineer (SRE), you will provide technical...Full timeTemporary workLocal areaFlexible hours2 days per week$207k - $301k
Develop strong, influential relationships with multiple stakeholders across the Site Reliability Engineering and Developer organizations.Serve as an expert on particular fields of knowledge related to rate limiting or sharding.Develop plans and lead projects on evolving...$210.6k - $305.1k
...Minimum Qualifications: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Full timeTemporary workLocal areaFlexible hours$207.4k - $259.2k
...differences, and supports and celebrates all of our team members.We are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you will be responsible for the reliability, scalability,...Permanent employmentLocal area$207k - $301k
...implementation of solutions to enhance the reliability of systems that support F1.Scale systems... ...for multiple teams.Engage in software engineering on services written in Java, C++, and Go... ...related technical field.Experience in a Site Reliability Engineering role.Experience...$262k - $365k
...automation, and evolve systems by pushing for changes that improve reliability and velocity.Practice sustainable incident response and... ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems...$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation...Full time$122.5k - $175k
...age, we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief...Full timeWork at officeLocal area3 days per week- ...powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base...Work at officeLocal areaWork from homeFlexible hours
$174k - $253k
...MINIMUM QUALIFICATIONS: Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5... ...s degree in Computer Science or Engineering. ABOUT THE JOB: Site Reliability Engineering (SRE) is what you get when you treat operations...$145k - $165k
...Your Ego : Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to...Work at officeImmediate start- ...Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots. FII provides customers with intelligent...Permanent employmentFull timeWork at officeLocal area
- ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud...Full time
- ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable...
$150k - $195k
...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the... ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation....Full timeWorldwide$64 - $68 per hour
...Akkodis is seeking a Site Reliability Engineer for a Contract with a client in Sunnyvale, CA/Austin, TX (Hybrid). The ideal candidate with experience maintaining highly available, scalable cloud infrastructure and driving operational excellence through automation...Hourly payContract workTemporary workLocal area$185k
...professionals. If the opportunity to build your career is compelling, read on for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance ofJuul’s hybrid cloud infrastructure (Nutanix, AWS/GCP)...Remote work- ...Site Reliability Engineer Location – San Jose, CA What You'll Do - Responsibilities Engage in and improve the whole lifecycle of services—from inception and design, through automated deployment, operation and refinement. Work with all relative teams to make...
$187.04k - $359.72k
...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum... ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas.... ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company...Temporary workLocal areaOverseasShift work$145k - $165k
...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...- ...that keep the world running. Location: 5 On-Site Days a Week in Sunnyvale, CA Headquarters Our Engineering team is driven by a culture that thrives on visionary... ...to-day basis, you will work on enhancing system reliability and scalability of Illumio SaaS products, and...Work experience placementImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
Related searches
- engineering aide Santa Clara, CA
- technology administrator Santa Clara, CA
- senior staff engineer Santa Clara, CA
- staff engineer Santa Clara, CA
- senior staff systems engineer Santa Clara, CA
- assistant engineer Santa Clara, CA
- software engineer staff Santa Clara, CA
- site reliability engineer Santa Clara, CA
- site reliability engineer sre Santa Clara, CA
- construction site safety Santa Clara, CA

