Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer

$152k - $228k

NightDragon Acquisition Corp.

About IonQ:

IonQ, Inc. [NYSE: IONQ] is the world’s leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ’s newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting-edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two-qubit gate fidelity, setting a world record in quantum computing performance. Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Canada, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space. IonQ is making quantum platforms more accessible and impactful than ever before.

Location: This role is based at our Santa Clara, CA office, with the option to work a few days a week remotely.
Travel: Up to 25%
Job ID: 1874

The Role

The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites.

The Site Reliability Engineering discipline keeps the platform stable and reliable, with a strong focus on service continuity and customer experience. It owns production reliability, service-level objectives, observability architecture, backup and disaster recovery, incident response, and resilience, and co-owns cloud security posture and runtime vulnerability management with DevSecOps.

As Staff Site Reliability Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing.

The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.

  • Production reliability, SLOs, and error budgets — the reliability of production services end to end, including standards, governance, and escalation for Tier-1 and Tier-2 services.
  • Observability architecture and standards — metrics, logs, distributed tracing, and profiles instrumented across production systems, with consistent platform-wide standards.
  • Chaos engineering and resilience — failure-injection experiments and validation of recovery mechanisms in pre-production and production environments.
  • Backup and disaster recovery — backup validation, disaster-recovery architecture, failover testing, and recovery verification against defined RTO and RPO objectives.
  • Cloud security posture — cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring, co-owned with DevSecOps.
  • Data and streaming platform reliability — reliability engineering for Postgres, Redis/Valkey, Kafka, OpenSearch, and other critical stateful services.
  • Capacity, efficiency, and AI Ops — resource rightsizing, predictive alerting, autonomous triage, remediation automation, and self-healing workflows.
  • Incident response and command — severity classification, incident command, executive communication, and blameless post-incident review for the highest-severity events.
  • On-call and escalation — rotation design, operational readiness, escalation policy, and clean follow-the-sun handoffs across regions.
Responsibilities
  • Production reliability — own service-level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions.
  • Engineer observability — design and operate the observability stack so production services are fully instrumented and define the standards platform and application teams follow.
  • Govern SLOs and error budgets — define and manage service-level objectives, run regular reviews with service owners, and drive corrective action when services consume error budgets unsafely.
  • Drive resilience — design and execute chaos experiments and validate that failure modes are covered by tested safeguards.
  • Lead incident response — define the incident process and serve as incident commander for the highest-severity incidents, including security incidents within the coverage window.
  • Run on-call and escalation — establish and manage rotations and escalation paths that provide continuous coverage with clean follow-the-sun handoffs.
  • Disaster recovery — own disaster-recovery testing and failover validation against defined recovery objectives and turn exercise findings into architectural and operational improvements.
  • Cloud security posture — co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring with DevSecOps.
  • Data, streaming, and AI Ops — own reliability of stateful and streaming services, capacity planning and rightsizing, and autonomous agents for triage, predictive alerting, remediation, and self-healing.
  • Scale the team and broaden impact — mentor engineers at different seniority levels, set standards adopted across teams, and align Architecture, DevSecOps, Cloud Operations, and Product Development behind a shared reliability roadmap.
Requirements
  • 7+ years of production engineering experience with recent hands-on reliability work.
  • Hands-on, recent experience operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Observability ownership — has instrumented production systems and governed service-level objectives and error budgets, not only installed dashboards.
  • Resilience practice — has designed and executed failure experiments or disaster-recovery exercises with real failover validation.
  • Incident command — has personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix.
  • Demonstrated ownership of reliability outcomes with measurable results, such as availability, mean time to recovery, and error-budget adherence.
  • Evidence of multi-team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team.
Preferred Qualifications
  • Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments.
  • Strong experience prioritizing risk using identity, workload, and exposure-path context to focus remediation on issues that materially increase attack likelihood and operational impact.
  • Experience with autonomous remediation and self-healing workflows powered by AIOps, including Amazon Bedrock Agent Core or equivalent agentic automation frameworks.
  • Hands-on experience in capacity management, resource rightsizing, efficiency engineering, and practical cost optimization based on FinOps principles.
  • Experience with load-balancing design and operations, including health-based failover, global traffic management, and performance optimization for highly available services.
  • Experience with AI traffic management via an LLM gateway, including request routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and observability for multi-model or multi-provider environments.
  • Ability to connect networking, security, and reliability considerations into cohesive platform design decisions that improve resilience, performance, and operability.

The total compensation package includes base, bonus, equity, and a range of benefit options found on our career site.

Wage Transparency

$152,000 — $228,000 USD

Compensation will vary based on individual factors such as education, qualifications, and experience of the final candidate(s), specific office location, and calibration against relevant market data and internal team equity. Posted base salary figures are subject to change as new market data becomes available. Our benefits include comprehensive medical, dental, and vision plans, matching 401(k), unlimited PTO and paid holidays, parental/adoption leave, legal insurance, and a home technology stipend. Details of participation in these benefit plans will be provided when a candidate receives an offer of employment.

At IonQ, we believe in fair treatment, access, opportunity, and advancement for all while striving to identify and eliminate barriers. We empower employees to thrive by fostering a culture of autonomy, productivity, and respect. We are dedicated to creating an environment where individuals can feel welcomed, respected, supported, and valued.

We are committed to equity and justice. We welcome different voices and viewpoints and do not discriminate on the basis of race, religion, ancestry, physical and/or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, transgender status, age, sexual orientation, military or veteran status, or any other basis protected by law. We are proud to be an Equal Employment Opportunity employer.

US Technical Jobs. The position you are applying for will require access to technology that is subject to U.S. export control and government contract restrictions. Employment with IonQ is contingent on either verifying “U.S. Person” (e.g., U.S. citizen, U.S. national, U.S. permanent resident, or lawfully admitted into the U.S. as a refugee or granted asylum) status for export controls and government contracts work, obtaining any necessary license, and/or confirming the availability of a license exception under U.S. export controls. Please note that in the absence of confirming you are a U.S. Person for export control and government contracts work purposes, IonQ may choose not to apply for a license or decline to use a license exception (if available) for you to access export-controlled technology that may require authorization, and similarly, you may not qualify for government contracts work that requires U.S. Persons, and IonQ may decline to proceed with your application on those bases alone. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law.

US Non-Technical Jobs. Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law.

#J-18808-Ljbffr
Vacancy posted 10 hours ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Santa Clara, CA vacancy
  • $170k - $200k

     ...We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance... 
    Suggested
    Full time

    Fortinet

    Sunnyvale, CA
    1 day ago
  • $150.4k - $277.6k

     ...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long...  ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced... 
    Suggested
    Relocation
    Day shift

    Apple

    Cupertino, CA
    1 day ago
  • $145k - $165k

     ...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining... 
    Suggested
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    3 days ago
  • $132.6k - $214.5k

     ..., you will collaborate closely with our engineering teams to develop innovative solutions that...  ...performance and health. As a Senior Staff SRE with the Cortex Observability team,...  ...operability of the product and ensure the reliability and availability of our services.... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $230k - $250k

     ...minds are shaping the future of network reliability, security, and AI‑ready operations. About...  ...you will be building the reliability engineering function at Forward — defining how we...  ...Looking For ~6+ years of experience in site reliability engineering, DevOps, or... 
    Suggested
    Night shift

    Forward

    Santa Clara, CA
    9 hours ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks Inc

    Santa Clara, CA
    1 day ago
  • $149.8k - $224.6k

     ...This hybrid role combines the hands-on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE). The ideal candidate has a strong technical foundation, thrives... 
    Local area

    F5

    San Jose, CA
    4 days ago
  •  ...Site Reliability Engineer Location – San Jose, CA What You'll Do - Responsibilities Engage in and improve the whole lifecycle of services—from inception and design, through automated deployment, operation and refinement. Work with all relative teams to make... 

    Netpace

    San Jose, CA
    1 day ago
  •  ...AWS Infra SRE/DevOps Engineer AWS Infra SRE/DevOps engineer with proven work experience ensuring reliability, availability and performance of cloud infra and platform. Specialist on Cisco Cloud run-on for infrastructure management, who can install, run, and maintain... 
    Work experience placement

    The Dignify Solutions, LLC

    San Jose, CA
    16 hours ago
  • $104.9k - $174.7k

     ...Site Reliability Engineer The Site Reliability Engineer role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and... 
    Temporary work
    Local area

    RELX

    San Jose, CA
    1 day ago
  •  ...Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots. FII provides customers with intelligent... 
    Permanent employment
    Full time
    Work at office
    Local area

    Foxconn Industrial Internet

    San Jose, CA
    3 days ago
  • $81.5k - $141.3k

     ...Site Reliability Engineer II Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices... 
    Remote work

    Abbott

    Sunnyvale, CA
    1 day ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response... 
    Contract work

    VDart

    Santa Clara, CA
    16 hours ago
  • Job Description : Need to have experience with ticket support, azure, Splunk, ServiceNow, and any Java experience is a plus. Ideally candidates that come from an Enterprise background Handling tickets for the Walmart environment. Splunk, Servicenow...

    3B Staffing LLC

    Sunnyvale, CA
    1 day ago
  •  ...Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job Description Deploy, support and monitor new and existing services, platforms, and application stacks. Use scale testing to measure, tune... 

    Zortech Solutions

    Cupertino, CA
    2 days ago
  •  ...Site Reliability Engineer (SRE) Location: Sunnyvale, CA (3x/ week onsite) Contract Responsibilities: Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure solutions. Operate, monitor, and... 
    Contract work

    AceStack LLC

    Sunnyvale, CA
    4 days ago
  • $110k - $130k

     ...interested in working with the World's leading AI-first Quality Engineering Company? Ready to advance your career, team up with global...  ...every day? Join us at QualityAI! We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United States... 
    Casual work
    Local area
    Flexible hours

    QualiTest Group

    Santa Clara, CA
    2 days ago
  •  ...Qualifications: 8+ years of software engineering experience, or equivalent...  ...and maintain scalable and reliable infrastructure on Google...  ...the client, IT management and staff, and other groups in Information...  .... Willingness to work on-site at stated location in the job... 
    For contractors
    Work experience placement

    Cedent Life Talent

    San Jose, CA
    1 day ago
  •  ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of...  ...are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    3 days ago
  •  ...Job Title : Site Reliability Engineer Location: San Jose, CA Duration: Contract Job Description: Extensive experience working with linux flavors like rhel/centos os, shells, filesystems and utilities Knowledge of distributed computing... 
    Contract work
    Immediate start

    Syntricate Technologies

    San Jose, CA
    16 hours ago
  • $65 - $85 per hour

     ...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    4 days ago
  • $60 - $62 per hour

     ...and improving existing processes to enhance overall system reliability. Key Responsibilities: Deploying software to cloud...  ...Computer Science or a related field. 3+ years of experience in Site Reliability Engineering. Proficiency with Kubernetes, Helm, Linux, AWS networking... 
    Hourly pay
    Contract work
    Remote work

    Akraya

    Santa Clara, CA
    3 days ago
  •  ...that keep the world running. Location: 5 On-Site Days a Week in Sunnyvale, CA Headquarters Our Engineering team is driven by a culture that thrives on visionary...  ...to-day basis, you will work on enhancing system reliability and scalability of Illumio SaaS products, and... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    3 days ago
  •  ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10+ Years Job Description: • WS application and CI/CD pipelines, Microsoft Server admin and workload support (Data Center and AWS... 
    Contract work
    Immediate start

    Syntricate Technologies

    Santa Clara, CA
    1 day ago
  •  ...Role This hybrid role combines the hands‑on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE). The ideal candidate has a strong technical foundation, thrives in... 
    Work at office
    Local area
    Remote work
    Work from home

    F5 Networks

    San Jose, CA
    1 day ago
  •  ...Site Reliability Engineer (SRE) Share Contractual Sunnyvale, CA PDT - 8450 8-10 Overview: *Must have Apple experience* • At least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Advanced experience with programming languages (Python... 

    Purple Drive

    Sunnyvale, CA
    1 day ago
  • $122.5k - $175k

     ...an impact at the company pioneering security transformation in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    1 day ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with... 
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...precision that drives great outcomes. Job Summary Key Responsibilities Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction, coaching, and career development. Own the reliability, availability, and... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    10 hours ago
  • $192.4k - $275.8k

     ...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Cisco

    San Jose, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!