Senior Staff Service Reliability and Operational Intelligence Engineer
$187.95k - $269.5kIonQ
About IonQ:
IonQ, Inc . [NYSE: IONQ] is the world’s leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ’s newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting-edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two-qubit gate fidelity, setting a world record in quantum computing performance .
Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Canada, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space. IonQ is making quantum platforms more accessible and impactful than ever before. Location: Santa Clara, CA
Travel: Up to 25%
Job ID: 1795
The Role:
The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites.
The Service Reliability and Operational Intelligence discipline ensures the platform remains stable and resilient, with focus on service continuity and seamless customer experience. It owns production reliability and resilience, observability architecture, service-level objectives, incident response, and implementation of AIOps workflows for triage, remediation, and self-healing.As a Senior Staff Service Reliability and Operational Intelligence Engineer, you define the technical direction for reliability across regions and services. You own the reliability strategy, establish the standards and mechanisms that guide production operations, and elevate excellence through design leadership, operational discipline, and mentorship. You stay deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, and building resilience and disaster-recovery automation.
The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.
Responsibilities :
- Own the technical strategy and multi-year roadmap for operational excellence and production readiness across development, pre-production, and production environments.
- Define and govern the New Service Introduction framework, including mandatory architecture, security, resilience, capacity, observability, supportability, and release-readiness reviews before services enter production.
- Establish organization-wide service ownership standards covering service catalog records, accountable owners, dependency maps, runbooks, support models, escalation paths, recovery objectives, and on-call readiness.
- Lead the architecture and evolution of the shared observability platform, establishing consistent standards for logs, metrics, distributed traces, and profiles across production systems.
- Define standards for dashboards, alert policies, synthetic monitoring, telemetry quality, retention, sampling, cardinality, and cost controls.
- Own the reliability governance model for production services, including SLIs, SLOs, error budgets, and escalation mechanisms.
- Connect service-health signals to customer and business impact, enabling early anomaly detection, service-degradation prevention, and rapid isolation of end-user-impacting events.
- Advance incident-management maturity through consistent severity classification, incident command, stakeholder and executive communications, automated evidence collection, and coordinated response to high-severity incidents.
- Establish blameless post-incident review practices, ensure remediation actions are tracked to completion, and drive systemic fixes for recurring failure modes.
- Lead operational capacity and efficiency management, including demand forecasting, cloud and Kubernetes capacity, performance testing, scaling thresholds, headroom policies, resource rightsizing, and capacity-risk reviews.
- Design and govern AI Ops capabilities for event correlation, alert-noise reduction, predictive detection, probable root-cause analysis, autonomous triage, assisted remediation, and controlled self-healing.
- Deliver secure AI-agent workflows across observability platforms, service catalog, Jira, Confluence, source control, and CI/CD.
- Improve on-call effectiveness through sustainable rotation design, operational-readiness standards, escalation policies, diagnostic automation, alert-quality management, and reliable follow-the-sun handoffs.
- Provide hands-on technical leadership during major incidents, complex reliability investigations, architectural reviews, resilience exercises, and critical service launches.
- Use operational data, incident trends, service-level performance, capacity signals, change outcomes, and automation effectiveness to prioritize continuous improvement.
Requirements:
- 12+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
- Recent experience designing and operating large-scale, fault-tolerant production systems on AWS or GCP.
- Deep understanding of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
- Demonstrated ownership of observability architecture, including instrumentation of production systems and governance of metrics, logs, traces, SLIs, SLOs, and error budgets.
- Proven experience establishing reliability and operational-readiness standards for business-critical services.
- Hands-on experience designing and executing failure experiments, disaster-recovery exercises, and validated service failovers.
- Experience personally commanding SEV1 or SEV2 incidents, coordinating technical and executive communications, and driving root causes through to systemic remediation.
- Demonstrated ownership of measurable reliability outcomes such as availability, latency, MTTR, change-failure rate, alert quality, and error-budget adherence.
- Experience with capacity forecasting, performance testing, scaling strategies, and cloud and Kubernetes resource management.
- Strong software engineering and automation skills using languages such as Python or Go, infrastructure as code, and modern delivery toolchains.
- Evidence of multi-team technical leadership through architecture reviews, standards, coaching, and mechanisms adopted beyond a single service or team.
- Ability to influence cross-functional stakeholders and deliver complex initiatives without relying on direct management authority.
Preferred Qualifications:
- Experience prioritizing operational risk using identity, workload, dependency, and exposure-path context to focus remediation on issues with material customer or business impact.
- Experience designing AI Ops capabilities for anomaly detection, event correlation, predictive alerting, root-cause analysis, and operational noise reduction.
- Hands-on experience with autonomous remediation and self-healing workflows using Amazon Bedrock AgentCore or comparable agentic automation frameworks.
- Experience integrating governed AI agents with operational platforms such as Jira, Confluence, source control, CI/CD, service catalogs, and observability systems.
- Practical experience with capacity optimization, resource rightsizing, efficiency engineering, telemetry cost management, and FinOps principles.
- Experience designing and operating load-balancing solutions, health-based failover, global traffic management, and performance optimization for highly available services.
- Ability to integrate networking, security, resilience, performance, and operability requirements into cohesive platform architecture decisions.
- Experience with progressive-delivery techniques such as canary deployments, blue-green deployments, automated rollback, and feature-flag governance.
- Experience establishing sustainable global on-call models and follow-the-sun operational practices.
The total compensation package includes base, bonus, equity, and a range of benefit options found on our career site.
If this role has a commission structure, the compensation range below just reflects the base compensation range.
Wage Transparency:
$187,945—$269,503 USD
Compensation will vary based on individual factors such as education, qualifications, and experience of the final candidate(s), specific office location, and calibration against relevant market data and internal team equity. Posted base salary figures are subject to change as new market data becomes available. Our benefits include comprehensive medical, dental, and vision plans, matching 401(k), unlimited PTO and paid holidays, parental/adoption leave, legal insurance, and a home technology stipend. Details of participation in these benefit plans will be provided when a candidate receives an offer of employment.
At IonQ, we believe in fair treatment, access, opportunity, and advancement for all while striving to identify and eliminate barriers. We empower employees to thrive by fostering a culture of autonomy, productivity, and respect. We are dedicated to creating an environment where individuals can feel welcomed, respected, supported, and valued.
We are committed to equity and justice. We welcome different voices and viewpoints and do not discriminate on the basis of race, religion, ancestry, physical and/or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, transgender status, age, sexual orientation, military or veteran status, or any other basis protected by law. We are proud to be an Equal Employment Opportunity employer.
US Technical Jobs. The position you are applying for will require access to technology that is subject to U.S. export control and government contract restrictions. Employment with IonQ is contingent on either verifying “U.S. Person” (e.g., U.S. citizen, U.S. national, U.S. permanent resident, or lawfully admitted into the U.S. as a refugee or granted asylum) status for export controls and government contracts work, obtaining any necessary license, and/or confirming the availability of a license exception under U.S. export controls. Please note that in the absence of confirming you are a U.S. Person for export control and government contracts work purposes, IonQ may choose not to apply for a license or decline to use a license exception (if available) for you to access export-controlled technology that may require authorization, and similarly, you may not qualify for government contracts work that requires U.S. Persons, and IonQ may decline to proceed with your application on those bases alone. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law.
US Non-Technical Jobs. Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum. Accordingly, we will have some additional questions regarding your immigration status that will be used for export control and compliance purposes, and the answers will be reviewed by compliance personnel to ensure compliance with federal law.
If you are interested in being a part of our team and mission, we encourage you to apply!
$149.9k - $202.8k
...are looking for a Senior Business Intelligence Engineer to build and run the... ...you turn them into reliable data and clear answers... ...across sourcing, operations, finance, and... ..., supervisors, and staff; adhere to standards... ...exceptional customer service; and follow all federal...SeniorLocal areaFlexible hoursShift work$185.9k - $300.68k
...outcomes.Job SummaryAs the Senior Manager of Product Management for Operations & Reliability, you will build and... ...our AI Firewall services meet a 99.99% SLA. This... ...work across Product, Engineering, SRE, and Support to... ...proactive monitoring, intelligent alerting, and...SeniorFull timeWork at office- ...and multi-year roadmap for operational excellence and production readiness... ...Define and govern the New Service Introduction framework and... ..., and cost controls. Own reliability governance involving SLIs,... ...~12+ years of production engineering, site reliability engineering...SeniorFull timeContract work
$245k - $312k
...currently seeking a dynamic FortiGuard Senior Threat Intelligence Research Engineer to contribute to the success of... .... This role is designed to operate at the strategic intersection of cybersecurity... ...-impact opportunities.FortiGuard Services offer broad security solutions...SeniorLocal areaRemote work- STMicroelectronics is seeking an experienced operations leader to serve as the primary interface... ...ST for orders, forecast alignment, and service performance. You will drive order... ...and divisions to translate demand into reliable supply commitments, manage risk and escalations...Senior
- About The Team Our Business Intelligence team exists at the intersection of data and strategy... ...between systems with precision and reliability. As part of the DoorDash family, we are... ...& Analytics Team under the Strategy & Operations - In-Store organization. You’re Excited...SeniorHourly payWork at officeLocal areaRemote workFlexible hours
- ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure... ...has over 8 years of experience in wireless network operations and a strong background in wireless technologies like...Senior
- Senior Director, Artificial Intelligence Engineering Full-time Halo Group is a premier provider of IT talent. We place... ...our client and with consulting services who are working with customers... ...analytic application development, and operations. This is an Engineering...SeniorFull time
- ...cloud inference services. This order of magnitude... ...and increasing intelligence via additional... ...and data center operations programs supporting... ...with Hardware Engineering, Inference Engineering... ...systems are reliably deployed, operated... ...operational risks to senior...Senior
$300 per month
...abundance of energy and intelligence . As the only... ...up, we own and operate each layer of... ..., and cloud services. If you want... ...execution. This senior individual-contributor... ..., AI-assisted reliability engineering, and support... ...operations, chief-of-staff, or related...SeniorTemporary workShift work$256k - $385.25k
...more than 25 years. As a Manager, Product Reliability, you will be at the forefront of... ...you to lead and build the future of our engineering projects, making a lasting impact on the... ...engineering, validation, manufacturing, operations, and field teams to resolve issues and...SeniorFull time- Apple Inc. is seeking a Senior Site Reliability Engineer in Cupertino to drive reliability, scalability, and observability of our cloud... ...with developers and architects to build and operate mission-critical services at scale, ensuring uptime and performance for millions...Senior
- Palo Alto Networks Senior Manager of Product Management for Operations & Reliability will build and lead a team ensuring our AI Firewall services meet a 99.99% SLA and scale globally. You will... ...capabilities across Product, Engineering, SRE, and Support to enable seamless...Senior
- ...Sunnyvale, CA is seeking a Data Center Operations Technician L3 to support deployment and... ...handle hardware bring-up, diagnostics, and reliability improvements across enterprise-scale... .... You will collaborate with engineering and infrastructure teams, use Linux command...Senior
$232k - $368k
...volume production. We are hiring a Senior Manager to lead our Test, Manufacturability, Reliability & Quality (TMRQ) organization.... ...every program. Establish operational objectives, and refine closed-... ...retain your manager and lead engineers across multiple geographies —...SeniorFull time- ...premier telecommunications client in Sunnyvale, CA. The role focuses on operations of a 140,000-square-foot facility, coordinating repairs, vendor relations, and budget oversight to deliver reliable service. You will perform site inspections, ensure regulatory compliance,...Senior
- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn... ...striving for sustained operational excellence of Sumo’s... ...Compute, Storage, and managed services.Competency with modern CI... ...data through its Intelligent Operations Platform. Built...SeniorFlexible hours
$262k - $364k
...improve the whole lifecycle of services—from inception and design, through to deployment, operation and refinement.Support... ...pushing for changes that improve reliability and velocity.Practice sustainable... ...in Computer Science or Engineering.Site Reliability Engineering...Senior$168k - $264.5k
...: to amplify human inventiveness and intelligence. Make the choice to join us today. We... ...an outstanding candidate for Silicon Reliability Engineer to drive and utilize cutting edge technologies... ..., Advanced Technology Group and Operations to develop process reliability...SeniorFull time$168k - $264.5k
...Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa... ...our Digital Marketing Services reliable, fast, and... ...efficient origin offload and intelligent content delivery.Quickly... ...experience supporting technical operations in a live-site production...SeniorFull time$152k - $241.5k
.... Success in this role requires both operational precision along with developing and supporting... ...improvements in observability, service reliability, and automation, ensuring the EDA... ...measurable, and aligned with long-term engineering demands.What you'll be doing:Manage,...SeniorFull time- Everpure in Santa Clara is seeking a Senior Software Engineer in Production Engineering to own the... ...This is a developer-first role with operational responsibility for code merges and... ...automate quality gates, and improve reliability with observability and data-driven insights...Senior
- Programmers.io in California's Santa Clara region seeks a reliability engineer with a BS in Engineering and 5+ years in high-tech environments... ...will understand reliability testing processes, program and operate HALT/HASS, environmental chambers, and mechanical shock/vibration...Senior
- NVIDIA is seeking a skilled Systems Operations & Administrator to join the Networking SW group in Santa Clara. You will... ...lab and data-center infrastructure to keep systems reliable, scalable, and ready for engineering use. The role requires strong problem-solving, Linux...Senior
- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...-based control plane services and dataplane software running... ...tooling and automation to reduce operational toil and improve... ...networking teams to improve service reliability and deployment workflowsDeploy...SeniorWork at officeLocal areaWork from homeFlexible hours
$240k - $379.5k
...future of technology. In the role of Senior Manager of Business Operations, you will join a proactive team... ...closely with internal data center engineering teams.Act as the primary point of... ...visibility. This enables scale and reliability in rapidly evolving data center environments...SeniorFull time$184k - $287.5k
...are looking for a Senior System Software Engineer, Software Defined Networking... ...design, build, and operate highly performant... ...production through reliability engineering, CI/CD,... ...-as-a-Service virtual network orchestration... ..., telemetry, intelligent metering, and performance...SeniorFull time$116k - $184k
...build the next era of computing!We're seeking an outstanding Senior HTOL Reliability Engineer to join our Santa Clara lab. This role requires deep... ...optimize HTOL test programs aligned with JEDEC standards.Operate and maintain HTOL ovens, ensuring efficient test conditions...SeniorFull time- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building... ...and maintain scalable control plane services, operators, and custom controllers for... ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations...SeniorWork at officeLocal areaWork from homeFlexible hours
- Q CELLS America Inc. seeks a Senior DevOps & SRE Manager to lead Platform Reliability & Global Operations. You will drive reliability, scalability, security, and operational excellence across multiple platforms including workflows, event streaming, and data pipelines....Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Staff Service Reliability and Operational Intelligence Engineer. Be the first to apply!
- community services specialist Santa Clara, CA
- professional services specialist Santa Clara, CA
- toyota service advisor Santa Clara, CA
- environmental services attendant Santa Clara, CA
- service administrator Santa Clara, CA
- health services administrator Santa Clara, CA
- car service advisor Santa Clara, CA
- facility service associate Santa Clara, CA
- nutrition services aide Santa Clara, CA
- service worker Santa Clara, CA


