Senior Staff Software Engineer - SRE & AIOps
$190.9k - $334.1kServiceNow
Company Description
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
Job Description
About the Role
ServiceNow is seeking a Senior Staff Reliability Engineer – SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, this technical leader will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.
This role combines deep hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with strategic influence across infrastructure teams. You will architect SRE tooling, develop auto-remediation capabilities, and establish patterns that allow ServiceNow's cloud platform to maintain industry-leading reliability while minimizing operational toil across follow-the-sun global teams.
What you get to do in this role:
- Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments, establishing governance, scaling policies, and operational practices that support high-velocity application deployments at 99.99%+ availability targets.
- Architect and implement closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures, trigger preventive actions, and continuously reduce MTTR and on-call burden.
- Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations, that support global follow-the-sun on-call operations and enable data-driven incident response.
- Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, while developing automated runbooks and playbooks that empower on-call engineers to resolve issues autonomously.
- Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible, auditable infrastructure deployments across hybrid and multi-cloud environments with consistent security and compliance guardrails.
- Architect hybrid cloud and data center operations, spanning on-premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices across multi-region deployments.
- Drive adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls that enable rapid, safe release cycles.
- Design on-call rotation schedules, escalation policies, and incident command systems that span across different time zones, ensuring 24/7 incident response while driving post-incident review processes that capture learning and drive systemic improvements.
- Mentor and guide junior SRE engineers, infrastructure teams, and DevOps practitioners on advanced reliability patterns, incident investigation techniques, automation best practices, and agentic AI applications for infrastructure operations.
- Champion a culture of blameless incident analysis, data-driven decision-making, continuous improvement, and experimentation across engineering teams, establishing knowledge-sharing practices and technical documentation standards.
- Reduce operational toil through systematic automation of repetitive tasks, from infrastructure provisioning to incident response to cost optimization, directly improving team capacity and job satisfaction across globally distributed operations.
Qualifications
To be successful in this role you have:
- Kubernetes Mastery : Deep, hands-on expertise operating production Kubernetes clusters at scale, including cluster design, node management, pod orchestration, resource quotas, network policies, security controls, and troubleshooting complex runtime issues.
- Incident Auto-Remediation Expertise : Proven experience designing and implementing closed-loop automated remediation systems, including anomaly detection, alert correlation, runbook automation, and self-healing mechanisms, that measurably reduce MTTR and on-call burden.
- Cloud Platform Experience : Extensive hands-on experience with AWS (EKS, EC2, RDS, Lambda), Azure (AKS, VMs, CosmosDB), and GCP (GKE, Compute Engine, Cloud SQL), capable of architecting multi-region, multi-cloud solutions.
- DevOps & IaC Proficiency : Expert-level experience with Infrastructure-as-Code tools and GitOps platforms to drive reproducible, auditable infrastructure deployments.
- SRE Tooling Fluency : Strong working knowledge of observability platforms, incident management systems, and log aggregation.
- Distributed Systems Thinking : Deep understanding of distributed system challenges, eventual consistency, cascading failures, network partitions, Byzantine fault tolerance, and proven ability to design systems resilient to these conditions.
- On-Call Operations : Experience operating in follow-the-sun, 24/7 on-call models; ability to design escalation policies, runbooks, and communication patterns that balance responsiveness with operator well-being.
- Data Center & Hybrid Cloud Operations : Hands-on experience managing both on-premises infrastructure and public cloud environments, including hybrid networking, disaster recovery, and workload migration strategies.
- AI/ML Integration : Demonstrated ability to apply machine learning and AI-driven insights to infrastructure operations, including anomaly detection, predictive alerting, and intelligent remediation.
- Influence Without Authority : Proven ability to drive technical decisions across teams, influence architecture choices, and mentor engineers outside direct reporting structure through credibility and technical depth.
Qualifications
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
- 12+ years in software engineering or infrastructure operations, with 7+ years in senior SRE, DevOps, or cloud platform engineering roles managing large-scale distributed systems with a Bachelor's degree; or 8 years and a Master's degree; or a PhD with 5 years experience; or equivalent experience.
- 5+ years hands-on experience designing, deploying, and operating production Kubernetes clusters at enterprise scale.
- Proficiency in Infrastructure-as-Code : Terraform, CloudFormation, or equivalent tools used to manage infrastructure at scale.
- Public Cloud Expertise : Demonstrable experience across 2+ of the following: AWS, Azure, GCP, with deep knowledge of services relevant to SRE operations (compute, networking, storage, observability).
- On-Call Leadership : Experience designing or operating 24/7 follow-the-sun on-call models for globally distributed teams, including escalation policies, runbook development, and incident response.
- Incident Auto-Remediation : Proven ability to architect and implement automated remediation systems, from basic alert automation to sophisticated closed-loop AI-driven systems, that meaningfully reduce manual toil.
- Linux & Systems Programming : Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash).
- SRE Mindset : Demonstrated commitment to reliability through engineering, favoring durable automation over heroics, data-driven decision-making, and continuous learning.
- Bachelor’s degree in computer science , Computer Engineering, or related field (or equivalent professional experience).
Preferred:
- Kubernetes certification (CKA, CKAD, or equivalent).
- AI/ML certification or demonstrated expertise in agentic AI frameworks and machine learning applications for infrastructure operations.
- Experience with service mesh platforms or advanced networking in Kubernetes environments.
- Background in migrating workloads from on-premises data centers to public cloud environments.
- Experience with cost optimization practices in hybrid cloud environments (reserved instances, spot instances, resource right-sizing).
- Track record of mentoring or leading infrastructure engineering teams.
Why This Role?
This role offers the opportunity to eliminate operational toil at scale and build the automation-first infrastructure practices that define modern cloud operations. You will architect systems that allow ServiceNow's global teams to sleep soundly on-call, confident that auto-remediation systems are actively preventing and resolving failures. Your work will directly shape how the organization scales reliability as it grows, establishing patterns that persist long after you've architected them. This is a role for an engineer who wants technical depth, strategic influence, and the satisfaction of watching a system recover from failure without waking anyone up.
For positions in this location, we offer a base pay of $190,900 - $334,100 , plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.
Additional Information
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.
Employee Type: RegularRegion: AMS - North America and CanadaWork Persona: Flexible$148k - $235.75k
...world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume... ...depend on. You’ll partner Software Engineering and Systems Engineering... ...production distributed systems as SRE/DevOps/Platform Ops.Proven ownership...SeniorFull time$262k - $365k
...Influence and coach a distributed team of engineers.Lead a team of engineers focussed on... ...team, ML teams, hardware platform teams, SRE teams, Google Cloud).Minimum qualifications... ...practical experience.8 years of experience in software development.7 years of experience leading...Senior- ...Palo Alto Networks in Santa Clara seeks a visionary Senior Principal Engineer/Architect to serve as the technical authority for global SRE and Platform Engineering initiatives. You will be the primary architect and driver of our AI-driven Autonomous SRE transformation...Senior
$200k - $322k
We are looking for a highly skilled Senior Software Engineer to design and develop AIOps & Observability platforms at NVIDIA. The platforms are used by internal teams to monitor, diagnose, and optimize the products, millions of assets and services in cloud, on-prem, data...SeniorFull time- ...Sr Staff Software Engineer (Fullstack - Shadow IT) 1 day ago Be among the first 25 applicants Get AI-... ...looking for a seasoned and accomplished Senior Software Engineer to help scale out our... ...-functionally with Product Management, SRE, Software, and Quality Engineering teams...SeniorFull timeWork at office
$262k - $365k
...architecture and technical road map for the software stack on our accelerator platforms.Drive... ...with hardware, software, and SRE teams to deliver scalable solutions for... ...qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical...SeniorWorldwide$237k - $288k
...reliability, security, and operability — and we're hiring a founding engineer to lead our BMC firmware work. You'll set the technical... ...power, thermal, and RAS data flowing into Crusoe Cloud's ops and SRE stack — so the BMC layer is a first-class source of platform truth...SeniorFull timeTemporary work$101k - $161k
...artificial intelligence, and software-defined networking to provide... ...prestigious awards, such as Best Engineering Team, Best Company for... ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine... ...EngineeringExperience level: Mid-Senior LevelIndustry: Computer NetworkingSenior$349k
...of the future. We are seeking a Staff Engineer to help our development of our Managed... ...services for inference and AIOps. You'll work at the intersection... ...Qualifications10+ years of experience in software engineering, platform engineering, or SRE, with at least 5 years focused on...Work at officeLocal areaImmediate startWork from homeFlexible hours- ...design, development, and optimization of high‑performance C++ software for real‑time video processing and GPU‑intensive workloads. Build... ..., OS frameworks, and third‑party applications (including game engines). Act as a technical ambassador, collaborating with GPU vendors...SeniorFull time
- NVIDIA is seeking a highly skilled Senior Software Engineer to design and develop AIOps and Observability platforms. You will mentor engineers and collaborate with product managers to define strategy and roadmaps for monitoring millions of assets across cloud, on-prem...Senior
- ..., every role at AMD contributes to something bigger — technology that moves the world forward.THE ROLEAMD is seeking a Senior Staff Software Engineer to lead compiler development and GPU performance optimization for next-generation AI and compute platforms. You will design...Senior
- ...the ground up, and we're doing it with speed. We're building software that orchestrates every go-to-market motion, enabling B2B teams... ...alongside the best! About the Role We are looking for Software Engineers who are excited to solve challenging problems and build...SeniorFull timeImmediate start
- ...Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE: AMD is looking for an influential software engineer who is passionate about improving the performance of key applications and benchmarks. You will be a member of a core team of...Senior
$262k - $365k
Design, develop, test, deploy, maintain, and enhance large scale software solutions.Provide technical leadership on high-impact projects... ..., and timelines. Influence and coach a distributed team of engineers.Drive technical project strategy, lead large-scale ML infrastructure...Senior$150k - $225k
...Research and Development - Runtime Platform /Full-time /HybridPlusAI is a Physical AI company pioneering AI-based virtual driver software for factory-built autonomous trucks. Headquartered in Silicon Valley with operations in the United States and Europe, Plus was named...SeniorFull time$262k - $365k
...approach to coding and system design while setting the standard for engineering excellence.Act as a technical multiplier. Navigate complex... ...or equivalent practical experience.8 years of experience in software development.5 years of experience testing, and launching software...Senior$262k - $364k
...Infrastructure layer of the GDC product line.Drive and influence engineering excellence across the organization, establish Service Level... ...degree or equivalent practical experience.8 years of experience in software development.5 years of experience with design and architecture...Senior$262k - $365k
..., Multi-Dimensional Roadmaps collaborating with Distinguished engineers and Google Fellows working on Exabyte scale Storage Solutions... ...experience with design and architecture; and testing/launching software products.Experience with distributed computing, big data analytics...SeniorWorldwide$223.61k - $322k
...world, through our leading products OKX, OKX Wallet, OKLink and more.About the OpportunityOur engineering team is comprised of creative problem solvers and strategic software engineers who collaborate closely with product, design, and business teams. We sit at the intersection...SeniorWorldwide$262k - $365k
...degree or equivalent practical experience.8 years of experience in software development.7 years of experience leading technical project... ...mobile GPU.Preferred qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical field.Experience...Senior$126k - $204.5k
...enterprise environments is critical. We're looking for a Senior Staff Software Engineer to build the automation frameworks, performance tooling, and... ..., Platform Engineering, Site Reliability Engineering (SRE), or a related technical discipline. Strong software engineering...SeniorFull timeWork at officeVisa sponsorshipWork visa3 days per week$198k - $326k
...robust, scalable frameworks that empower engineers and researchers to rigorously quantify... ..., and performant at scale.As a Sr. Staff Software Engineer, you will help define and build... ...posters: Full-timeFunction: SalesExperience level: Mid-Senior LevelIndustry: InternetSeniorFor contractorsWork at officeFlexible hoursShift work- ...Senior Staff Software Engineer, Developer and Qualification ToolsAt d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of...Senior
- ...versatility and a passion for building high-quality enterprise-class software systems. Primary Duties and Responsibilities Lead in... ...evaluation approaches Work with a team of architects and engineers to develop proof-of-concept systems and components Write...Senior
- ...Design and implement scalable software solutions for a B2B revenue stack, turning conceptual ideas into production-ready products. Collaborate... ...teams to drive the 0?1 creation of new products and refine engineering standards. Requirements Requires at least 5 years of software...Senior
- ...Senior Staff Engineer, SoftwareThe Senior Staff Engineer, Software develops, debugs, tests, deploys and supports code to be deployed in systems/products/equipment for various applications. They write, debug, maintain, and test software in various common languages and...SeniorWork at office
- ...Senior Staff Software EngineerAt FloQast, we aren't just building tools; we are crafting the "Operating System for Accounting." As a Senior Staff Software Engineer on the Close team, you are a primary architect of this vision. This is a role for a technical visionary who...SeniorWork at office
$190.9k - $334.1k
...Senior Staff Software EngineerWe are looking for a software engineer and technical lead with deep software design and architecture skills, excellent coding ability, and the technical judgment to build reliable production systems at enterprise scale.ServiceNow's core platform...SeniorFlexible hours$140k - $210k
...Senior Staff Software Engineer Credo is seeking a Senior Staff Software Engineer to lead the development of scalable software systems for automated validation of 800G and 1.6T OSFP and QSFP-DD optical modules. You'll architect lab automation frameworks, integrate advanced...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Staff Software Engineer - SRE & AIOps. Be the first to apply!
- srs distribution Santa Clara, CA
- senior associate architect Santa Clara, CA
- senior dynamics crm developer Santa Clara, CA
- senior application security Santa Clara, CA
- senior account director Santa Clara, CA
- senior supervisor Santa Clara, CA
- senior plumbing designer Santa Clara, CA
- senior advisor Santa Clara, CA
- senior cloud data engineer Santa Clara, CA
- senior customer service manager Santa Clara, CA



