Senior Software Engineer - SRE & AIOps
$143.2k - $243.4kServiceNow
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started. Join us to put AI to work for people. Job Description About the role: ServiceNow is seeking a Senior Software Engineer - SRE & AIOps to contribute to infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale. This role combines solid hands‑on technical expertise in Kubernetes, cloud platforms, and DevOps practices with growing technical leadership capabilities. You will contribute to SRE tooling design, develop auto-remediation capabilities, and help establish patterns that maintain ServiceNow's cloud platform reliability while minimizing operational toil across follow‑the‑sun global teams. What you get to do in this role: Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments, maintaining operational standards and supporting high-velocity application deployments. Implement and maintain closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures, leveraging automation frameworks and machine learning insights to reduce MTTR and on-call burden. Contribute to the design and evolution of SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations that support global on-call operations. Develop and maintain SLO frameworks, alerting policies, and automated runbooks that empower on-call engineers to resolve issues autonomously while managing alert fatigue. Build and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible infrastructure deployments across hybrid and multi-cloud environments with security and compliance guardrails. Support hybrid cloud and data center operations, including on‑premises infrastructure, public cloud environments, and workload optimization across multi-region deployments. Contribute to adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices and network security controls. Support on‑call rotation operations and incident response processes across different time zones, helping develop runbooks and contributing to post‑incident reviews that drive continuous improvement. Share knowledge and mentor junior SRE engineers on reliability patterns, incident investigation techniques, and automation best practices. Champion a culture of blameless incident analysis, data-driven decision-making, and continuous improvement through knowledge sharing and documentation. Identify and systematically automate repetitive operational tasks, from infrastructure provisioning to incident response, improving team efficiency and capacity. Qualifications To be successful in this role you have: Kubernetes Proficiency: Solid hands‑on experience operating production Kubernetes clusters, including deployment models, pod orchestration, resource management, network policies, and troubleshooting runtime issues. Incident Remediation Experience: Demonstrated experience designing and implementing automated remediation systems, including alert automation, runbook development, and self‑healing mechanisms. Cloud Platform Knowledge: Strong hands‑on experience with AWS (EKS, EC2, RDS) and/or Azure (AKS, VMs) or GCP (GKE), with understanding of core SRE‑related services. DevOps & IaC Skills: Solid experience with Infrastructure‑as‑Code tools (Terraform, CloudFormation) and GitOps practices. SRE Tooling Familiarity: Working knowledge of observability platforms, incident management systems, and log aggregation tools. Distributed Systems Understanding: Understanding of distributed system challenges, fault tolerance, and resilience patterns. On‑Call Operations: Experience participating in on‑call rotations and understanding 24/7 operational models, runbook development, and escalation procedures. Cloud & Hybrid Operations: Hands‑on experience working with cloud infrastructure and understanding hybrid cloud concepts. Systems Administration: Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, or Bash). Collaborative Mindset: Ability to work effectively with infrastructure and application teams, contribute to technical discussions, and help drive reliability improvements. Qualifications Experience in leveraging or critically thinking about how to integrate AI into work processes, decision‑making, or problem‑solving. This may include using AI‑powered tools, automating workflows, analyzing AI‑driven insights, or exploring AI's potential impact on the function or industry. 5+ years in software engineering or infrastructure operations, with 3+ years in SRE, DevOps, or cloud platform engineering roles with a Bachelor's degree; or 3 years and a Master's degree; or a PhD without experience; or equivalent work experience. 2+ years of hands‑on experience working with production Kubernetes clusters. Proficiency in at least one Infrastructure‑as‑Code tool: Terraform, CloudFormation, or equivalent. Demonstrable hands‑on experience with at least one major cloud platform: AWS, Azure, or GCP. Experience operating in on‑call environments and participating in incident response. Experience implementing or improving automated remediation and alert systems. Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash). Demonstrated commitment to reliability engineering and continuous improvement through hands‑on contributions. Bachelor's degree in computer science, Computer Engineering, or related field (or equivalent professional experience). Preferred: Kubernetes certification (CKA, CKAD, or equivalent). Experience with service mesh technologies or advanced Kubernetes networking. Background in cloud migration or infrastructure modernization projects. Experience with cost optimization in cloud environments. Track record of implementing automation solutions that significantly reduced operational toil. Why This Role? This role offers the opportunity to work with infrastructure automation and reliability engineering at scale. You will implement systems and practices that directly reduce operational burden across ServiceNow's global engineering teams. Your contributions will help establish reliable, automated infrastructure operations and provide a clear career path toward senior technical leadership. This is a role for an engineer who enjoys solving complex operational challenges, continuous learning, and working collaboratively to improve how systems operate. For positions in this location, we offer a base pay of $143,200 - $243,400 , plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location. Additional Information We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third‑party service. Equal Opportunity Employer ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. Accommodations We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance. Export Control Regulations For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. #J-18808-Ljbffr ServiceNow
- ServiceNow in Santa Clara, CA, is seeking a Senior Software Engineer - SRE & AIOps to strengthen reliability across hybrid cloud and data center operations. You will implement automation-first systems to reduce toil and accelerate incident remediation for global engineering...Senior
$200k - $322k
We are looking for a highly skilled Senior Software Engineer to design and develop AIOps & Observability platforms at NVIDIA. The platforms are used by internal teams to monitor, diagnose, and optimize the products, millions of assets and services in cloud, on-prem, data...SeniorFull time$148k - $235.75k
...world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume... ...depend on. You’ll partner Software Engineering and Systems Engineering... ...production distributed systems as SRE/DevOps/Platform Ops.Proven ownership...SeniorFull time$101k - $161k
...artificial intelligence, and software-defined networking to provide... ...prestigious awards, such as Best Engineering Team, Best Company for... ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine... ...EngineeringExperience level: Mid-Senior LevelIndustry: Computer NetworkingSenior$172.43k - $209k
...Senior Software Engineer The Developer Experience (DevX) team is a high-impact group charged with fueling the future of Crusoe's engineering... ...Proven experience in DevOps, Site Reliability Engineering (SRE), Release Engineering, or a similar productivity-focused discipline...SeniorFull timeTemporary workLocal area- ...home day is currently Tuesday.About the RoleWe are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a... ..., and higher-level platform services for inference and AIOps. You'll work at the intersection of distributed systems, GPU...SeniorWork at officeLocal areaWork from homeFlexible hours
- ...Palo Alto Networks is seeking a Senior Principal Engineer/Architect to serve as the technical authority for global SRE and Platform Engineering initiatives across the US and India. You will architect AI-driven, self-healing platform capabilities and partner with product...Senior
$224k - $356.5k
NVIDIA is hiring engineers to scale up the introduction of next generation architecture into... ...), distributed systems, familiarity with software testing and deployment, and excellent communication... ...at scale infrastructure, DevOps and/or SRE practices and/or Platform Engineering....SeniorFull time$152k - $241.5k
...passionate about building reliable systems software for cloud-scale GPU infrastructure, we... ...to our team. We are looking for a Senior Software Engineer to join our DGX Cloud / Fleet... ...Collaborate with backend, infrastructure, SRE, security, and datacenter operations teams...SeniorFull timeLocal areaRemote work$170k - $277k
...Career We are looking for a visionary Senior Principal Engineer/Architect to serve as the technical authority for our global SRE and Platform Engineering initiatives across... ...coupled with a deep empathy for internal software developers as your primary customers. The...SeniorFull timeWork at officeVisa sponsorshipWork visaFlexible hours$164k - $205k
Join to apply for the Senior Software Engineer role at Cohesity Join to apply for the Senior Software Engineer role at Cohesity Get AI-powered... ....00-$355,000.00 1 week ago Sr Principal Engineer Software (AIOps for NGFW) Senior Software Engineer, Fabric Networking - GPU...SeniorFull timeWork at officeRemote work2 days per week3 days per week$184k - $287.5k
...building world-class reliability systems? Join NVIDIA as a Senior Software Engineer - Resilience Engineering, DGX Cloud, and be a pivotal part of... ...Experience within a world-class reliability function like Google SRE or Meta production engineering.Expertise in operating GPU,...SeniorFull time$168k - $270.25k
The NVIDIA DSX organization is looking for software engineering talent to build NVIDIA’s NICo technology. This software assists in the rapid bring... ...protocols (mutual-TLS, IPsec, or similar).Knowledge of SRE principles (observability, SLOs, logging, etc.)Ways to stand...SeniorFull timeRemote work$152k - $241.5k
...concurrency AI workloads.Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments.Use modern software engineering practices, including AI-assisted and agentic development workflows,...SeniorFull timeRemote work- Company Description Mindlance is a national recruiting company which partners with many of the leading employers across the country. Feel free to check us out at Job Description Job Description: We are currently seeking highly-motivated individuals who...SeniorFull time
$150k - $190k
...We are seeking an experienced Senior Software Engineer with a strong background in EVPN VXLAN technologies for data center and service provider environments. The ideal candidate will have hands-on experience working on a large-scale network operating system for Layer...SeniorWorldwide$150k - $180k
...we can accomplish anything. Come do your best work and live your best life as part of the IPI team! We are looking for a Senior Software Engineer to join our US team. This will be Full time opportunity based out of our office in Santa Clara, California. It will be a...SeniorFull timeWork at office$208.8k - $229.7k
...home devices, get support, change their service subscription and pay their bill. Role Description As a Senior iOS Software Engineer on the Customer Experience Mobile Team, you will play a key role in building and enhancing features for our native GFiber...SeniorFull time$98.9k - $228.7k
...infrastructure to build the systems and frameworks that set the standard for how the team operates. This is a hands-on, on-call engineering role: you will own the systems you build, respond to production incidents with urgency, and deliver permanent improvements.About...SeniorPermanent employmentFull timeWork at officeRemote work$134k - $167k
...CI/CD infrastructure to support end-to-end automation of lab software Work with internal groups to create test automation functionality... ...in a related field ~ BS in Computer Science, Computer Engineering or related field ~ Proficiency in Python or C++ ~ Experience...SeniorFull timeLocal area- ...patients worldwide. We’re a team of engineers, clinicians, and innovators united by one... ...Job Description We are seeking a Senior Unity Engineer, with an emphasis on the... ...timelines to consistently ship high-quality software. Responsibilities Evaluate the real-...SeniorFull timeLocal areaWorldwideFlexible hours
$137.1k - $188.3k
...initiative for innovative Dolby Imaging/Video algorithms and software starting with fresh proof of concept to delivering high quality... ...n * Completed Bachelor\u2019s in Computer Science, Electrical Engineering or equivalent\n * Passion for multimedia technologies and creating...SeniorFull timeLocal areaWorldwideFlexible hours$166k - $244k
A leading technology company located in Sunnyvale, California, is seeking a Site Reliability Engineer responsible for building and maintaining large-scale systems. The ideal candidate should possess a degree in Computer Science and have significant experience in programming...Senior$280k - $380k
...of disciplines.About the TeamOur DevOps/SRE team runs an active-active, multi-cloud... ...focus on reliability and automation, engineering systems that perform under stress and continuously... .../SRE (Site Reliability Engineering) Senior Software Engineer to join our dynamic team. The...SeniorWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$130k - $180k
...investors , we're positioned at the forefront of the AI-powered data engineering revolution. You can read more about us in a recently published... ...: Innovation at the Forefront : Push the boundaries of software engineering by combining traditional techniques with cutting-...SeniorFull timeWorldwide$152k - $241.5k
...make a lasting impact on the world.Join NVIDIA's dynamic team and immerse yourself in the forefront of technology! As a Senior Tegra Software Engineer, you’ll collaborate with a hardworking group, moving forward innovation in graphics, computing, and AI. This role...SeniorFull time$152k - $241.5k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.We are looking for a highly motivated senior software engineer for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced...SeniorFull time$126k - $203.5k
...define the future of operations. As a Senior DevOps Engineer, you will design and operate the platforms... ...learning to transform how we approach SRE/DevOps. This isn't just about... ...powered tools (e.g., LLM-based agents, AIOps platforms) into SRE workflows for intelligent...SeniorFull timeWork at office$184k - $287.5k
...teams with the smartest people in the world. Join us at the forefront of technological advancement.Are you a motivated system software engineer with a deep understanding of device drivers who has phenomenal C/C++ skills? If so, this role might be for you. We are looking...SeniorFull time$184k - $287.5k
We are hiring software engineers to work on the CUDA driver, a core component of our platform for accelerating general purpose computation on the GPU. Our team delivers features and improvements to better realize the potential of NVIDIA hardware for a growing range of...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer - SRE & AIOps. Be the first to apply!
- software engineer internship Santa Clara, CA
- software development engineer aws Santa Clara, CA
- software developer internship no experience Santa Clara, CA
- real time software engineer Santa Clara, CA
- financial software developer Santa Clara, CA
- graduate software developer Santa Clara, CA
- software engineer travel Santa Clara, CA
- experienced software developer Santa Clara, CA
- remote entry level software developer Santa Clara, CA
- information technology software engineer Santa Clara, CA



