Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer - Cloud

$168k - $264.5k

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.Are you ready to be part of something outstanding? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA, you will have a meaningful role in keeping our Digital Marketing Services reliable, fast, and efficient. You'll use innovative technology and work alongside skilled professionals who continuously explore new possibilities.What you'll be doing:Build and deploy large-scale dynamic URL redirects using Akamai Edge Redirector Cloudlets for promotional efforts and site migrations.Configure Akamai Forward Rewrite Cloudlets to map inbound requests to SEO-friendly paths.Provide on-call support for production-grade applications, responding to incidents, prioritizing issues, and driving resolution across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate shared and non-shared Cloudlet Policies via the Akamai Cloudlets Policy Manager.Maintain custom match criteria — including Geo, Device Characteristics, RegEx, and Query Strings — to ensure efficient origin offload and intelligent content delivery.Quickly identify and address user-reported problems throughout the Digital Marketing Organization ecosystem.On-board new applications, AI/ML services, and model endpoints on AWS Infrastructure.Implement monitors, alerts, and SOPs to ensure early detection and accurate response to service-impacting issues, including tracking model drift and inference latency.What we need to see:MS or BS in Computer Science/Engineering or a related field, or equivalent experience.8+ years’ experience supporting technical operations in a live-site production environment with a real passion for CDN automation, tooling, and infrastructure supporting AI applications.Strong knowledge of the Kubernetes Platform, deployments, and cloud-native automation.Proven strengths in problem-solving and root-causing issues, while continuously seeking ways to drive optimization, efficiency, and the bottom line.Advanced level experience with scripting and development in Python, fully automating operational steps with “one-click” rapid solutions.Key participation in the incident management process for early recognition of all service-impacting issues, accurate triage, partner communication, impact containment, service restoration, and post-incident follow-up. SRE On call experience is a must.Ways to stand out from the crowd:Strong Akamai CDN Support skills and deep understanding of edge computing/edge AI.Solid experience with the AWS Cloud Platform and Kubernetes as a platform. SRE on-call experience.Hands-on experience deploying and scaling Generative AI/LLM applications, integrating vector databases, or managing GPU-accelerated infrastructure.Excellent communication, presentation, and analytical skills; the ability to communicate sophisticated infrastructure and AI concepts clearly across different audiences and varying levels of the organization.Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 264,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 7, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer - Cloud in Santa Clara, CA vacancy
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our...  ...orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually...  ...optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    1 day ago
  • $174k - $253k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Sunnyvale, CA
    1 day ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...networking teams to improve service reliability and deployment...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 hours ago
  •  ...world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is designed for a... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    3 days ago
  • $101k - $161k

     ...leader in data-driven, client-to-cloud networking for large data...  ...awards, such as Best Engineering Team, Best Company for Diversity...  ...Work WithWe’re looking for Site Reliability Engineers to join our growing...  ...EngineeringExperience level: Mid-Senior LevelIndustry: Computer... 
    Senior

    Arista Networks

    Santa Clara, CA
    1 day ago
  • $90k - $180k

     ...serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale...  ...business demands, including distributed systems and cloud deployments in Azure. Work closely with software engineering... 
    Senior
    Remote work

    Abbott

    Sunnyvale, CA
    5 days ago
  • $152k - $241.5k

     ...intelligence.We’re looking for a Senior SRE to join our Compute Farm...  ...globally distributed, multi‑cloud hybrid environment - On‑prem,...  ...lifecycle management, fleet reliability/auto-healing, E2E...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $262k - $365k

     ...automation, and evolve systems by pushing for changes that improve reliability and velocity.Practice sustainable incident response and...  ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems... 
    Senior

    Google

    San Jose, CA
    1 day ago
  • $210.6k - $305.1k

     ...own. Powered by AI and an unmatched set of cloud, internet and enterprise network...  ...:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    5 hours ago
  •  ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 
    Senior

    TechDigital Group

    Santa Clara, CA
    3 days ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    3 days ago
  • $262k - $365k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and...  ...Software Engineer chose to join SRE.As the Senior Engineering Manager for Collaboration... 
    Senior

    Google

    Sunnyvale, CA
    5 hours ago
  • $187.04k - $359.72k

     ...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum...  ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas....  ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company... 
    Senior
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    3 days ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior...  ...as Python, Go, or JavaDeep expertise in cloud platforms, particularly AWS, and container... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    Palo Alto, CA
    3 days ago
  • $207.4k - $259.2k

     ...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you...  ...implement and maintain highly available, scalable, and secure cloud-native infrastructure on Amazon Elastic Kubernetes Service... 
    Senior
    Permanent employment
    Local area

    Archer Aviation

    San Jose, CA
    2 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally...  ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at...  ...systems, microservices architecture, and cloud-native technologiesHands-on... 
    Senior

    JP Morgan Chase

    Palo Alto, CA
    3 days ago
  • $184k - $287.5k

    Our Autonomous Vehicles Platform team is searching for engineers to develop and bring NVIDIA's automotive platform out to the world. You will participate in a focused effort to develop and productize ground-breaking solutions that will revolutionize the world of transportation... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous vehicle platform to new markets! This role involves developing and productizing innovative solutions that will transform transportation and the field of... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...Vehicles Platform team is looking for a hands-on System Software Engineer. As part of our team, you will work on our Autonomous Driving...  ...teams across the stack, from platform and embedded software to cloud infrastructure, underpinned by safety and performance. It extends... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...hardware design and review HW architecture & schematics.What we need to see:A Bachelor of Science Degree (or higher) in Electrical Engineering or Computer Science or equivalent experience.8+ years of experience.Domain expertise in BMC Firmware development on X86 or ARM... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...their best work. Come join the team and see how you can make a lasting impact on the world.We are now looking for a Senior Software Performance Engineer for Autonomous Vehicles! Our team builds NVIDIA’s end-to-end autonomous driving applications. We are seeking senior... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...creative Factory System Software and Diagnostics Integration engineer to join the Datacenter Platform Software team. You will play a...  ...identify and address issues, and propose solutions to enhance system reliability and efficiency. Collaborate with vendors and external partners... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are looking for a senior systems software engineer to improve the operation and user experience of distributed system infrastructure using AI. We...  ...requests and jobs on thousands of servers efficiently, reliably, and securely.What you will be doing:Design, build, test,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $152k - $241.5k

     ...to assess the performance of future GPU hardware features.What we need to see:Masters or PhD degree in Computer Science, Computer Engineering, or related field (or equivalent experience).3+ years of relevant industry experience.Strong proficiency in C++ programming and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

    GeForce NOW is Nvidia’s Cloud Gaming service, streaming games at the highest quality...  ...details, see We are looking for a Senior System Software Engineer for Cloud who sees the big picture...  ..., and enhance overall platform reliability.Influence the technology stack, architecture... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets! This role...  ..., debuggable, and production-ready.Operational Reliability: Implement atomic update semantics and safe failure... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    5 days ago
  • $176k - $276k

     ...looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team of innovative engineers who develop and maintain software facilitating GPU communication, driving groundbreaking solutions in High Performance Computing and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...for hardware-specific inference issues related to partner concernsWhat we need to see:BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.12+ years of software engineering with depth in GPU computing, ML systems, or... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...scaling for HPC and generative AI workload. Scale out is inherent to the design of this massive superchip. We are looking for expert engineers to come and help design rack level solutions for next generation scaling AI supercomputing platforms.Join us at the forefront of... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer - Cloud. Be the first to apply!