Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

VDart Inc

Job Title : Senior Site Reliability Engineer

Location : Santa Clara, CA

Contract

ENGAGEMENT SUMMARY

The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response, service health, and operational automation. This role is best suited to a senior hands-on engineer who can improve availability while remaining effective in detailed production troubleshooting.

WHAT THIS CANDIDATE WILL BE DOING

  • Operate and improve reliability of AI platform services, cluster dependencies, and shared infrastructure components.
  • Lead or support incident triage for service degradation involving Kubernetes, Linux hosts, storage, network,scheduling, job orchestration, or dependency failures.
  • Define and refine SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident actions.
  • Analyze recurring failure patterns and convert manual operations into automation and preventive controls.
  • Build observability across system, service, workload, and dependency layers using metrics, logs, traces, and event correlation.
  • Troubleshoot performance and availability issues affecting training jobs, inference services, internal platforms, and support tooling.
  • Partner with infrastructure and validation teams to improve production readiness and change safety.
  • Drive operational reviews, readiness criteria, and resilience testing.

WHAT WE NEE D TO SEE

  • 7+ years in SRE, production operations, or reliability-focused infrastructure engineering.
  • Strong hands-on troubleshooting across Linux, Kubernetes, networking, and distributed systems.
  • Experience building observability, alerting, and response workflows in complex production environments.
  • Ability to balance urgent operational response with medium-term reliability engineering improvements.
  • Strong scripting and automation skills, with experience reducing toil through tooling.
  • Experience participating in incident management, root cause analysis, and post-incident follow-through.
  • Strong communication skill with the ability to summarize technical issues clearly for cross-functional teams.

PREFERRED EXPERIENCE

  • Experience in AI platforms, ML infrastructure, or large-scale HPC-like service environments.
  • Familiarity with Prometheus, Grafana, ELK/OpenSearch, Loki, PagerDuty, and incident tooling.
  • Experience defining error budgets and applying SRE practices in environments with heavy batch and service traffic.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Santa Clara, CA vacancy
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    5 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $160k - $240k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our...  ...operations or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    1 day ago
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We... 
    Senior
    Remote work

    Abbott

    Sunnyvale, CA
    4 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    5 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 days ago
  • $192.4k - $275.8k

     ...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines...  ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    2 days ago
  • $200k - $322k

     ...best work.We are seeking a highly skilled Senior Staff SRE to join our dynamic team. Our...  ...includes building for performance and reliability at global scale, covering automation, monitoring...  ...with NVIDIA leadership, senior engineers, program managers, and product managers... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  •  ...Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $187.04k - $359.72k

     ...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum...  ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas....  ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company... 
    Senior
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    2 days ago
  • $166k - $244k

    Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have... 
    Senior
    Full time

    Google

    Sunnyvale, CA
    2 days ago
  • $128k - $216k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe...  ...a difference at Fiserv.Job TitleSr. Site Reliability EngineerAbout...  ...with confidence.What does a successful Senior Site Reliability Engineer do at Fiserv?As a Senior Site Reliability... 
    Senior
    Full time
    Worldwide

    Fiserv

    Sunnyvale, CA
    2 days ago
  • $222k - $300.5k

     ...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational...  .... The Fintech Platform Systems Engineering team builds and operates the AWS-based...  ...negotiable.The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to... 
    Senior
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    4 days ago
  • $262k - $364k

     ...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with...  ...capacity and performance.Build creative engineering solutions to operations and...  ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software... 
    Senior

    Google

    Mountain View, CA
    5 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms... 
    Senior

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  •  ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available database...  ...data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and... 
    Senior

    Neshent Technologies

    Los Gatos, CA
    1 day ago
  • $166k - $244k

    A leading technology company located in Sunnyvale, California, is seeking a Site Reliability Engineer responsible for building and maintaining large-scale systems. The ideal candidate should possess a degree in Computer Science and have significant experience in programming... 
    Senior

    Google

    Sunnyvale, CA
    2 days ago
  •  ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will... 
    Senior

    Kody

    Palo Alto, CA
    10 days ago
  • $120k - $200k

    Sr Site Reliability Engineer (Prisma Access) 2 days ago Be among the first 25 applicants Job Description This role requires US Citizenship. Your Career Palo Alto Networks runs a large infrastructure and is one of the biggest GCP customers. As a Principal SRE, you'll be... 
    Senior
    Rotating shift

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  • $128.6k - $184.9k

     ...private datacenters and AWS while ensuring reliable operations, resilience, and zero-...  ...private cloud environments. Lead cloud engineering initiatives using Terraform, Ansible, and...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Santa Clara, CA
    6 days ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    2 days ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Full time

    Google

    Sunnyvale, CA
    2 days ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

     ...the world.Join the Simulation Software team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding opportunity to work on...  ...development by enabling Chips Simulation as a trusted and reliable virtual platform.What you will be doing:Drive early... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!