Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

ReqRoute,Inc

Job Title: Site Reliability Engineer



Location: Sunnyvale, CA



Experience Required: 10+ years



Employment Type: W2



Work Authorization: All visa types considered except H1B

Job Summary:
Seeking an experienced engineer who can analyze, diagnose, and optimize the performance and reliability of large-scale distributed systems. This role requires deep technical understanding across the entire application stack, the ability to read and reason about code, and the capability to provide data-backed answers to both engineering teams and business stakeholders.

This role goes beyond traditional operations or DevOps. The successful candidate will think like a software engineer, act like a systems engineer, and operate with a production-first mindset.

Key Responsibilities:

Performance & Reliability Engineering

  • Analyze and resolve performance issues such as high latency, slow login, throughput degradation, and system instability.
  • Perform deep, end-to-end investigations across the full stack, including load balancers and traffic routing, web server and application runtime configurations, middleware and messaging systems, database performance (queries, indexing, pooling), Kubernetes clusters (pods, resources, scaling behavior), and Linux OS tuning (CPU, memory, I/O, ulimits, networking).
  • Identify root causes and propose clear, actionable engineering solutions.

Distributed Systems Design

  • Design, review, and influence high-performance, highly available distributed architectures.
  • Evaluate trade-offs related to scalability, latency, fault tolerance, and cost.
  • Partner with development teams early to prevent reliability and performance issues before production.

Capacity Planning & Scalability

  • Assess system readiness for growth scenarios (e.g., onboarding 10,000 additional users within six months).
  • Perform capacity and scale analysis across application tiers, databases, messaging systems, and Kubernetes compute and storage.
  • Provide evidence-based recommendations supported by metrics, benchmarks, and production data.

Engineering Collaboration

  • Work with software engineering teams to review performance-critical code paths and propose improvements at the code, configuration, or infrastructure level.
  • Improve system observability (metrics, logs, traces).
  • Communicate complex technical findings clearly to both engineers and business stakeholders.

Required Technical Skills:

  • Strong understanding of distributed systems and performance engineering.
  • Ability to read, analyze, and troubleshoot Java code.
  • Hands-on experience with Kubernetes (resource management, scaling, container behavior), Linux internals and tuning, and PostgreSQL (queries, indexing, performance optimization).
  • Proven experience building or operating high-availability, high-throughput systems.
  • Hands-on experience with Azure cloud services.
  • Strong scripting skills (Python, Bash, PowerShell, or similar).
  • Experience with deployment pipelines, automation, and monitoring tools.
  • Solid understanding of cloud infrastructure, networking, and application operations.
  • Strong analytical and problem-solving skills with a data-driven approach.

LLM & AI Experience:

  • Practical experience working with Large Language Models (LLMs).
  • Familiarity with applying LLMs to engineering or operational workflows.

Nice to Have:

  • Messaging systems such as ActiveMQ.
  • Load testing and benchmarking experience.
  • Background in SRE, Performance Engineering, or Platform Engineering roles.

Professional Attributes:

  • Strong desire to learn and deeply understand complex systems.
  • Self-starter with the ability to take ownership and drive initiatives independently.
  • Demonstrates leadership, accountability, and a problem-solving mindset.
  • Strong collaboration and communication skills.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Sunnyvale, CA vacancy
  • $160k - $240k

     ...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit...  ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in... 
    Suggested
    Full time

    Fiserv

    Sunnyvale, CA
    1 day ago
  • $128.6k - $184.9k

     ...private datacenters and AWS while ensuring reliable operations, resilience, and zero-...  ...private cloud environments. Lead cloud engineering initiatives using Terraform, Ansible, and...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Santa Clara, CA
    6 days ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Suggested
    Night shift

    Forward Networks

    Santa Clara, CA
    2 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Suggested
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  • $128k - $216k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe your...  ...make a difference at Fiserv.Job TitleSr. Site Reliability EngineerAbout CloverClover is...  ...does a successful Senior Site Reliability Engineer do at Fiserv?As a Senior Site Reliability... 
    Full time
    Worldwide

    Fiserv

    Sunnyvale, CA
    2 days ago
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are... 
    Remote work

    Abbott

    Sunnyvale, CA
    4 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $101k - $161k

     ...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-...  ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-... 

    Arista Networks

    Santa Clara, CA
    5 days ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Full time

    Google

    Sunnyvale, CA
    2 days ago
  • $276.1k - $311.4k

     ...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a... 
    Permanent employment
    Full time
    Work at office
    Work from home

    Lindus Health

    Sunnyvale, CA
    4 days ago
  • $120k - $180k

     ...cybersecurity starts with you.About the Role:At CrowdStrike, our engineering organization depends on shared infrastructure platforms that...  ...platforms require dedicated engineering ownership to operate reliably, scale safely, harden for security, and mature into self-service... 
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    3 days ago
  • $145k - $165k

     ...Your Ego : Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    3 days ago
  •  ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 

    TechDigital Group

    Santa Clara, CA
    2 days ago
  •  ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable... 

    Epic Games

    Sunnyvale, CA
    2 days ago
  •  ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud... 
    Full time

    Saransh

    Sunnyvale, CA
    2 days ago
  • $150k - $195k

     ...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the...  ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation.... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    1 day ago
  •  ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data... 

    Tik Tok

    Mountain View, CA
    2 days ago
  • $150.4k - $277.6k

     ...Technical Operations & Site Reliability Engineer, Customer SystemsAt Apple, Customer Experience is at the forefront of everything we do. The Customer Systems Operations team is looking for a highly skilled and motivated TechOps Engineer (Technical Operations & Site Reliability... 
    Work experience placement
    Relocation

    Apple

    Sunnyvale, CA
    3 days ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 

    Bolt Graphics, Inc.

    Sunnyvale, CA
    2 days ago
  • $262k - $364k

     ...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with...  ...capacity and performance.Build creative engineering solutions to operations and...  ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software... 

    Google

    Mountain View, CA
    5 days ago
  • $222k - $300.5k

     ...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps...  ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    4 days ago
  •  ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the... 
    Work at office

    Foxconn Industrial Internet - FII

    Sunnyvale, CA
    10 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response... 
    Contract work

    VDart Inc

    Santa Clara, CA
    3 days ago
  •  ...of Huobi globe spanning infrastructure. •       Work with engineering teams to make sure new features and changes are deployed quickly...  .... •       Constantly improve our system performance and reliability through better tools, process and monitoring system. •... 
    Worldwide

    Cryptoware Technologies Inc

    Santa Clara, CA
    10 days ago
  • $120k - $200k

    Sr Site Reliability Engineer (Prisma Access) 2 days ago Be among the first 25 applicants Job Description This role requires US Citizenship. Your Career Palo Alto Networks runs a large infrastructure and is one of the biggest GCP customers. As a Principal SRE, you'll be... 
    Rotating shift

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $200k - $322k

     ...PTP, DHCP, and LDAP. This includes building for performance and reliability at global scale, covering automation, monitoring, high...  ...alerting, monitoring.Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop compelling... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $207k - $300k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...Science or Engineering.1 year of people management experience. Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    Mountain View, CA
    5 days ago
  • $255.7k - $300k

    Lead a team of engineers to maintain service uptime while managing global on-call rotations...  ...improve operational practices to drive reliability, maintainability, and stakeholder alignment...  ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation... 
    Full time
    Work at office

    Google

    Sunnyvale, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!