Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

ReqRoute,Inc

Position: Site Reliability Engineer-10+ Year exp required
Location : Sunnyvale CA

Job Summary

We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems. This role requires deep technical understanding across the entire application stack, the ability to read and reason about code, and the capability to provide data-backed answers to both engineering teams and business stakeholders.

This role goes beyond traditional operations or DevOps. The successful candidate will think like a software engineer, act like a systems engineer, and operate with a production-first mindset.

Key Responsibilities

Performance & Reliability Engineering

  • Analyze and resolve performance issues such as high latency, slow login, throughput degradation, and system instability.
  • Perform deep, end-to-end investigations across the full stack including:
    • Load balancers and traffic routing
    • Web server and application runtime configurations
    • Middleware and messaging systems
    • Database performance (queries, indexing, pooling)
    • Kubernetes clusters (pods, resources, scaling behavior)
    • Linux OS tuning (CPU, memory, I/O, ulimits, networking)
  • Identify root causes and propose clear, actionable engineering solutions.

Distributed Systems Design

  • Design, review, and influence high-performance, highly-available distributed architectures.
  • Evaluate trade-offs related to scalability, latency, fault tolerance, and cost.
  • Partner with development teams early to prevent reliability and performance issues before production.

Capacity Planning & Scalability

  • Assess system readiness for growth scenarios such as:
    • "We plan to onboard 10,000 users in 6 months - can the system support it?"
  • Perform capacity and scale analysis for:
    • Application tiers
    • Databases
    • Messaging systems
    • Kubernetes compute and storage
  • Provide evidence-based recommendations supported by metrics, benchmarks, and production data.

Engineering Collaboration

  • Work closely with software engineering teams to:
    • Review performance-critical code paths
    • Propose improvements at code, configuration, or infrastructure level
    • Improve system observability (metrics, logs, traces)
  • Communicate complex technical findings clearly to both engineers and business stakeholders.

Required Technical Skills

  • Strong understanding of distributed systems and performance engineering
  • Ability to read, analyze, and troubleshoot Java code
  • Hands-on experience with:
    • Kubernetes (resource management, scaling, container behavior)
    • Linux internals and tuning
    • PostgreSQL (queries, indexing, performance optimization)
  • Proven experience building or operating high-availability, high-throughput systems
  • Strong analytical and problem-solving skills with a data-driven approach

Nice to Have

  • Experience with Azure cloud services
  • Messaging systems such as ActiveMQ
  • Load testing and benchmarking experience
  • Background in roles such as SRE, Performance Engineering, Platform Engineering

Required Skills & Qualifications

Technical Skills

  • Hands-on experience with cloud platforms (Azure.
  • Strong scripting skills (e.g., Python, Bash, PowerShell, or similar).
  • Experience with deployment pipelines, automation, and monitoring tools.
  • Solid understanding of cloud infrastructure, networking, and application operations.

LLM & AI Experience

  • Practical experience working with Large Language Models (LLMs).
  • Familiarity with applying LLMs to engineering or operational workflows is required.

Professional Attributes

  • Strong desire to learn and deeply understand complex systems.
  • Self-starter with the ability to take ownership and drive initiatives independently.
  • Demonstrates leadership, accountability, and problem-solving mindset.
  • Strong collaboration and communication skills
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Sunnyvale, CA vacancy
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Suggested
    Night shift

    Forward Networks Inc

    Santa Clara, CA
    3 days ago
  • $145k - $165k

     ...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining... 
    Suggested
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    1 day ago
  •  ...Job Description Job Description Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our cloud infrastructure (GCP or AWS, on-prem). Build, maintain, and optimize Kubernetes... 
    Suggested

    Amiri Recruiting

    Mountain View, CA
    5 days ago
  •  ...of Huobi globe spanning infrastructure. •       Work with engineering teams to make sure new features and changes are deployed quickly...  .... •       Constantly improve our system performance and reliability through better tools, process and monitoring system. •... 
    Suggested
    Worldwide

    Cryptoware Technologies Inc

    Santa Clara, CA
    a month ago
  •  ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the... 
    Suggested
    Work at office

    Foxconn Industrial Internet - FII

    Sunnyvale, CA
    a month ago
  • $150k - $195k

     ...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the...  ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation.... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  • $60 - $62 per hour

     ...and improving existing processes to enhance overall system reliability. Key Responsibilities: Deploying software to cloud...  ...Computer Science or a related field. 3+ years of experience in Site Reliability Engineering. Proficiency with Kubernetes, Helm, Linux, AWS networking... 
    Hourly pay
    Contract work
    Remote work

    Akraya

    Santa Clara, CA
    1 day ago
  • $110k - $130k

     ...interested in working with the World's leading AI-first Quality Engineering Company? Ready to advance your career, team up with global...  ...every day? Join us at QualityAI! We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United States... 
    Casual work
    Local area
    Flexible hours

    QualiTest Group

    Santa Clara, CA
    5 days ago
  • $170k - $200k

     ...Site Reliability Engineer We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high... 
    Full time
    Worldwide

    Edelman

    Sunnyvale, CA
    3 days ago
  • $200k - $260k

     ...for enterprise trust, as we bring Work AI to every employee, in every company. About the Role: Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing,... 
    Work at office
    Home office
    Flexible hours

    Glean.info

    Mountain View, CA
    1 day ago
  • $132.6k - $214.5k

     ...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and...  ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications DevOps... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  • $195k - $285k

     ...purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we build. This role builds and leads d-Matrix's Site Reliability Engineering function from the ground up, owning the... 
    Full time
    Remote work

    d-Matrix

    Santa Clara, CA
    2 days ago
  •  ...Role :- Site Reliability Engineer (SRE) Infrastructure & Agentic Automation Location :- Santa Clara, CA (Hybrid) Work Authorization: USC/GC only Position Summary & Job Description:- Client is looking for an experienced Site Reliability Engineer (SRE) to... 

    ReqRoute,Inc

    Santa Clara, CA
    2 days ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    1 day ago
  •  ...Job Description Job Description Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots... 
    Permanent employment
    Full time
    Work at office
    Local area

    Foxconn Industrial Internet - FII

    San Jose, CA
    a month ago
  •  ...Site Reliability Engineer Location – San Jose, CA What You'll Do - Responsibilities Engage in and improve the whole lifecycle of services—from inception and design, through automated deployment, operation and refinement. Work with all relative teams to make... 

    Netpace

    San Jose, CA
    3 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    1 day ago
  • $165k - $280k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we're leveraging our experience in building rockets and spacecraft to deploy Starlink, the world's... 
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  •  ...Must Have Technical/Functional Skills: 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role supporting cloud-based production environments. Practical vulnerability-management experience; familiarity with... 
    Full time
    Worldwide

    SFE

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (...  ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Job Description Job Description Site Reliability Engineer II Bay Area, offices in San Jose · Hybrid · 24/7 FedRAMP Operations · Rotational Shift · Initial Contract till March 27. KEY REQUIREMENT This role requires US citizenship and residence on US soil.... 
    Hourly pay
    Contract work
    For contractors
    Shift work
    Night shift
    Weekend work

    C-Serv

    San Jose, CA
    25 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...AI, IBM and Accern. Position Summary We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE practices... 
    Shift work

    Wand AI

    Palo Alto, CA
    1 day ago
  •  ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently... 
    Shift work

    AI Chopping Block

    Menlo Park, CA
    3 days ago
  •  ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We...  ...Kubernetes platforms. This role is focused on infrastructure, reliability, and automation , with Java exposure as a supporting skill.... 

    Eitacies Inc

    Santa Clara, CA
    a month ago
  •  ...Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE Devops Engineer Fulltime Role for one of our direct client. Role: SRE Devops Engineer Location :- Santa Clara,CA (Remote Looking Local... 
    Full time
    Local area
    Remote work

    Rootshell Enterprise Technologies

    Santa Clara, CA
    3 days ago
  •  ...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the... 

    Outshift by Cisco

    San Jose, CA
    5 days ago
  • $125k - $160k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we're leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation... 
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    55 minutes ago
  •  ...Job Title: Senior Site Reliability Engineer Kubernetes Platform Location: Remote Duration : Full Time Job Description Must Have Technical/Functional Skills: ~10+ years of experience in SRE, DevOps, or infrastructure engineering ~ Strong... 
    Full time
    Remote work

    SFE

    San Jose, CA
    2 days ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Full time

    Google

    Sunnyvale, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!