Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$230k - $250k

Forward Networks

Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment.Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.If you thrive in environments where you're handed a problem rather than a playbook this role is for you.What You'll OwnDefine and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually useDrive the reliability and operational excellence of the Forward SaaS platformBuild and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers doLead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twicePartner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviewsHelp define and build the SRE team as the company scales — this is a foundational hire with a path to leadershipWhat We're Looking For6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environmentProven experience building or significantly maturing an SRE function — not just operating within one someone else builtStrong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plusHands-on experience with Kubernetes and container orchestration in production environmentsDeep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similarStrong scripting and automation skills in Python, Bash, or similarExperience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvementsAbility to communicate clearly with both engineering teams and non-technical stakeholders — you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing themNice to HaveExperience supporting enterprise or federal government customers with high availability requirementsExperience in a foundational or early SRE hire capacity at a growth stage companyWhat This Role Is NotA pure ops or NOC role — you are building and engineering, not just monitoringA siloed function — you will be deeply embedded with product and engineering teamsA ticket-taker — you will be proactively identifying and solving reliability problems before they become incidentsWhy ForwardYou'll be building something from scratch at a company with real enterprise traction and world-class investors behind itOur customers include some of the most complex network environments on the planet — the reliability bar is high and the work is genuinely interestingPeople-centric culture built by Stanford Ph.D.s who care deeply about doing things the right wayCompetitive compensation, equity, and the opportunity to grow into a leadership role as the SRE function scalesThe base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Santa Clara, CA vacancy
  •  ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation... 
    Suggested

    Oracle

    Santa Clara, CA
    2 days ago
  •  ...function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.As a... 
    Suggested
    Shift work

    CEREBRAS SYSTEMS INC.

    Sunnyvale, CA
    4 days ago
  •  ...in Cupertino, California, invites an experienced CDN Solutions Engineer to join the Content Delivery Network Solutions team. You will...  ...and collaborate with engineering groups across Apple to ensure reliable delivery at scale. The ideal candidate has 4+ years in CDNs and... 
    Suggested

    Apple

    Cupertino, CA
    4 days ago
  •  ...ServiceNow in Santa Clara, CA, seeks a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, resilience, and toil...  ...for global engineering teams. Embedded within the Site Reliability & Database Engineering organization, you will architect SRE tooling... 
    Suggested

    ServiceNow

    Santa Clara, CA
    2 days ago
  •  ...automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and... 
    Suggested
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    2 days ago
  • $276.1k - $311.4k

     ...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a... 
    Permanent employment
    Full time
    Work at office
    Work from home

    Lindus Health

    Sunnyvale, CA
    1 day ago
  • $170k - $200k

     ...We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    1 day ago
  •  ...The RoleThis hybrid role combines the hands-on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE).The ideal candidate has a strong technical foundation, thrives in a... 
    Full time
    Local area

    F5 Networks

    San Jose, CA
    1 day ago
  • $145k - $175k

     ...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data... 
    Full time
    Remote work

    GrabJobs

    Santa Clara, CA
    4 days ago
  • $187.04k - $359.72k

     ...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum...  ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas....  ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company... 
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    3 days ago
  •  ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable... 

    Epic Games

    Sunnyvale, CA
    3 days ago
  •  ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud... 
    Full time

    Saransh

    Sunnyvale, CA
    3 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    a month ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    a month ago
  •  ...Job Description Job Description Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots... 
    Permanent employment
    Full time
    Work at office
    Local area

    Foxconn Industrial Internet - FII

    San Jose, CA
    26 days ago
  •  ...of Huobi globe spanning infrastructure. •       Work with engineering teams to make sure new features and changes are deployed quickly...  .... •       Constantly improve our system performance and reliability through better tools, process and monitoring system. •... 
    Worldwide

    Cryptoware Technologies Inc

    Santa Clara, CA
    26 days ago
  • $160k - $240k

     ...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit...  ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in... 
    Full time

    Fiserv

    Sunnyvale, CA
    18 days ago
  •  ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the... 
    Work at office

    Foxconn Industrial Internet - FII

    Sunnyvale, CA
    26 days ago
  •  ...to join IBM in a full‑time role between December 2027 and August 2028 upon successful completion of their degree. As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client... 
    Full time
    Contract work
    Part time
    Fixed term contract
    Internship
    Worldwide
    Flexible hours
    Shift work

    IBM

    San Jose, CA
    4 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with... 
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $122.5k - $175k

     ...impact at the company pioneering security transformation in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    1 day ago
  •  ...Job Description Job Description Site Reliability Engineer II San Fran Bay · Hybrid · 24/7 FedRAMP Operations · Rotational Shift · Initial Contract till March 27. KEY REQUIREMENT This role requires US citizenship and residence on US soil. It sits within a... 
    Hourly pay
    Contract work
    For contractors
    Shift work
    Night shift
    Weekend work

    C-Serv

    San Jose, CA
    2 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    25 days ago
  •  ...Must Have Technical/Functional Skills: 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role supporting cloud-based production environments. Practical vulnerability-management experience; familiarity with... 
    Full time
    Worldwide

    Siri InfoSolutions Inc

    San Jose, CA
    2 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Full time

    Nvidia

    Santa Clara, CA
    25 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    22 days ago
  • $207k - $300k

     ...areas within SU SRE, mentoring team members to enhance system reliability and efficiency.Initiate, own, and lead large-scale,...  ...Design for Reliability techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading projects.3 years of experience... 

    Google

    San Jose, CA
    7 days ago
  • $207.4k - $259.2k

     ...built specifically for aviation. We’re seeking exceptional engineers, operators and builders to join us on our mission to build the...  ...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role... 
    Permanent employment
    Local area
    Visa sponsorship
    Night shift

    Archer Aviation

    San Jose, CA
    7 days ago
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized... 
    Full time

    Nvidia

    Santa Clara, CA
    11 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!