Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$230k - $250k

Forward Networks

Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment.Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.If you thrive in environments where you're handed a problem rather than a playbook this role is for you.What You'll OwnDefine and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually useDrive the reliability and operational excellence of the Forward SaaS platformBuild and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers doLead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twicePartner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviewsHelp define and build the SRE team as the company scales — this is a foundational hire with a path to leadershipWhat We're Looking For6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environmentProven experience building or significantly maturing an SRE function — not just operating within one someone else builtStrong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plusHands-on experience with Kubernetes and container orchestration in production environmentsDeep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similarStrong scripting and automation skills in Python, Bash, or similarExperience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvementsAbility to communicate clearly with both engineering teams and non-technical stakeholders — you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing themNice to HaveExperience supporting enterprise or federal government customers with high availability requirementsExperience in a foundational or early SRE hire capacity at a growth stage companyWhat This Role Is NotA pure ops or NOC role — you are building and engineering, not just monitoringA siloed function — you will be deeply embedded with product and engineering teamsA ticket-taker — you will be proactively identifying and solving reliability problems before they become incidentsWhy ForwardYou'll be building something from scratch at a company with real enterprise traction and world-class investors behind itOur customers include some of the most complex network environments on the planet — the reliability bar is high and the work is genuinely interestingPeople-centric culture built by Stanford Ph.D.s who care deeply about doing things the right wayCompetitive compensation, equity, and the opportunity to grow into a leadership role as the SRE function scalesThe base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Santa Clara, CA vacancy
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are... 
    Suggested
    Remote work

    Abbott

    Sunnyvale, CA
    3 days ago
  • $128k - $216k

     ...consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay...  ...scale, come make a difference at Fiserv. Job Title Sr. Site Reliability Engineer About Clover Clover is a pioneer in the fintech space,... 
    Suggested
    Full time
    Worldwide

    Fiserv

    Sunnyvale, CA
    4 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Suggested
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $160k - $240k

     ...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit...  ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in... 
    Suggested
    Full time

    Fiserv

    Sunnyvale, CA
    5 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Suggested
    Flexible hours

    Sumo Logic

    San Jose, CA
    4 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    3 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    1 day ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Role: Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job Description: Deploy, support and monitor new and existing services, platforms, and application stacks.... 

    Zortech Solutions

    Santa Clara, CA
    4 days ago
  •  ...Overview: *Must have Apple experience* • At least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Advanced experience with programming languages (Python, Java) • Passion for designing and building reliable systems • Strong sense... 

    Purple Drive

    Sunnyvale, CA
    5 days ago
  •  ...on glass. Hands on the pipeline. Real ownership from day one. This isn't a watch-and-wait monitoring seat. Our client needs engineers who can read a Kibana query at 3am, know the difference between a blip and a breach, and act on it, on a FedRAMP-authorised cloud... 
    Hourly pay
    For contractors
    Shift work
    Night shift
    Weekend work

    C-Serv

    San Jose, CA
    2 days ago
  •  ...Senior Site Reliability Engineer Location: Remote Duration: 12 month contract to start IV Process: 1-3 Round IV process International Tech Top Skills: Java Python NodeJS -DevOps Engineer should work here too Main Responsibilities:... 
    Contract work
    Local area
    Remote work

    My3Tech Inc

    Sunnyvale, CA
    4 days ago
  • $145k - $165k

     ...Your Ego : Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    2 days ago
  •  ...Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots. FII provides customers with intelligent... 
    Permanent employment
    Full time
    Work at office
    Local area

    Foxconn Industrial Internet

    San Jose, CA
    2 days ago
  •  ...AWS Infra SRE/DevOps Engineer AWS Infra SRE/DevOps engineer with proven work experience ensuring reliability, availability and performance of cloud infra and platform. Specialist on Cisco Cloud run-on for infrastructure management, who can install, run, and maintain... 
    Work experience placement

    The Dignify Solutions, LLC

    San Jose, CA
    4 days ago
  •  ...Site Reliability Engineer Sunnyvale, CA Site Reliability Eng. Must have LinkedIn profile. • t least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Strong AWS and Linux Operating System, standard networking protocols, component... 

    Q1 Technologies

    Sunnyvale, CA
    1 day ago
  • $128.6k - $184.9k

     ...global cloud platform. As a team of six engineers distributed across the US, Canada, and the...  ...with a strong focus on automation, reliability, and operational excellence. We are one...  ...Qualifications ~7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    San Jose, CA
    5 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $101k - $161k

     ...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-...  ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-... 

    Arista Networks

    Santa Clara, CA
    4 days ago
  •  ...Qualifications: 8+ years of software engineering experience, or equivalent demonstrated through...  ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform...  ...vendor resources. Willingness to work on-site at stated location in the job opening.... 
    For contractors
    Work experience placement

    Cedent Life Talent

    San Jose, CA
    5 days ago
  •  ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of...  ...are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    1 day ago
  • $65 - $85 per hour

     ...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    2 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $122.5k - $175k

     ...we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    2 days ago
  • $210.6k - $305.1k

     ...Minimum Qualifications:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure...  ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    3 days ago
  • $192.4k - $275.8k

     ...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    22 hours ago
  • $207k - $300k

     ...areas within SU SRE, mentoring team members to enhance system reliability and efficiency.Initiate, own, and lead large-scale,...  ...Design for Reliability techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading projects.3 years of experience... 

    Google

    San Jose, CA
    22 hours ago
  • $207k - $300k

     ...automation, and evolve systems by pushing for changes that improve reliability and velocity.Practice sustainable incident response and...  ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems... 

    Google

    San Jose, CA
    22 hours ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!