Site Reliability Engineer
$230k - $250kForward Networks
Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment.Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.If you thrive in environments where you're handed a problem rather than a playbook this role is for you.What You'll OwnDefine and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually useDrive the reliability and operational excellence of the Forward SaaS platformBuild and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers doLead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twicePartner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviewsHelp define and build the SRE team as the company scales — this is a foundational hire with a path to leadershipWhat We're Looking For6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environmentProven experience building or significantly maturing an SRE function — not just operating within one someone else builtStrong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plusHands-on experience with Kubernetes and container orchestration in production environmentsDeep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similarStrong scripting and automation skills in Python, Bash, or similarExperience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvementsAbility to communicate clearly with both engineering teams and non-technical stakeholders — you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing themNice to HaveExperience supporting enterprise or federal government customers with high availability requirementsExperience in a foundational or early SRE hire capacity at a growth stage companyWhat This Role Is NotA pure ops or NOC role — you are building and engineering, not just monitoringA siloed function — you will be deeply embedded with product and engineering teamsA ticket-taker — you will be proactively identifying and solving reliability problems before they become incidentsWhy ForwardYou'll be building something from scratch at a company with real enterprise traction and world-class investors behind itOur customers include some of the most complex network environments on the planet — the reliability bar is high and the work is genuinely interestingPeople-centric culture built by Stanford Ph.D.s who care deeply about doing things the right wayCompetitive compensation, equity, and the opportunity to grow into a leadership role as the SRE function scalesThe base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location
- ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation...Suggested
- ...function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.As a...SuggestedShift work
- ...in Cupertino, California, invites an experienced CDN Solutions Engineer to join the Content Delivery Network Solutions team. You will... ...and collaborate with engineering groups across Apple to ensure reliable delivery at scale. The ideal candidate has 4+ years in CDNs and...Suggested
- ...ServiceNow in Santa Clara, CA, seeks a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, resilience, and toil... ...for global engineering teams. Embedded within the Site Reliability & Database Engineering organization, you will architect SRE tooling...Suggested
- ...automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and...SuggestedLocal area
$276.1k - $311.4k
...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a...Permanent employmentFull timeWork at officeWork from home$170k - $200k
...We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance...Full timeWorldwide- ...The RoleThis hybrid role combines the hands-on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE).The ideal candidate has a strong technical foundation, thrives in a...Full timeLocal area
$145k - $175k
...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data...Full timeRemote work$187.04k - $359.72k
...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum... ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas.... ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company...Temporary workLocal areaOverseasShift work- ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable...
- ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud...Full time
- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...Work at officeLocal areaWork from homeFlexible hours
$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...Full time- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...Full timeWork at office2 days per week
- ...Job Description Job Description Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots...Permanent employmentFull timeWork at officeLocal area
- ...of Huobi globe spanning infrastructure. • Work with engineering teams to make sure new features and changes are deployed quickly... .... • Constantly improve our system performance and reliability through better tools, process and monitoring system. •...Worldwide
$160k - $240k
...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit... ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in...Full time- ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the...Work at office
- ...to join IBM in a full‑time role between December 2027 and August 2028 upon successful completion of their degree. As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client...Full timeContract workPart timeFixed term contractInternshipWorldwideFlexible hoursShift work
$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with...Full time$122.5k - $175k
...impact at the company pioneering security transformation in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect...Full timeWork at officeLocal area3 days per week- ...Job Description Job Description Site Reliability Engineer II San Fran Bay · Hybrid · 24/7 FedRAMP Operations · Rotational Shift · Initial Contract till March 27. KEY REQUIREMENT This role requires US citizenship and residence on US soil. It sits within a...Hourly payContract workFor contractorsShift workNight shiftWeekend work
$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...Work experience placementWork at officeLocal areaWork from homeFlexible hours- ...Must Have Technical/Functional Skills: 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role supporting cloud-based production environments. Practical vulnerability-management experience; familiarity with...Full timeWorldwide
$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...Full time- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...Work at officeLocal areaWork from homeFlexible hours
$207k - $300k
...areas within SU SRE, mentoring team members to enhance system reliability and efficiency.Initiate, own, and lead large-scale,... ...Design for Reliability techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading projects.3 years of experience...$207.4k - $259.2k
...built specifically for aviation. We’re seeking exceptional engineers, operators and builders to join us on our mission to build the... ...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role...Permanent employmentLocal areaVisa sponsorshipNight shift$168k - $270.25k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Santa Clara, CA
- site reliability engineer Santa Clara, CA
- site services specialist Santa Clara, CA
- junior website developer Santa Clara, CA
- official site Santa Clara, CA
- site leader Santa Clara, CA
- historic site Santa Clara, CA
- construction site safety Santa Clara, CA
- IT site lead Santa Clara, CA
- website content developer Santa Clara, CA



