Site Reliability Engineer
CEREBRAS SYSTEMS INC.
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.About the RoleWe are building a high-performance SRE function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.As a Principal SRE, you will define and drive the technical architecture for scaling our inference fleet through self-service delivery, shared observability, capacity orchestration, rollout safety, and operational automation. This role starts with 2–3 weeks of hands-on operational immersion to build deep context on the current stack, production pain points, and high-stakes workflows.From there, your mandate shifts to architecting the “tomorrow” layer: a unified capacity management and production control plane that enables reliable capacity planning, workload placement, rollout safety, validation, and operational decision-making across large-scale inference infrastructure.Success in the first year means core engineering teams, product managers, external customers, and cluster stakeholders can execute critical operational workflows through self-service systems with strong guardrails, clear ownership, and minimal dependency on expert SRE operators.You will collaborate with the tech leads and the leadership team across core, cluster, cloud, and product stakeholders. This work will shift reliability from an ops-only burden to a shared engineering discipline that underpins frontier AI inference at scale.If you are a proven Principal engineer who enjoys turning complexity into elegant reliability at scale, this is your chance to lead this transformation from the front.This role does not require 24/7 on-call rotations.Key ResponsibilitiesDefine and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.Architect self-service platforms and internal tooling that let product teams, external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.Define and evolve reliability practices for inference workloads, including SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi-datacenter and on-prem environments.Mentor senior SREs, support critical incident escalations, and use production pain points to prioritize the highest-leverage automation work.Measure and drive impact through clear metrics, including toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.Required Experience & Skills15+ years in SRE, infrastructure engineering, or platform engineering, with a record of setting technical direction and delivering reliability improvements at large scale in FAANG, hyperscaler, frontier AI, or similarly demanding production environments.Deep experience with large-scale compute fleets, internal control planes, schedulers, orchestration systems, capacity management, and reliability automation.Experience defining and driving cross-team architecture for production control planes, capacity orchestration, fleet management, or self-service infrastructure platforms with clear operational ownership.Strong judgment in converging fragmented workflows, tools, and teams into coherent architectures that improve reliability, efficiency, and operational leverage.Ability to lead complex, ambiguous technical programs end to end; influence senior cross-functional stakeholders; mentor senior engineers; and communicate technical strategy clearly.Hands-on experience with production observability, incident response, and SLO-based reliability management across metrics, logs, traces, alerting, dashboards, and operational review loops.Nice-to-HavesExperience with Bazel or other large-scale build systems in production.Background in AI/ML inference systems, including model serving runtimes, disaggregated inference, GPU orchestration, latency and accuracy SLOs, or drift monitoring.Prior work on predictive autoscaling, chaos engineering, or cost-aware capacity management for compute-intensive workloads.LocationSF Bay AreaTorontoWhy Join CerebrasPeople who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:Build a breakthrough AI platform beyond the constraints of the GPU.Publish and open source their cutting-edge AI research.Work on one of the fastest AI supercomputers in the world.Enjoy job stability with startup vitality.Our simple, non-corporate work culture that respects individual beliefs.Find out more about what it's like to work at Cerebras here! Apply today and become part of the forefront of groundbreaking advancements in AI!Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.LocationSunnyvale, CAEmployment TypeFull timeLocation TypeHybridDepartmentSoftware Engineering
$150.4k - $277.6k
...Technical Operations & Site Reliability Engineer, Customer SystemsAt Apple, Customer Experience is at the forefront of everything we do. The Customer Systems Operations team is looking for a highly skilled and motivated TechOps Engineer (Technical Operations & Site Reliability...SuggestedWork experience placementRelocation- ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation...Suggested
$276.1k - $311.4k
...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a...SuggestedPermanent employmentFull timeWork at officeWork from home- ...in Cupertino, California, invites an experienced CDN Solutions Engineer to join the Content Delivery Network Solutions team. You will... ...and collaborate with engineering groups across Apple to ensure reliable delivery at scale. The ideal candidate has 4+ years in CDNs and...Suggested
- ...ServiceNow in Santa Clara, CA, seeks a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, resilience, and toil... ...for global engineering teams. Embedded within the Site Reliability & Database Engineering organization, you will architect SRE tooling...Suggested
- ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud...Full time
- ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable...
$170k - $200k
...We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance...Full timeWorldwide- ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data...
$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...Full time- ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the...Work at office
- ...of Huobi globe spanning infrastructure. • Work with engineering teams to make sure new features and changes are deployed quickly... .... • Constantly improve our system performance and reliability through better tools, process and monitoring system. •...Worldwide
$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with...Full time$101k - $161k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-... ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-...- ...Job Description Job Description We are hiring Site Reliability Engineer Manager- Hybrid for a Contract To Hire position in santa clara, CA The Role You will build and lead the Site Reliability Engineering team, owning the infrastructure that development, validation...Contract workRemote work
$255.7k - $300k
...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering ~ link Copy link corporate_fare Google place Sunnyvale, CA, USA Advanced Experience owning outcomes and decision making, solving ambiguous...Full timeWork at office- ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available... ...and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and infrastructure...
- ...exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of... ...and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase...
- ...automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and...Local area
$187.04k - $359.72k
...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum... ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas.... ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company...Temporary workLocal areaOverseasShift work$169k - $338k
...advanced agentic AI systems that can autonomously handle complex reliability engineering workflows, predictive failure analysis, and self-... ...associates, or business operations across any Walmart system.Site Reliability Engineering Technical Excellence:Design, write and...Full timeTemporary workPart time- ...The RoleThis hybrid role combines the hands-on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE).The ideal candidate has a strong technical foundation, thrives in a...Full timeLocal area
$165k - $280k
...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s...Permanent employmentTemporary workWorldwideWeekend work- ...Overview We are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of operational excellence... ...company supporting sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps will embrace our commitment...For contractorsWork at officeFlexible hours
- ...Job Description Job Description Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots...Permanent employmentFull timeWork at officeLocal area
$175k - $263k
...our platforms remain at the cutting edge of performance and reliability. WHAT YOU'LL DO Architect High-Performance Networking:... ...interact directly with ASIC capabilities. Availability-Focused Engineering: Understanding of non-disruptive upgrade (NDU) technologies...Work at officeFlexible hours- ...to join IBM in a full‑time role between December 2027 and August 2028 upon successful completion of their degree. As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client...Full timeContract workPart timeFixed term contractInternshipWorldwideFlexible hoursShift work
$122.5k - $175k
...impact at the company pioneering security transformation in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect...Full timeWork at officeLocal area3 days per week$100k - $200k
OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time$160k - $240k
...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit... ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Sunnyvale, CA
- site recruiter Sunnyvale, CA
- junior website developer Sunnyvale, CA
- official site Sunnyvale, CA
- site leader Sunnyvale, CA
- construction site safety Sunnyvale, CA
- IT site lead Sunnyvale, CA
- website content developer Sunnyvale, CA
- on-site clinical research associate (traveling/remote) Sunnyvale, CA
- site safety Sunnyvale, CA




