Principal SRE - AI Inference
Cerebras Systems
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.About the RoleWe are building a high-performance SRE function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.As a Principal SRE, you will define and drive the technical architecture for scaling our inference fleet through self-service delivery, shared observability, capacity orchestration, rollout safety, and operational automation. This role starts with 2–3 weeks of hands-on operational immersion to build deep context on the current stack, production pain points, and high-stakes workflows.From there, your mandate shifts to architecting the “tomorrow” layer: a unified capacity management and production control plane that enables reliable capacity planning, workload placement, rollout safety, validation, and operational decision-making across large-scale inference infrastructure.Success in the first year means core engineering teams, product managers, external customers, and cluster stakeholders can execute critical operational workflows through self-service systems with strong guardrails, clear ownership, and minimal dependency on expert SRE operators.You will collaborate with the tech leads and the leadership team across core, cluster, cloud, and product stakeholders. This work will shift reliability from an ops-only burden to a shared engineering discipline that underpins frontier AI inference at scale.If you are a proven Principal engineer who enjoys turning complexity into elegant reliability at scale, this is your chance to lead this transformation from the front.This role does not require 24/7 on-call rotations.Key ResponsibilitiesDefine and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.Architect self-service platforms and internal tooling that let product teams, external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.Define and evolve reliability practices for inference workloads, including SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi-datacenter and on-prem environments.Mentor senior SREs, support critical incident escalations, and use production pain points to prioritize the highest-leverage automation work.Measure and drive impact through clear metrics, including toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.Required Experience & Skills15+ years in SRE, infrastructure engineering, or platform engineering, with a record of setting technical direction and delivering reliability improvements at large scale in FAANG, hyperscaler, frontier AI, or similarly demanding production environments.Deep experience with large-scale compute fleets, internal control planes, schedulers, orchestration systems, capacity management, and reliability automation.Experience defining and driving cross-team architecture for production control planes, capacity orchestration, fleet management, or self-service infrastructure platforms with clear operational ownership.Strong judgment in converging fragmented workflows, tools, and teams into coherent architectures that improve reliability, efficiency, and operational leverage.Ability to lead complex, ambiguous technical programs end to end; influence senior cross-functional stakeholders; mentor senior engineers; and communicate technical strategy clearly.Hands-on experience with production observability, incident response, and SLO-based reliability management across metrics, logs, traces, alerting, dashboards, and operational review loops.Nice-to-HavesExperience with Bazel or other large-scale build systems in production.Background in AI/ML inference systems, including model serving runtimes, disaggregated inference, GPU orchestration, latency and accuracy SLOs, or drift monitoring.Prior work on predictive autoscaling, chaos engineering, or cost-aware capacity management for compute-intensive workloads.LocationSF Bay AreaTorontoWhy Join CerebrasPeople who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:Build a breakthrough AI platform beyond the constraints of the GPU.Publish and open source their cutting-edge AI research.Work on one of the fastest AI supercomputers in the world.Enjoy job stability with startup vitality.Our simple, non-corporate work culture that respects individual beliefs.Find out more about what it's like to work at Cerebras here! Apply today and become part of the forefront of groundbreaking advancements in AI!Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.LocationHeadquarters/Sunnyvale OfficeEmployment TypeFull timeLocation TypeOn-siteDepartmentSoftware
- ...are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the... ...Clara, CA, headquarters 3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires you to be...Principal3 days per week
- Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This... ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based... ...inference.About The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform...Principal
$182.5k - $260.5k
...is a leader in modern security and networking for the cloud and AI era. We secure and accelerate cloud, data, and AI in real time,... ...the roleAs a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast,...Principal$278.1k - $347.6k
...opportunity We are building the next generation of AI-driven game experiences, running generative models on-... ...deployed and accelerated entirely within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you will be the foremost engineering...PrincipalWork at officeWorldwideRelocation package- ...that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded... ...Together, we advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications...Principal
$152k - $241.5k
...are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping... ...13, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering...Full time$152k - $241.5k
NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing... ....Collaborate with model, platform, and SRE teams to translate production requirements...Full timeRemote work$148k - $235.75k
...in our rapidly growing data center business and pivotal in our inference marketing. You will be focused on working with engineering to understand... ...marketing strategy to showcase our leadership position in AI inference.Want to join a fun, creative company that is at the...Full time$272k - $431.25k
...large scale workloads for ETL, SQL, and ML/DL model training and inference pipelines, spanning many domains and use cases. NVIDIA GPUs... ...accelerate Apache Spark with GPUs. You will apply the latest ML/AI methods to empower enterprises to migrate Spark workloads onto GPUs...PrincipalFull time$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry...Full time$207k - $301k
...Collaboration: Partner closely with Research, SRE, Product, and Core GPU library teams to... ...team efforts with broader organizational AI priorities.Minimum qualifications:Bachelor... ...internationally.The Distributed Cloud (DSC) AI Inference Platform team operates at the critical...$224k - $356.5k
...groups around the world are using GPUs to power a revolution in AI, enabling breakthroughs in problems from image classification to... ...language processing. We are a fast-paced team building Generative AI inference platform to make design and deployment of new AI models easier...Full time$193.3k - $261.5k
...popular ML frameworks like PyTorch and JAX enabling unparalleled ML inference and training performance.The Inference Enablement and... ...with ML expertise to push the boundaries of what's possible in AI acceleration.As part of the broader Neuron organization, our team...Work experience placementInternshipLocal areaFlexible hours- ...next-generation computing experiences—from AI and data centers, to PCs, gaming and... ...Together, we advance your career. THE ROLE:As a Principal Engineer, you will spearhead the next... ...x performance gains in both training and inference pipelines through innovative system design...PrincipalRemote work
$272k - $431.25k
...amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU... ...clusters used for distributed Deep Learning LLM training and inference. Your primary focus will be collectives communication and networking...PrincipalFull timeRemote work$185k - $299k
...with voices, and find opportunity. As a Principal Product Manager on the Feed Relevance team... ...Product Manager to lead the next generation of AI-driven ranking and recommendation systems... ...teams to scale model training, inference, and evaluation infrastructure that supports...PrincipalFor contractorsWork at officeFlexible hours$143k - $286k
...Summary...About the RoleWe are looking for a Principal Data Scientist to lead high-impact... ...building scalable, end-to-end Data Science and AI solutions that improve customer... ...machine learning, experimentation, causal inference, and data science methodologies.Experience...PrincipalFull timeTemporary workPart time- ...products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded... ...models, and enabling RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node systems. You will collaborate...
$110k - $220k
...OfficePosition Summary...What you'll do...The Principal Data Scientist in the Data Science team... ...a crucial role as an architect, leading AI explorations, research, algorithmic... ...experiences via Insights, frameworks, causal inference solutions and machine learning prototypes...PrincipalFull timeTemporary workPart timeShift work$132k - $264k
...classification models, regression models, NLP, forecasting, unsupervised models, optimization, graph ML, causal inference, causal ML, statistical learning, experimentation, and Gen-AI.In Gen-AI, it is desirable to have experience in embedding generation from training materials,...PrincipalFull timeTemporary workPart timeFlexible hours$253.1k - $342.3k
...set AWS’s services and features apart in the industry. Come optimize and deploy inference models on Trainium, Amazon's custom cloud-scale machine learning accelerators that power the latest AI modelsAs a Sr. SDM for the Inference Team, you will lead a strong team of...Local areaFlexible hours$110k - $220k
...Segment: Home OfficePosition Summary...As a Principal Data Scientist at Walmart, you will... ...experimentation science, large-scale data systems, and AI evaluation. You will own the scientific... ...expertise in experimentation, causal inference, and statistical decision-making, with a...PrincipalFull timeTemporary workPart time- Overview As a Principal Machine Learning Engineer, you will drive the development and implementation... ..., and analytics teams to build AI functionalities into Atlassian products... ...features for offline training and online inference at scaleOversee end-to-end deployment of...PrincipalLocal areaRemote work
$152k - $241.5k
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery... ...fundamentals: concise, well-tested code; fluent with AI-assisted toolingExperience with ML accelerators with a basic understanding...Full time$208.3k - $281.8k
...chips in production, used for training and inference of frontier models. AWS Neuron is the... ...customers to run deep learning and generative AI workloads with optimal performance and cost efficiency. AWS Neuron is hiring a Principal Technical Product Manager to define and drive...PrincipalLocal areaFlexible hours$272k - $431.25k
We're looking for a Principal Engineer to join our CSP Engagements team as the technical focal... ...production environmentsUnderstanding of inference workload performance dynamics (vLLM,... ...is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA...PrincipalFull timeRemote work- ...next-generation computing experiences—from AI and data centers, to PCs, gaming and... ...Together, we advance your career. THE ROLE:As a Principal AI Infrastructure Solution Engineer, you... ...to enable large‑scale LLM training and inference on AMD Instinct GPUs. You will design and...Principal
$272k - $431.25k
...NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments... ...of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap for...PrincipalFull timeLocal areaRemote work$152k - $241.5k
...parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain... ...advancement of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices, automotive...Full timeRemote work$296.3k
...location three times a week, at minimum.The Role:We are seeking a Principal AI Engineer to lead the design and advancement of our AI... ...the infrastructure that powers large-scale training and cloud inference. This includes accelerating training throughput, scaling multi...PrincipalFull timeLocal areaRemote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal SRE - AI Inference. Be the first to apply!

