Director/Sr. Manager, AI Inference Model Scaling
Cerebras Systems
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.Sunnyvale, CA or Toronto, Canada (Hybrid)About the TeamThe Inference Model Scaling team enables state-of-the-art foundation models and generative AI workloads to run efficiently on Cerebras' Wafer-Scale Engine (WSE). We build the compiler frontend, model transformation pipeline, graph optimization infrastructure, high-performance kernel enablement, and runtime integration that together make next-generation AI models execute with industry-leading performance.The team works at the intersection of machine learning frameworks, compiler technologies, distributed systems, hardware architecture, and model optimization. We collaborate closely with hardware architects, runtime engineers, cloud platform teams, AI researchers, and strategic customers to rapidly bring new model architectures into production.About the RoleWe're looking for an experienced engineering leader to build and scale our Inference Model Scaling organization.You will define the technical vision, organizational strategy, and execution roadmap for a globally distributed engineering team responsible for enabling the latest foundation models on Cerebras hardware. You will lead the engineering organization responsible for ML model compilation and optimization as well as development of high-performance kernels.This role combines deep technical leadership with organizational excellence. You will partner across compiler, runtime, cloud infrastructure, hardware architecture, product management, and AI research teams while helping shape the future of AI inference at Cerebras.ResponsibilitiesTechnical LeadershipDefine the technical roadmap and strategy for the team.Establish technical direction across multiple teams and engineering leaders.Lead design reviews and establish engineering standards.Drive support for emerging LLM architectures and inference workloads.Team LeadershipHire, mentor, and grow a high-performing engineering team.Develop future technical leaders and managers.Drive organizational planning, headcount strategy, and investment priorities.Foster a strong engineering culture focused on execution, quality, and innovation.Scale engineering processes while maintaining execution velocity.Cross-Functional CollaborationPartner with Cloud Platform, ML, and Hardware teams in planning and delivering for end-to-end service enablement in Cloud and On-Premise settingsWork with Product Management to prioritize model enablement and customer needs.Collaborate closely with customers and solution architects on new model bring-up.Influence future hardware/software co-design through ML model enablement and optimization insights.Delivery & ExecutionOwn planning, prioritization, and execution across multiple concurrent initiatives.Balance rapid model support with long-term ML Compiler architecture.Drive predictable delivery for strategic customer commitments.Required QualificationsBS, MS, or PhD in Computer Science, Computer Engineering or related field.12+ years building compiler, ML systems, or infrastructure software.5+ years leading engineering teams.Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).Strong understanding of graph compilation and optimization.Experience with Python and C++.Experience delivering production-quality software.Strong communication and cross-functional leadership skills.Preferred QualificationsExperience building compiler frontends for AI accelerators.Experience supporting PyTorch, JAX, TensorFlow, or ONNX.Experience with LLM inference or training systems.Familiarity with distributed compilation.Experience working with hardware architects.Experience leading teams through rapid growth.Why Join CerebrasPeople who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:Build a breakthrough AI platform beyond the constraints of the GPU.Publish and open source their cutting-edge AI research.Work on one of the fastest AI supercomputers in the world.Enjoy job stability with startup vitality.Our simple, non-corporate work culture that respects individual beliefs.Find out more about what it's like to work at Cerebras here! Apply today and become part of the forefront of groundbreaking advancements in AI!Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.LocationSunnyvale, CAEmployment TypeFull timeLocation TypeHybridDepartmentSoftware Engineering
$193.3k - $261.5k
..., Amazon's custom cloud-scale machine learning accelerators... ...to optimize the latest models to run really fast on... ...Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you... ...engineers, and product managers to architect and deliver...SeniorInternshipLocal areaFlexible hours$193.3k - $261.5k
...enabling unparalleled ML inference and training... ...running a wide range of models and supporting novel architecture... ...of what's possible in AI acceleration.As part of... ...peak performance at scale for customers and developers... ...engineers, and product managers to deliver state-of-the...SeniorWork experience placementInternshipLocal areaFlexible hours$219k - $351k
...becoming a memory-bandwidth business. As models scale past what any single GPU can hold — KV... ...treat memory as the core product of AI inference, not an afterthought.We are looking... ...class engines — including their memory-management internals, not just their flags.Drive...SuggestedWork at officeRemote workFlexible hours$200k - $322k
...real-world product outcomes at scale. We are hiring a Senior TPM... ...The exceptional hire also uses AI deliberately: they have hands-... ...experience with AI-powered program management tools (e.g., automated status,... ...working with AI/ML teams on inference infrastructure and aligning...SeniorFull time3 days per week$171k - $280k
...groundbreaking initiative to develop agentic models. We are building advanced AI agents that exhibit reasoning,... ...of work.The Technical Program Manager (level is TBD ) will have the opportunity... ...and executing intermediate to large scale, cross-functional and company-wide...SeniorTemporary workFor contractorsWork at officeFlexible hours- ...work. Today, ServiceNow is the AI control tower for business reinvention... .... As an Outbound Product Manager, you’ll play a key role in... ...for all employees.The Role As a Sr. Staff Outbound Product Manager... ...Demonstrated experience shipping and scaling integrations/connectors,...SeniorWork at officeImmediate startRemote workFlexible hours
$300k - $333k
...style, tone, and behavior with the model specifications and taxonomy, scaling processes for consistent evaluations... ...relying on direct authority while managing high-visibility executive reviews,... ...tolerance for ambiguity inherent in AI research and operates at a fast pace...Senior$166k - $271k
...Engineering Technical Program Management Group is looking for a Senior... ...org builds and operates massive-scale systems: physical data centers... ..., observability, and Data and AI/ML platforms.Engineering Team... ...fundamentals, including consistency models, durability, fault tolerance,...SeniorFor contractorsWork at officeFlexible hours$189.8k - $256.16k
...About Databricks Databricks is the data and AI company. More than 10,000 organizations... ...exceptional Senior Staff Technical Program Manager (TPM) for Reliability to lead the strategy... ...engineering teams at Databricks. As Databricks scales to support thousands of customers and the...SeniorLocal areaWorldwide$143.56k - $215k
...Across enterprise, cloud and AI, and carrier architectures, our... ....What You Can ExpectAs a Sr. Staff Manager, Product Engineering in Marvell... ...action on poor performance.· Model and hold others accountable for... ...and supply chain to scale production capacity, qualify...SeniorPermanent employmentInternshipWork from home$152k - $230k
...into the unlimited potential of AI to define the next era of... ...build with open, customizable models. We seek a Senior Product Marketing... ...develop market strategy and manage the entire launch-to-adoption... ...Experience marketing AI models, inference platforms, developer software,...SeniorFull time$192k - $264k
...the exciting technologies that literally connect our world - like AI and IoT. If you want to push the boundaries of materials science... ...approval guidelines and leadership to reporting engineering managers in the area of long-term program, strategy, and process design.Defines...SeniorFull time$184k - $253k
...Sonatus, we’re driving the transformation to AI-enabled software-defined vehicles.... ...agility of a fast-growing company with the scale and impact of an established partner. Backed... ...for a highly motivated Technical Program Manager to define, manage, and support Sonatus Engineering...SeniorWork at officeWorldwideFlexible hoursShift work$173.9k - $235.2k
...eliminating the carbon footprint of our products at a scale and speed that manual processes simply can't match.... ...slow, manual sustainability analysis into fast, AI-assisted work, so our scientists and program managers can focus on the highest-impact carbon-reduction work...SeniorLocal areaFlexible hours$190.9k - $334.1k
...Today, ServiceNow is the AI control tower for... ...Staff Technical Program Manager to lead our most critical... ...portfolio over time as the BU scales. This is a senior,... ...logs that engineering directors and VPs actually read.... ...you will serve as the model for how to operate. Set...SeniorWork at officeImmediate startRemote workFlexible hours$168k - $258.75k
At NVIDIA, we are at the forefront of AI and accelerated computing, redefining the world... ...society. As part of the Customer Program Management Team, you will lead strategic technical... ..., delivering unmatched data center-scale performance.What you'll be doing:Leading...SeniorFull time$143.5k - $205k
...future of work is Human + AI and are building an... ...are looking for a Sr. Staff Technical Program Manager - Service Health to join... ...to the Senior Director of Site Reliability Engineering... ...at global scale.What you’ll do (Role... ...Zscaler's hybrid working model and benefits here.By...SeniorFull timeWork at officeLocal area$336k
...time-sensitive products to support growing AI infrastructure needs.Look across... ...into unified, executable roadmaps.Lead, scale, and manage a global team of Ops PMs, SCPMs, and PMEs... ...fastest experience possible.As the Senior Director, Centralized Manufacturing Technical Operations...SeniorContract work$300k - $333k
...safety mitigations directly with modeling teams, ensuring all features... ...partnerships with Product Manager, Engineering, Legal, Safety, Privacy... ...cybersecurity within a high-scale consumer tech or enterprise... ...Processing (NLP), conversational AI principles, and model...Senior- ...and business teams in Sunnyvale. The role blends hands-on and strategic work to cultivate an innovative community, drive growth, and scale人才 development programs. You will partner across HR functions, coach leaders, and lead talent strategies, while leveraging data to...Senior
$152k - $230k
We are looking for a Senior Product Marketing Manager for the AI Infrastructure Solutions product marketing team. You will be alert to the... ...knowledge of GPU/CPU architectures, as well as training and inference performance metricsProven track record to lead and drive cross...SeniorFull timeWork experience placement$272k - $431.25k
...generation of interactive world-model systems. With this... ...fidelity, real-time inference performance in world... ...talk.What you'll be doing:Scale the engineering team... ...experience, CI/CD, and release management.Partner with research,... ...vacancy. NVIDIA uses AI tools in its recruiting...SeniorFull time$200k - $322k
...Infrastructure is seeking a Senior Technical Program Manager to lead the strategy and execution of... ...and partners of other business units to scale the EDA Infrastructure charter. They will... ...is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA...SeniorFull time$168k - $258.75k
...for a highly-motivated Technical Program Manager (TPM) to join our Applied Systems Engineering... ...for the next generation of NVIDIA AI supercomputing systems. This TPM will play... ...the lifecycle of the latest AI systems at scale, from datacenter design and requirements definition...SeniorFull timeRemote work$220k - $270k
Sr. Manager, Technical Marketing & Applications (DC-DC Power Modules) page is loaded## Sr. Manager... ...ideal for someone who thrives in a large-scale environment but brings a startup mindset... ...high-impact markets such as data center, AI infrastructure, and industrial systems....SeniorRemote workWorldwide$262k - $364k
...test, deploy, maintain, and enhance large-scale software solutions.Refine and improve the... ...industry ML infrastructure (e.g., model deployment, model evaluation, data processing... ...forward.With your technical expertise you will manage project priorities, deadlines, and...Senior$190.9k - $334.1k
...meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our... ...Intelligent Services is seeking a Senior Manager of Product Management for Search AI to lead... ...next wave of growth and innovation as we scale to $30B+ in revenue. Own outcomes, not...SeniorFull timeWork at officeImmediate startRemote workFlexible hours$190.9k - $334.1k
...work. Today, ServiceNow is the AI control tower for business... ...production with marquee customers and scaling fast, competing head-to-head... ...for a Senior Staff Product Manager to lead the conversational core... ...The emerging voice-to-voice model integration and deployment strategy...SeniorFull timeWork at officeImmediate startRemote workFlexible hours$190.9k - $334.1k
...work. Today, ServiceNow is the AI control tower for business... ...exceptional Senior Staff Product Manager to own and drive the strategic... ...knowledge of the AI market—emerging models, techniques, competitive moves... ...expertise building and scaling conversational AI, agentic, or...SeniorFull timeWork at officeImmediate startRemote workFlexible hours$120k - $160k
...news and information powered by advanced AI, recommendation systems, and adtech.Recognized... ...about solving meaningful challenges at scale.Together, we reached unicorn status in 20... ..., including onboarding, performance management, employee development, organizational planning...SeniorFull timeWork at officeLocal areaWork from homeMonday to Friday
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director/Sr. Manager, AI Inference Model Scaling. Be the first to apply!
- ai director Sunnyvale, CA
- director of aviation Sunnyvale, CA
- director medical information Sunnyvale, CA
- director talent management Sunnyvale, CA
- director of automation Sunnyvale, CA
- senior director epidemiology Sunnyvale, CA
- residence director Sunnyvale, CA
- director of billing Sunnyvale, CA
- director of practice management Sunnyvale, CA
- director Sunnyvale, CA



