Director/Sr. Manager, AI Inference Model Scaling
Cerebras
Cerebras Systems builds the world’s largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
Sunnyvale, CA or Toronto, Canada (Hybrid)
About the Team
The Inference Model Scaling team enables state-of-the-art foundation models and generative AI workloads to run efficiently on Cerebras' Wafer-Scale Engine (WSE). We build the compiler frontend, model transformation pipeline, graph optimization infrastructure, high-performance kernel enablement, and runtime integration that together make next-generation AI models execute with industry-leading performance.
The team works at the intersection of machine learning frameworks, compiler technologies, distributed systems, hardware architecture, and model optimization. We collaborate closely with hardware architects, runtime engineers, cloud platform teams, AI researchers, and strategic customers to rapidly bring new model architectures into production.
About the Role
We're looking for an experienced engineering leader to build and scale our Inference Model Scaling organization.
You will define the technical vision, organizational strategy, and execution roadmap for a globally distributed engineering team responsible for enabling the latest foundation models on Cerebras hardware. You will lead the engineering organization responsible for ML model compilation and optimization as well as development of high-performance kernels.
This role combines deep technical leadership with organizational excellence. You will partner across compiler, runtime, cloud infrastructure, hardware architecture, product management, and AI research teams while helping shape the future of AI inference at Cerebras.
Responsibilities
Technical Leadership
- Define the technical roadmap and strategy for the team.
- Establish technical direction across multiple teams and engineering leaders.
- Lead design reviews and establish engineering standards.
- Drive support for emerging LLM architectures and inference workloads.
Team Leadership
- Hire, mentor, and grow a high-performing engineering team.
- Develop future technical leaders and managers.
- Drive organizational planning, headcount strategy, and investment priorities.
- Foster a strong engineering culture focused on execution, quality, and innovation.
- Scale engineering processes while maintaining execution velocity.
Cross-Functional Collaboration
- Partner with Cloud Platform, ML, and Hardware teams in planning and delivering for end-to-end service enablement in Cloud and On-Premise settings
- Work with Product Management to prioritize model enablement and customer needs.
- Collaborate closely with customers and solution architects on new model bring-up.
- Influence future hardware/software co-design through ML model enablement and optimization insights.
Delivery & Execution
- Own planning, prioritization, and execution across multiple concurrent initiatives.
- Balance rapid model support with long-term ML Compiler architecture.
- Drive predictable delivery for strategic customer commitments.
Required Qualifications
- BS, MS, or PhD in Computer Science, Computer Engineering or related field.
- 12+ years building compiler, ML systems, or infrastructure software.
- 5+ years leading engineering teams.
- Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).
- Strong understanding of graph compilation and optimization.
- Experience with Python and C++.
- Experience delivering production-quality software.
- Strong communication and cross-functional leadership skills.
Preferred Qualifications
- Experience building compiler frontends for AI accelerators.
- Experience supporting PyTorch, JAX, TensorFlow, or ONNX.
- Experience with LLM inference or training systems.
- Familiarity with distributed compilation.
- Experience working with hardware architects.
- Experience leading teams through rapid growth.
Why Join Cerebras
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
Find out more about what it's like to work at Cerebras here!
Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.
#J-18808-Ljbffr- ...Cerebras Systems is seeking an engineering leader to build and scale the Inference Model Scaling organization. You will define the technical... ...Work across compiler, runtime, cloud, hardware, product management, and AI research to shape the future of AI inference at Cerebras...Suggested
$193.3k - $261.5k
..., Amazon's custom cloud-scale machine learning accelerators... ...to optimize the latest models to run really fast on... ...Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you... ...engineers, and product managers to architect and deliver...SeniorInternshipLocal areaFlexible hours$148k - $235.75k
...Senior Technical Product Marketing Manager. This role will be located in... ...business and pivotal in our inference marketing. You will be focused... ..., networking, CUDA libraries, model architectures and deployment... ...showcase our leadership position in AI inference.Want to join a fun,...SeniorFull time$193.3k - $261.5k
...enabling unparalleled ML inference and training... ...running a wide range of models and supporting novel architecture... ...of what's possible in AI acceleration.As part of... ...peak performance at scale for customers and developers... ...engineers, and product managers to deliver state-of-the...SeniorWork experience placementInternshipLocal areaFlexible hours- ...builds the world's largest AI chip, 56 times larger... ...-leading training and inference speeds; over 10 times... ...works with the leading model labs, global enterprises... ...deploy 750 megawatts of scale, transforming key... ...visibility and aggressively manage schedule compression.Cross...SeniorContract workRemote work
$190.9k - $334.1k
...Today, ServiceNow is the AI control tower for business... ...transforming IT Service Management with AI by delivering on... ...Motion. We're looking for a Sr Staff Product Manager to define and scale this AI-Native Product-... ...Product Qualified Lead (PQL) models and product-to-sales...SeniorWork at officeImmediate startRemote workFlexible hours$191.4k - $252.72k
...building and operating the world's best data and AI infrastructure platform, enabling our... ...Senior Staff Technical Program Manager (TPM) for Reliability to lead the strategy... ...engineering teams at Databricks. As Databricks scales to support thousands of customers and the...SeniorLocal areaWorldwide$166k - $271k
...Engineering Technical Program Management Group is looking for a Senior... ...org builds and operates massive-scale systems: physical data centers... ..., observability, and Data and AI/ML platforms.Engineering Team... ...fundamentals, including consistency models, durability, fault tolerance,...SeniorFor contractorsWork at officeFlexible hours$184k - $287.5k
...aren't just powering the AI revolution—we're... ...accelerating it. The TensorRT inference platform is the... ...cutting-edge deep learning models on every NVIDIA GPU. With... ...and driven Engineering Manager to take the lead in developing... ...ability to lead and scale high-performing...Full time$199k - $269.5k
...delivering awesome enterprise-level outcomes at a scale that helps transform the lives of our... ...PMO as a Senior Staff Technical Program Manager. The Fintech PMO exists to accelerate the... ...portfolios end-to-end. Leverage and design AI capabilities within daily program...SeniorShift work$202k - $273.5k
...the Tech Strategic Programs team as a Sr Staff Technical Program Manager (TPM) focused on driving planning for... .... In this role, you will drive an AI-native approach that leverages our newly... ...AI-native planning mechanisms at scale Design operating mechanisms that drive...Senior$148k - $235.75k
...looking for a Senior Technical Marketing Manager to join our GeForce team — a high-impact... ...inform product decisionsHands-on use of AI tools — coding assistants, agentic... ...of modern AI systems: neural networks, model inference, transformer architectures, and GPU-accelerated...SeniorFull time- ...builds the world's largest AI chip, 56 times larger... ...-leading training and inference speeds; over 10 times... ...works with the leading model labs, global enterprises... ...deploy 750 megawatts of scale, transforming key... ...Senior Product Marketing Manager, you'll own realtime product...SeniorShift workNight shift
$332k
...into the unlimited potential of AI to define the next era of... ...hardware and software. "Embed and Scale" is our key GTM strategy and... ...multi-functionally with Product Management, Product Marketing, Developer... ...Define and Implement a leadership Inference go-to-market strategy!...SeniorFull timeWorldwide$168k - $258.75k
...real-world product outcomes at scale. This is not a coordination role... ...The exceptional hire also uses AI deliberately, not as a... ...tape-out and silicon correlation.Manage schedule headroom and surface resource... ...— and evolve the operating model so each successor program runs...SeniorFull time$192k - $264k
...the exciting technologies that literally connect our world - like AI and IoT. If you want to push the boundaries of materials science... ...approval guidelines and leadership to reporting engineering managers in the area of long-term program, strategy, and process design.Defines...SeniorFull time$262k - $364k
...evaluation metrics, or mathematical models.Use custom data infrastructure... ..., machine learning, and AI algorithms to develop models and... ...PhD degree.Experience in team management.Google Ads is helping power... ...Google to engage with customers at scale.Google Ads is helping power...SeniorWork experience placement$184k - $253k
...Sonatus, we’re driving the transformation to AI-enabled software-defined vehicles.... ...agility of a fast-growing company with the scale and impact of an established partner. Backed... ...for a highly motivated Technical Program Manager to define, manage, and support Sonatus Engineering...SeniorWork at officeWorldwideFlexible hoursShift work$175k - $236.8k
...ground floor of new projects to bring Agentic AI to more developers worldwide. You will be... ...for customers.As a Senior Product Manager Technical ES, you will be part of the larger... ...) experience- Experience delivering large-scale SaaS, PaaS or LaaS products where you are...SeniorLocal areaWorldwideFlexible hours$190.9k - $334.1k
...work. Today, ServiceNow is the AI control tower for business... ...innovations are conceived, built, and scaled — and where the programs you... ...Staff Technical Program Manager to join the APEX Strategic Program... ...merely adopting AI tools but modeling the ways of working we expect...SeniorWork at officeImmediate startRemote workFlexible hours$168k - $258.75k
At NVIDIA, we are at the forefront of AI and accelerated computing, redefining the world... ...society. As part of the Customer Program Management Team, you will lead strategic technical... ..., delivering unmatched data center-scale performance.What you'll be doing:Leading...SeniorFull time$212.7k - $287.7k
...hardware.As an SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize... .... You should be capable of managing demanding, fast-changing... ...design patterns, reliability and scaling) of new and existing systems experience...Local areaFlexible hours$272k - $431.25k
...generation of interactive world-model systems. With this... ...fidelity, real-time inference performance in world... ...talk.What you'll be doing:Scale the engineering team... ...experience, CI/CD, and release management.Partner with research,... ...vacancy. NVIDIA uses AI tools in its recruiting...SeniorFull time$152k - $230k
We are looking for a Senior Product Marketing Manager for the AI Infrastructure Solutions product marketing team. You will be alert to the... ...knowledge of GPU/CPU architectures, as well as training and inference performance metricsProven track record to lead and drive cross...SeniorFull timeWork experience placement$205.5k - $278k
...Strategic Programs as a Senior Staff Product Manager for AI Applications & Tools. Recognized by the... ...teams from experimentation through scaled adoption. This role requires an AI-first... ...saved.The team establishes a repeatable model for turning AI innovation into trusted,...SeniorWorldwide$224k - $356.5k
...Technical sales leader to drive our Data Center AI Factory business. We are searching for a... ...making revenue generating opportunity management, promotions, building sales strategies... ...techniques & software and the ability to scale up technical knowledge to serve the needs...SeniorFull timeFlexible hours$168k - $258.75k
...for a highly-motivated Technical Program Manager (TPM) to join our Applied Systems Engineering... ...for the next generation of NVIDIA AI supercomputing systems. This TPM will play... ...the lifecycle of the latest AI systems at scale, from datacenter design and requirements definition...SeniorFull timeRemote work- ...seeking a Senior Staff Technical Program Manager for the Fintech PMO to accelerate technology... ...partner with cross-functional teams to scale delivery. The role requires deep expertise... ...software development life cycles, architecture, and AI capabilities. #J-18808-Ljbffr IntuitSenior
$200k - $322k
NVIDIA is seeking a Senior Technical Program Manager to lead Trust Services programs for DGX Cloud. DGX Cloud powers large-scale AI infrastructure across NVIDIA, cloud service... ...firmware security, or hardware/software trust models.Working knowledge of major cloud platforms....SeniorFull time- ...builds the world's largest AI chip, 56 times larger... ...GPUs. Our novel wafer‑scale architecture provides... ...industry‑leading training and inference speeds and empowers... ...without the hassle of managing hundreds of GPUs or... ...customers include top model labs, global enterprises...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director/Sr. Manager, AI Inference Model Scaling. Be the first to apply!
- director of corporate relations Sunnyvale, CA
- director of culinary Sunnyvale, CA
- director of equity and inclusion Sunnyvale, CA
- residence director Sunnyvale, CA
- director of grants Sunnyvale, CA
- rehabilitation director Sunnyvale, CA
- dance director Sunnyvale, CA
- director revenue management Sunnyvale, CA
- director talent acquisition Sunnyvale, CA
- director biology Sunnyvale, CA


