Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Core - SDET Technical Lead, Release Integration Testing

Full-time

Cerebras Systems

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.About the RoleWe are looking for a hands-on SDET Technical Lead to establish and lead Release Integration Testing within Release & Feature Qualification for AI Inference Core.The Production Engine for Inference Core — turning integrated features into reliable production releases.You will define the quality strategy across the pre-release and release cycle, from feature and model integration through branch stability, release qualification, deployment, and post-release learning. You will work across AI frameworks, runtime, compiler, kernels, distributed systems, infrastructure, and hardware to make release risk visible and actionable.This is a technical-leadership role, not a coordination-only position. You will design test architecture, lead difficult debugging and release decisions, mentor engineers, and write software and automation alongside the team.Release Integration Testing (RIT) is the bridge between feature qualification and release qualification. Feature teams retain ownership of feature design, feature-level qualification, and feature regression. Release Integration Testing owns inference-core integration strategy, inference-path readiness approval, integrated cross-stack validation, and first-pass rollout triage.What Makes This Role DistinctDedicated Release Integration Testing ownership: Engage before feature qualification completes while keeping the boundary clear: feature teams own feature behavior and qualification; Release Integration Testing owns integration strategy, readiness approval, integrated validation, and first-pass rollout triage.Inference-path readiness gate: Require evidence across unit, simulation, benchmark, feature, and integration testing, with explicit coverage gaps before release entry.Cross-stack test strategy: Define risk-based E2E and regression coverage for features spanning components, organizations, software layers, infrastructure, and hardware.Branch and rollout leadership: Establish measurable health standards for master and release branches, and coordinate inference-impacting rollout across multiple product and release projects.Hands-on technical authority: Lead through code, test architecture, difficult debugging, quality metrics, and evidence-based release decisions.Team multiplier: Raise the technical bar, mentor engineers, and align feature, infrastructure, integration, qualification, and release teams.What You Will DoDefine the Release Integration Testing strategy, engagement criteria, ownership boundaries, entry and exit criteria, coverage expectations, and escalation thresholds for AI Inference Core.Engage early on high-risk inference changes; identify dependencies and interaction risks across runtime, host, device programming, memory, scheduling, model execution, infrastructure, and hardware.Own the inference-path readiness gate by reviewing unit, simulation, benchmark, feature-test, and integration evidence, documenting gaps, and approving integration readiness before release entry.Lead integrated inference E2E validation across features and the cloud-to-wafer stack; promote durable feature tests and add risk-based scenarios to release regression.Improve master and release-branch stability through actionable health metrics, failure classification, release-quality reporting, dashboards, qualification workflows, and release pipelines.Lead first-pass regression and rollout triage, coordinate owners through resolution, drive RCA, place missing coverage at the correct layer, and plan rollout across multiple product and release projects.Partner with and mentor SDETs, feature teams, Integration, Core Infra, release owners, and deployment teams; between active engagements, advance automation efficiency, diagnostics, probes, and roadmap test planning.Minimum Skills & QualificationsStrong software-engineering fundamentals and programming ability in Python Go, or a similar language.Demonstrated technical leadership in software quality, test infrastructure, systems validation, release engineering, or complex software integration.Experience designing automation and test architecture for distributed, systems-level, infrastructure, or AI software.Proven ability to break down ambiguous cross-stack failures, form hypotheses, gather evidence, and drive issues to resolution.Strong understanding of risk-based testing, release readiness, regression strategy, failure analysis, and quality metrics.Ability to influence and align multiple engineering teams without relying solely on organizational authority.Clear communication and sound judgment during high-pressure release situations, including the ability to explain technical risk to engineering and leadership audiences.Preferred SkillsExperience with software/hardware co-design, hardware accelerators, compilers, kernels, runtimes, or low-level systems.Experience with AI infrastructure, model deployment, LLMs, multimodal workloads, or large-scale compute clusters.Experience building test frameworks, distributed test systems, release pipelines, dashboards, or internal developer tooling.Experience with performance testing, profiling, observability, fault injection, reliability, or production failure analysis.Experience in a startup or similarly fast-moving, resource-constrained engineering environment.Track record of taking a quality or release capability from zero to one and scaling it across teams.Familiarity with containers, cluster orchestration, cloud infrastructure, CI/CD, or high-performance computing.What Success Looks LikeRelease readiness is based on explicit criteria and high-signal evidence rather than intuition.Fewer inference-path integration defects are first discovered in final release qualification or production.Cross-component risks are found earlier, debug cycles are shorter, and coverage ownership is explicit.Master and release-branch health is measurable, actionable, and steadily improving.Test automation and release infrastructure shorten feedback loops without sacrificing signal quality.Release metrics and reports drive clear decisions, ownership, and predictable feature rollout.Engineers across the organization are more effective because Release Integration Testing provides strong technical direction, tooling, and mentorship.LocationThis role requires in-office presence, at least three days per week. Fully remote work is not available.Office locations: Sunnyvale, CA or Toronto, ON.Why Join CerebrasPeople who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:Build a breakthrough AI platform beyond the constraints of the GPU.Publish and open source their cutting-edge AI research.Work on one of the fastest AI supercomputers in the world.Enjoy job stability with startup vitality.Our simple, non-corporate work culture that respects individual beliefs.Find out more about what it's like to work at Cerebras here! Apply today and become part of the forefront of groundbreaking advancements in AI!Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.LocationSunnyvale, CAEmployment TypeFull timeLocation TypeHybridDepartmentSoftware Engineering

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Inference Core - SDET Technical Lead, Release Integration Testing in Sunnyvale, CA vacancy
  •  ...the world's largest AI chip, 56 times larger...  ...to deliver industry-leading training and inference speeds; over 10...  ...looking for a hands-on SDET Technical Lead to establish and lead Release Integration Testing within Release & Feature...  ...for AI Inference Core.The Production Engine... 
    Suggested
    Work at office
    Remote work
    3 days per week

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  •  ...world's largest AI chip, 56...  ...deliver industry-leading training and inference speeds; over...  ...production release shipped to Cerebras...  ...Engineer in Test for the ML...  ...as a key technical leader in delivering...  ...feature integration quality and drive...  ...guide junior SDETs on testing... 
    Suggested
    Work at office
    Remote work
    Shift work
    3 days per week

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  •  ...world's largest AI chip, 56 times larger...  ...deliver industry-leading training and inference speeds; over 10...  ...Platform SDET to join the Inference...  ...build, and maintain test infrastructure and...  ...features and releases.Improve observability...  ..., or platform integration.Strong problem-solving... 
    Suggested
    Work at office

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...world's largest AI chip, 56 times...  ...industry-leading training and inference speeds; over 1...  ...GPU Inference SDET, you will be the...  ...the end-to-end release qualification and automated test ecosystem for...  ...be the primary technical anchor ensuring...  ...& CI/CD Integration: Integrate automated... 
    Suggested

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  •  ...the world's largest AI chip, 56 times...  ...deliver industry-leading training and inference speeds; over 10 times...  ...About the TeamThe Core Infrastructure team...  ...systems, test infrastructure, developer...  ...supporting build, test, integration, qualification, and release workflows.Build... 
    Suggested

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $119.8k - $234.7k

     ...SCHIE delivers the core infrastructure...  ...Thanks to its integrated design, this...  ...designing, building, testing and deploying...  .../firmware releases, and the fixes...  ...diagnostics tools, and AI agents to...  ...issues. Provides technical leadership to...  ...project schedules. Leads the team by... 
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Santa Clara, CA
    5 days ago
  • $186.9k - $267.7k

     ...platforms for Cisco's core Switching, Routing, and...  ..., developing and testing some of the most complex...  ...Your ImpactAs an ASIC Technical Lead, you will stand at the...  ...custom DFT logic & IP integration; familiarity with functional...  ...organizations in the AI era - and beyond. We’... 
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    San Jose, CA
    1 day ago
  • $165.2k - $223.6k

     ...that seamlessly integrates with popular ML...  ...ML inference and training performance...  ...'s possible in AI acceleration.As...  ...role will help lead the efforts in...  ...optimizations, testing and production...  ...deployment and releases through pipelines...  ...decisions with your technical input. You will... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    19 hours ago
  • $184k - $287.5k

     ...foundational to modern HPC and AI. At the center of this platform are CUDA Core Libraries that enable...  ...for GPU computing while integrating closely with native C/...  ...design, implementation, testing, profiling, benchmarking...  ..., packaging, release, and maintenance.Improve... 
    Full time

    Nvidia

    Santa Clara, CA
    19 hours ago
  • $184k - $287.5k

     ...foundational to modern HPC and AI. At the center of this platform are CUDA Core Libraries that provide...  ...design, implementation, testing, profiling, benchmarking, documentation, release, and maintenance.Improve...  ..., examples, build integration, tests, benchmarks, and continuous... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...to modern HPC and AI. At the center of...  ...platform are CUDA Core Libraries that enable...  ....Develop and integrate the native C/C++ components...  ..., implementation, testing, profiling,...  ...benchmarking, documentation, release, and long-term...  ...specifications, technical designs, and user... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $60 - $82 per hour

     ...Job Description SDET Android System...  ...execute comprehensive test strategies,...  ...including functional, integration, regression, and...  ...with a focus on core Android...  ...during monthly releases Co-ordinate with...  ...Developers and QA leads. Help in...  ...such as Gradle. AI/ML Testing: Experience... 
    Hourly pay

    Cypress HCM

    Mountain View, CA
    20 days ago
  • $165.2k - $223.6k

     ...right moment. As AI models outgrow...  ...of disaggregated inference: splitting LLM serving...  ...your scope and technical depth — on a team at the leading edge of AI/ML...  ...Labs, an integral part of AWS. Annapurna...  ...WLB as a core org tenet. The team...  ...build processes, testing, and operations... 
    Full time
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $193.3k - $261.5k

     ...right moment. As AI models outgrow any...  ...of disaggregated inference: splitting LLM serving...  ...is a role on the leading edge of AI/ML...  ...Annapurna Labs, an integral part of AWS. Annapurna...  ...respect WLB as a core org tenet. The team...  ...build processes, testing, and operations experience... 
    Full time
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $153.2k - $234.1k

     ...will be part of a core team that...  ..., and scalable releases of the Autonomous...  ...intelligent automation, AI-enabled...  ...effort, improving test and release...  ...validation pipelines, integrate simulation and...  ...’ll Be Doing  Lead the design and...  .... Communicate technical findings,... 
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    5 days ago
  • $240k - $320k

     ...product engineering of AI-based Autonomous...  ...efficiently trains, optimizes, tests, and releases various E2E AI-based...  ...execution of the technical roadmap and strategy...  ...functional tech leads (e.g. data engineering...  ...fast evaluation and integration of emerging E2E AI solutions... 
    Full time
    Work experience placement
    Local area
    Flexible hours

    Robert Bosch

    Sunnyvale, CA
    4 days ago
  • $180k - $300k

     ...potential of generative AI to power the...  ...you will both set the technical direction for test strategy and directly...  ...Strategy & Leadership• Lead and mentor junior QA...  ...regression test suite, integrated into CI/CD pipelines...  ...characteristics, and the inference stack — particularly... 

    d-Matrix

    Santa Clara, CA
    3 days ago
  • $202.5k - $274k

     ...you will set the technical direction for...  ...shape how we apply AI — both to...  ...investment in test automation, observability and release health- Work with...  ...design, solution integration and on-boarding...  ...platforms — Core ML, Vision, Natural...  ...on-device LLM inference- Practical... 
    Contract work
    Worldwide

    Intuit

    Mountain View, CA
    19 hours ago
  • $193.3k - $261.5k

     ...Development Engineer on the Inference Model Enablement team,...  ...libraries* Drive technical excellence in...  ...day in the lifeYou'll lead critical technical initiatives...  ...model enablement team releases its models in the vLLM...  ...management, build processes, testing, and operations... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $168k - $270.25k

     ...Senior Software Engineer in Test to join the Compute...  ...infrastructure, and AI assisted quality analysis...  ...NVIDIA Data Center GPUs!We lead test planning,...  ...platform bring up through key release stages. We collaborate...  ...readiness, CI/CD, API integrations, log analysis, root cause... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...builds the world's largest AI chip, 56 times larger than...  ...Cerebras to deliver industry-leading training and inference speeds; over 10 times...  ...software development engineer in Test, we are looking for a...  ...industry. With dozens of model releases and rapid growth, we’ve reached... 

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  • $184k - $287.5k

     ...Systems Software test (lead) Engineer to join...  ...role combines deep technical expertise from cluster...  ...training and inference platforms. You will...  ...validation plans for CSP integration milestones,...  ...test report for each release for the rack scale...  .... NVIDIA uses AI tools in its recruiting... 
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • $165.2k - $223.6k

     ...engineer to work on distributed AI/ML systems. This role involves...  ...joining is Annapurna Labs, an integral part of AWS and develops...  ...working hours, and respect WLB as a core org tenet. The team enjoys...  ...management, build processes, testing, and operations experience- Bachelor... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $193.3k - $261.5k

     ...to work on distributed AI/ML systems. This role involves...  ...is Annapurna Labs, an integral part of AWS and...  ..., and respect WLB as a core org tenet. The team enjoys...  ...- 5+ years of leading design or architecture...  ...management, build processes, testing, and operations experience... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $149.1k - $218.1k

     ...effectively. As part of the Security Core AI team, you will build and...  ...problems and serve as a technical resource within the team...  ...through implementation, testing, integration, deployment, and continuous...  ...documentation, debugging, release, and production analysis, with... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    11 hours ago
  • $79.1k - $166.1k

     ...an impact at the leading edge of cloud infrastructure...  ..., and OCI service integration layers.If yes,...  ...develop, debug, test, and improve RoT...  ..., validation, release, and operational support...  ...care. And with AI embedded across our...  ..., diagnose technical issues, and collaborate... 
    Temporary work
    Internship
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    11 hours ago
  • $100k

    Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance...  ...Interconnect / Signal Integrity Engineer to design and...  ...technologies for next-generation AI inference and training clusters....  ...cable specification and testing, as well as accelerated... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    3 days ago
  • $235.03k - $352.29k

     ...profound opportunity for AI to drive positive...  ...Price, and other leading investors.About the...  ...enthusiastic about integrating solutions...  ...WorkYou will:Own the core logic and software...  ...principled solutions, and testing and deployment.Lead and grow a technical team dedicated to this... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    11 hours ago
  • $152k - $204k

     ...Essential Cloud for AI™. Built for...  .... Trusted by leading AI labs, startups...  ...with deep technical expertise to accelerate...  ...-native inference platform and meet...  ...reliability improves release-over-release....  ...elevate coding/testing standards....  ...represented through our core values:  Be... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    28 days ago
  •  ...DescriptionAdvantest America, a leading Semiconductor Test and Measurement Company, is...  ...of a global initiative to integrate advances in electrical,...  ...mechanical, thermal, software, and AI/Machine Learning...  ...allowing for innovation and technical leadership. You will be involved... 
    For contractors

    Advantest

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Core - SDET Technical Lead, Release Integration Testing. Be the first to apply!