Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Orchestration Workload Engineer - ACE - AI Factory

Genentech

The PositionAs a Workload Orchestration Engineer within the Accelerated Compute Engineering (ACE) team, you will be recognised internally as an expert in workload orchestration, owning and advancing our scheduler tech stack across our High-Performance Computing (HPC) platforms. With the rapid expansion of our compute infrastructure, your broad expertise will drive the efficient scheduling, policy management, and resource optimization of our multi-node CPU and GPU environments.In this role, you will use your expertise to bridge traditional scientific computing with modern AI paradigms, while acting as a coach and mentor to help colleagues develop technical expertise. You will solve unique, unprecedented scheduling and infrastructure challenges that directly impact Roche’s compute architecture, ensuring our researchers, data scientists, and engineers can execute compute workloads reliably, efficiently, and successfully.Hosting and Infrastructure (HI) provides mission-critical on-premises infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise.The Value Streams - Accelerated Compute Engineering (ACE) Team acts as a center of excellence and delivery for High Performance Compute and AI Infrastructure across Roche. This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute—helping infrastructure consumers optimize for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI to achieve rapid time-to-value.The Opportunity:SLURM Architecture & Ecosystem LeadershipServe as the internal expert on the SLURM Workload Manager, architecting, scaling, and maintaining SLURM across heterogeneous HPC (and AI environments) to ensure high availability and dynamic resource distribution.Design and tune advanced SLURM configurations, including custom plugin integration, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and complex QoS/fair-share policies.Bridge HPC and cloud-native ecosystems by evaluating and implementing integrations between SLURM, Kubernetes, and orchestration platforms (e.g., SLURM Slinky or Run:ai) to streamline job submission workflows across architectures.Hybrid Workload & Kubernetes IntegrationIntegrate containerization standards across SLURM (using Singularity/Apptainer) while maintaining operational familiarity with Kubernetes container orchestration to support hybrid AI/HPC workloads.Solve unique, unprecedented multi-tenant bottlenecks, such as GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures.Technical Leadership, Mentorship & GovernanceLead large, global cross-functional initiatives across ACE, infrastructure, platform, scientific computing, and AI teams to establish workload orchestration standards, policies, and architectural patterns across Roche compute environments.Act as a technical mentor and coach for junior and mid-level engineers, driving skill development and continuous learning across the chapter.Partner with Observability Engineers to establish deep telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization, utilizing configuration-as-code to deploy policies uniformly.Who You Are:Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline.Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization.Demonstrated track record of leading complex technical initiatives and mentoring engineering peers.Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments.SLURM Architecture & Optimization: Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments, including scheduler/backfill tuning, partition and topology design, priority/fair-share/QoS policies, cgroups, HA architecture, plugin integration, and GRES/TRES modeling for GPUs and specialized resources.SLURM Operations, Accounting & Observability: Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions.Kubernetes & Container Knowledge: Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context.AI Infrastructure & Interconnects: Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL).Automation: Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines.Broad Platform Expertise: Apply broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management to solve orchestration problems spanning multiple technology domains.Domain Expertise & Problem Solving: Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads.Coaching & Collaboration: Strong leadership presence with a dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams.Cross-Organizational Coordination: Collaborative team player with demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders.Strategic Vision: Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows.#RDT2026Genentech is an equal opportunity employer. It is our policy and practice to employ, promote, and otherwise treat any and all employees and applicants on the basis of merit, qualifications, and competence. The company's policy prohibits unlawful discrimination, including but not limited to, discrimination on the basis of Protected Veteran status, individuals with disabilities status, and consistent with all federal, state, or local laws.If you have a disability and need an accommodation in relation to the online application process, please contact us by completing this form Accommodations for Applicants.Job SummaryJob number: 202609-123254Date posted : 2026-09-18Profession: Business Strategy & DeliveryEmployment type: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Orchestration Workload Engineer - ACE - AI Factory in South San Francisco, CA vacancy
  • $180k - $250k

     ...powering the next generation of AI products. We build the...  ...high-performance inference, orchestration, and observability come together...  ...are an experienced software engineer who thrives on building large...  ...platform: request routing, AI workload orchestration, scheduling, GPU... 
    Suggested
    Full time
    Currently hiring
    Remote work
    Relocation package

    Falò

    San Francisco, CA
    1 day ago
  • $140k - $210k

     ...Nimble Nimble is an AI robotics company...  ...everything from the inside of factories and warehouses to your...  ...As a Software Engineer on the Multi-Agent Systems...  ...reliable systems that orchestrate dozens or hundreds of...  ...break down orders and workloads into robot tasks while... 
    Suggested
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    18 days ago
  •  ...Overview \ We're hiring a Robotics Software Engineer to develop the real -time systems that power...  ...to make every tedious and dangerous warehouse/factory job virtual, safe, and semi -autonomous. \ \ With proven AI approaches and long distance teleoperation, you... 
    Suggested
    Full time
    Remote work
    Worldwide
    Relocation
    Long distance
    Flexible hours

    Avatar Robotics

    San Francisco, CA
    1 day ago
  • $150k - $237.5k

     ...batteries we already have. Senior Software Engineer, Energy Storage This position is on...  ..., simulation and operational control orchestration, and integration with energy markets. We...  ...improve it over time Apply AI tools to accelerate development while maintaining... 
    Suggested
    Full time

    Redwood Materials

    San Francisco, CA
    1 day ago
  •  ...that. You’ll help develop and deliver the factories, grids, transit systems, and public...  ...think big and act bold - project managers, engineers, technologists, and strategists who blend...  ...world experience with digital innovation and AI. Together, we’re transforming how capital... 
    Suggested
    Contract work
    Temporary work
    For contractors
    Work at office
    Local area
    Flexible hours

    Accenture Infrastructure and Capital Projects, LLC

    San Francisco, CA
    3 days ago
  •  ...that. You’ll help develop and deliver the factories, grids, transit systems, and public...  ...think big and act bold - project managers, engineers, technologists, and strategists who blend...  ...world experience with digital innovation and AI. Together, we’re transforming how capital... 
    Contract work
    Temporary work
    For contractors
    Work at office
    Local area
    Flexible hours

    Accenture Infrastructure and Capital Projects, LLC

    San Francisco, CA
    18 days ago
  • $127k - $184k

    Orchestrate a structured, end-to-end deployment plan across customer,...  ...initial and ongoing ramp of workloads, moving customers from agreement...  ...management, data analytics, AI, networking, migrations or security...  ...(GCP) Outcome Customer Engineer (OCE), you will drive initial... 
    Contract work

    Google

    San Francisco, CA
    4 days ago
  • $347k

     ...organization builds and evaluates the systems that power advanced AI workloads. We work closely with hardware, modeling, and architecture...  ...About the RoleWe are seeking a Workload Porting & Performance Engineer to evaluate new hardware platforms by porting benchmarks and real... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    4 days ago
  • $170k - $210k

     ...Through intelligent automation, we give factories newfound flexibility, scalability, and resilience...  ...you.  ABOUT THE ROLE: Software Engineers at Bright Machines are responsible for...  ...Bright Machines is a next-generation, AI-enabled manufacturer focused on data center... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Bright Machines

    San Francisco, CA
    1 day ago
  •  ...Description Job Description Robotics Software Engineer – Motion Planning, Vision & Deployment...  ...automation, computer vision, and AI-driven robotics. Their technology addresses...  ...excited about building systems that operate on factory floors, not just in simulation. This... 
    Permanent employment
    Full time

    MRINetwork Jobs

    San Francisco, CA
    27 days ago
  • $200k - $400k

    Senior Software Engineer - Agentic Systems We are partnered with a highly technical AI research company building advanced AI systems capable of operating across complex...  ...to design and build the core planning and orchestration layers that allow LLM‑based agents to reliably... 
    Full time
    Immediate start

    Strativ Group

    San Francisco, CA
    4 days ago
  •  ...can do the impossible at record breaking speeds.About You and The Role As an Industrial & Operations Engineer on the Manufacturing Engineering team, you will own factory performance and scaling.You will take a factory-level view of production—connecting demand, labor,... 
    Local area
    Shift work

    Zipline

    South San Francisco, CA
    6 hours ago
  • $300 per month

     ...only vertically integrated AI infrastructure company built...  ...the world's most ambitious AI workloads. When you join Crusoe, you...  ...Role:At Crusoe, our Production Engineering team ensures the reliability...  ...Kubernetes or container orchestration platformsStrong collaboration... 
    Temporary work

    Crusoe

    San Francisco, CA
    6 hours ago
  •  ...Mission Dedalus Labs is an AI Neolab building the compute substrate...  ..., networking, scheduling, orchestration, and low-level runtime...  .... We are looking for systems engineers who want to understand computers...  ...systems for secure multi-tenant workloads. Persistent state, storage,... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Dedalus Labs

    San Francisco, CA
    5 days ago
  •  ...the Brain Health and Neurodegeneration (BHN) model, bridging quantitative biology and AI-driven software. You will develop ODE-based mechanistic models and collaborate with engineers to translate discoveries into experiments and publications. Travel to the South San Francisco... 
    Work at office
    Remote work

    Fulcrum Neuroscience, Inc.

    South San Francisco, CA
    2 days ago
  •  ...designed for the unique demands of advanced AI workloads. The team is responsible for building...  ...We are seeking a Manufacturing Test Engineer to own and drive manufacturing test strategy...  ...across both engineering development and factory execution, and can bridge the gap... 
    Full time
    Contract work

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...edge research —including deep learning, generative AI, and reinforcement learning techniques— with large-scale engineering to bridge experimentation and production; you'll...  ...assessment, comprehensive guides, FAQs, and modules designed to help you ace the hiring process.... 
    Full time

    Roblox

    San Mateo, CA
    1 day ago
  • $300k

     ...mode startup building out their AI and cloud platform, powered...  ...or inference. As a Platform Engineer/Senior Site Reliability Engineer...  ..., ensuring seamless orchestration across environments managed by...  ...infrastructure for frontier AI workloads, automate systems at petascale... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $125k - $170k

     ...speeds.About You and The RoleWe are hiring a Senior Production Test Engineer to own electrical, functional, and automated test systems that...  ...In this role you will lead end-to-end test system delivery for factory floors that support high-throughput manufacturing and rapid... 
    Local area
    Shift work

    Zipline

    South San Francisco, CA
    3 days ago
  • $243.29k - $295.25k

     ...the Infrastructure Foundation Hardware Engineering team, you will help develop and validate...  ...and performance under production-scale workloads.Firmware & Fleet Enablement: Support BIOS...  ...Architectures: Familiarity with GPU platforms, AI accelerators, or high-performance... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    6 hours ago
  • $243.29k - $295.25k

     ...the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling...  ...be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible...  ...Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    6 hours ago
  •  ...Combinator and top-tier Silicon Valley investors, we’re turning physical AI into reality — helping industries facing critical labor...  ...The Role We’re looking for a talented Robotics Software Engineer to design and build the core systems that power Verne’s robots in... 
    Full time
    Work experience placement
    Immediate start

    Verne Robotics

    South San Francisco, CA
    1 day ago
  • $350k

     ...knowledge and tools to make AI work for their unique needs and...  ...goals.  We are scientists, engineers, and builders who’ve created...  ...Kubernetes clusters with GPU workloads, or building infrastructure to...  ...scale clusters and container orchestration systems (e.g. Kubernetes or... 
    Full time
    Local area
    Immediate start
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours

    Thinking Machines Lab

    San Francisco, CA
    1 day ago
  •  ...modern applications by helping them modernize legacy workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB...  ...scalable yet observable system for customers and engineers. The Atlas Search product is quickly gaining traction... 
    Full time
    Work at office
    Local area
    Worldwide

    Mongodb

    San Francisco, CA
    1 day ago
  • $200k

     ...Senior Software Engineer — Lakehouse Systems Location: Mountain View, CA — On-site About Granica Granica builds AI infrastructure for enterprises operating massive data environments...  ..., and compute layers Develop workload-aware table optimization systems that... 
    Full time
    Work at office

    Granica

    San Francisco, CA
    1 day ago
  •  ...seamless experience for even the largest workloads Work with a collaborative team that prioritizes...  ...love and that we are proud of as engineers Have the opportunity to lead projects...  .... We have redefined the database for the AI era, enabling innovators to create, transform... 
    Full time
    Work at office
    Local area
    Worldwide

    Mongodb

    San Francisco, CA
    1 day ago
  • Phylo is an applied research lab accelerating discovery through AI agents. We seek an engineer to build and evaluate systems that make agents capable, reliable, and better. You will own production harnesses, incorporate cutting-edge research, and drive experiments to measure... 

    Phylo

    South San Francisco, CA
    2 days ago
  • $160k - $200k

    Vast.ai Engineering RoleVast.ai runs one of the largest GPU marketplaces in the world. Over 20,000 GPUs, from RTX 4090s up to B300s, rented...  ...new model actually runs, keep reading.What You'll DoRun real workloads on Vast every week. Deploy, fine-tune, and serve open source... 
    Weekend work

    Vast

    San Francisco, CA
    6 days ago
  • Modal Ai InfrastructureAI needs a new infrastructure layer. We're building it at Modal.Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases...  ...medalists, and experienced engineering and product leaders with decades... 
    Work at office

    Modal

    San Francisco, CA
    6 days ago
  • Teliolabs Communication Private Limited in California seeks an Oracle OSM Developer to design, develop, and maintain OSM cartridges and orchestration flows that automate order fulfillment processes across telecom and digital service ecosystems. You will implement cartridges for... 

    Teliolabs Communication Private Limited

    South San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Orchestration Workload Engineer - ACE - AI Factory. Be the first to apply!