Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI

Magic AI Corp.

Magic AI, Inc. is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and inference. You will implement Terraform-driven IaC across cloud and hybrid environments, manage Kubernetes clusters, and ensure reproducibility and reliability of thousands of GPUs. This role offers visa sponsorship and relocation support to San Francisco. #J-18808-Ljbffr Magic AI Corp.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI in San Francisco, CA vacancy
  • Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building... 
    Suggested

    Linuxcareers

    San Francisco, CA
    4 days ago
  • Salesforce, Inc. is seeking a Senior Software Engineer to join the team responsible for the voice infrastructure at scale. The role focuses on deploying, maintaining, and monitoring voice services across multiple regions, ensuring high availability and quality for customers... 
    Suggested

    Salesforce

    San Francisco, CA
    10 hours ago
  • A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design...  ...training systems and optimize GPU utilization while collaborating with... 
    Suggested

    Baseten

    San Francisco, CA
    3 days ago
  •  ...innovative company is seeking a talented software engineer to join their dynamic Inference team....  ...and implementing infrastructure for large-scale multimodal models, focusing on high-performance...  ...product teams to push the boundaries of AI technology, ensuring reliable production... 
    Suggested

    Jobleads-US

    San Francisco, CA
    3 days ago
  • Crusoe is seeking a Staff TPM to own deployment programs for new sites and capacity expansions in our GPU‑powered AI cloud. You will define targets, gating criteria, and DRI matrices...  ...junior TPMs, shaping how we deploy at scale across hyperscaler customers and internal... 
    Suggested

    Crusoe

    San Francisco, CA
    2 days ago
  • $350k

    Mirendil is looking for engineers to build infrastructure for frontier reasoning models at their San Francisco location. This role focuses on large-scale reinforcement learning (RL) model training and requires a solid understanding of engineering principles. The ideal candidate... 

    Mirendil

    San Francisco, CA
    3 days ago
  • Applied Compute is seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning stack in a San Francisco office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable... 
    Work at office

    Applied Compute

    San Francisco, CA
    4 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    1 day ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...a next-generation GPU platform designed for...  ...looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments...  ...-Code tooling such as Terraform and Ansible Proven... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 

    Anthropic

    San Francisco, CA
    2 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 

    Baseten

    San Francisco, CA
    3 days ago
  • Rox Data Corp is hiring a Foundations Engineer (Deep Infra) to design and operate the systems powering Rox’s agent runtime. You’ll work at...  ...will build and maintain streaming, real-time analytics, and large-scale data infrastructure that powers production deployments, ensuring... 

    Rox Data Corp

    San Francisco, CA
    10 hours ago
  • $179k - $218k

     ...only vertically integrated AI infrastructure company...  ...urgency, who believe in the scale of our ambition and thrive...  ...bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to...  ...ROCm). Experience using large datasets or basic ML... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $300k

     ...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to...  ...team is building low-latency AI systems where milliseconds...  ...inference execution — not general infra, not model research. You’ll...  ...runtime layers, profiling large-scale speech and multimodal... 
    Relocation
    Visa sponsorship
    Free visa

    techire ai

    San Francisco, CA
    3 days ago
  • $250k

    A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include... 

    Acceler8 Talent

    San Francisco, CA
    1 day ago
  • Imprezia is building the ads infra for the first 1 billion AI users. Besides cars and perhaps cellphones, the only products to...  ...advertising across the AI ecosystem. We are scaling rapidly and are hiring a generalist engineer who can take ownership of important systems and... 

    Socket.dev

    San Francisco, CA
    20 hours ago
  • Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress. Collaborate with researchers to... 

    Kindredventures

    San Francisco, CA
    3 days ago
  • Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates have... 
    Full time

    Vast.ai

    San Francisco, CA
    1 day ago
  • A leading technology firm in San Francisco is seeking a Software Architect to design high-performance infrastructure and foster engineering best practices. The ideal candidate will have experience with AWS Serverless products, including Lambda and Step Functions, and be... 
    Full time

    SafetyKit

    San Francisco, CA
    4 days ago
  • $232k - $319k

     ...Identity, from AI to HumanIdentity...  ...us continue to scale the service...  ...Edge networking, K8s platform, Observability...  ...doing Lead the Infra platform and...  ...and product engineering by developing...  ...experience running large-scale...  ..., Nginx), IaC (Terraform), Splunk, Grafana... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  •  ...for the world's most dynamic AI companies, like Cursor, Notion...  ...us and help build the platform engineers turn to to ship AI products....  ...LLM and multi-modal workloads scale, the network is the computer....  ...foundational engineers to lead our GPU Networking efforts, making... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    20 hours ago
  • Doordashusa is seeking an Infrastructure Software Engineer to enhance simulation and machine learning infrastructures. You will work...  ...dynamic environment and collaborate with various teams to support large-scale operations. Ideal candidates will hold degrees in relevant... 

    Doordashusa

    San Francisco, CA
    20 hours ago
  • $300 per month

     ...vertically integrated AI infrastructure...  ...who believe in the scale of our ambition and...  ...platform — and Production Engineering sits at the heart...  ...of Crusoe’s GPU cloud that powers next...  ...problems, improving large-scale distributed systems...  ...tools such as Terraform or AnsibleScripting... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

     ...Circle to power trusted, internet-scale financial innovation. Learn...  ...:As a Senior Site Reliability Engineer on Circle’s platform team, you...  ...critical digital-assets, AI, and application workloads. You...  ...Kubernetes platforms, and using Terraform to make infrastructure repeatable... 
    Flexible hours

    Circle

    San Francisco, CA
    20 hours ago
  • $170k - $250k

     ...infrastructure company operating in the AI space, backed by a leading...  ...revenue within six months and is scaling rapidly with a small, high-performing...  ...This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role... 
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    1 day ago
  •  ...Cloud, is a leader in AI cloud...  .... One person, one GPU.If you'd like to build...  ...for building and scaling the internal systems...  ...company—Finance, GTM, Engineering, and People—to implement...  ..., and methods for large-scale distributed...  ...(Chef, Ansible, Terraform, GitHub Actions,... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    2 days ago
  • $250k - $300k

     ...vertically integrated AI infrastructure...  ...who believe in the scale of our ambition...  ...looking for a Senior Staff Network Architect...  ...Architect and network engineering leadership to...  ...Ansible, Nornir, or Terraform, driven by the...  ...-team initiatives.Large-Scale Data Center... 
    Temporary work
    Work at office

    Crusoe

    San Francisco, CA
    20 hours ago
  • $250k - $300k

     ...vertically integrated AI infrastructure...  ...believe in the scale of our ambition and...  ...Role:As a Senior Staff/Principal Deployment Automation Engineer for the Compute...  ...testing automation of large-scale, multi-node GPU clusters. You...  ..., Docker, Terraform, and Postgres.CI/... 
    Temporary work

    Crusoe

    San Francisco, CA
    15 hours ago
  • $200k

     ...high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing,...  ...deep experience with Ansible, Terraform, or SaltStack. High-Scale Networking: A strong foundation...  ...Solver" DNA: A background in large-scale DataCenter environments... 
    Full time
    San Francisco, CA
    more than 2 months ago
  • $204k - $306k

     ...Every Identity, from AI to HumanIdentity is...  ...Site Reliability Engineering GroupOkta authenticates...  ...us continue to scale the service with great...  ...Edge networking, K8s platform, CI/CD,...  ...scaleExperience running large-scale...  ...(Kubernetes), IaC (Terraform), and CI/CD pipelinesStrong... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    20 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI. Be the first to apply!