Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI
Magic AI Corp.
Magic AI, Inc. is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and inference. You will implement Terraform-driven IaC across cloud and hybrid environments, manage Kubernetes clusters, and ensure reproducibility and reliability of thousands of GPUs. This role offers visa sponsorship and relocation support to San Francisco. #J-18808-Ljbffr Magic AI Corp.
- Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building...Suggested
- Salesforce, Inc. is seeking a Senior Software Engineer to join the team responsible for the voice infrastructure at scale. The role focuses on deploying, maintaining, and monitoring voice services across multiple regions, ensuring high availability and quality for customers...Suggested
- A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design... ...training systems and optimize GPU utilization while collaborating with...Suggested
- ...innovative company is seeking a talented software engineer to join their dynamic Inference team.... ...and implementing infrastructure for large-scale multimodal models, focusing on high-performance... ...product teams to push the boundaries of AI technology, ensuring reliable production...Suggested
- Crusoe is seeking a Staff TPM to own deployment programs for new sites and capacity expansions in our GPU‑powered AI cloud. You will define targets, gating criteria, and DRI matrices... ...junior TPMs, shaping how we deploy at scale across hyperscaler customers and internal...Suggested
$350k
Mirendil is looking for engineers to build infrastructure for frontier reasoning models at their San Francisco location. This role focuses on large-scale reinforcement learning (RL) model training and requires a solid understanding of engineering principles. The ideal candidate...- Applied Compute is seeking a research scientist to design, implement, and optimize the large-scale training infrastructure powering our reinforcement learning stack in a San Francisco office. You’ll work with researchers to ensure the RL system is fast, reliable, and capable...Work at office
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...
$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed for... ...looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... ...-Code tooling such as Terraform and Ansible Proven...Full timeRemote work$315k
A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams...- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...
- Rox Data Corp is hiring a Foundations Engineer (Deep Infra) to design and operate the systems powering Rox’s agent runtime. You’ll work at... ...will build and maintain streaming, real-time analytics, and large-scale data infrastructure that powers production deployments, ensuring...
$179k - $218k
...only vertically integrated AI infrastructure company... ...urgency, who believe in the scale of our ambition and thrive... ...bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to... ...ROCm). Experience using large datasets or basic ML...Temporary work$300k
...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to... ...team is building low-latency AI systems where milliseconds... ...inference execution — not general infra, not model research. You’ll... ...runtime layers, profiling large-scale speech and multimodal...RelocationVisa sponsorshipFree visa$250k
A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include...- Imprezia is building the ads infra for the first 1 billion AI users. Besides cars and perhaps cellphones, the only products to... ...advertising across the AI ecosystem. We are scaling rapidly and are hiring a generalist engineer who can take ownership of important systems and...
- Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress. Collaborate with researchers to...
- Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates have...Full time
- A leading technology firm in San Francisco is seeking a Software Architect to design high-performance infrastructure and foster engineering best practices. The ideal candidate will have experience with AWS Serverless products, including Lambda and Step Functions, and be...Full time
$232k - $319k
...Identity, from AI to HumanIdentity... ...us continue to scale the service... ...Edge networking, K8s platform, Observability... ...doing Lead the Infra platform and... ...and product engineering by developing... ...experience running large-scale... ..., Nginx), IaC (Terraform), Splunk, Grafana...Permanent employmentLocal areaWorldwideFlexible hours- ...for the world's most dynamic AI companies, like Cursor, Notion... ...us and help build the platform engineers turn to to ship AI products.... ...LLM and multi-modal workloads scale, the network is the computer.... ...foundational engineers to lead our GPU Networking efforts, making...Full timeFlexible hours
- Doordashusa is seeking an Infrastructure Software Engineer to enhance simulation and machine learning infrastructures. You will work... ...dynamic environment and collaborate with various teams to support large-scale operations. Ideal candidates will hold degrees in relevant...
$300 per month
...vertically integrated AI infrastructure... ...who believe in the scale of our ambition and... ...platform — and Production Engineering sits at the heart... ...of Crusoe’s GPU cloud that powers next... ...problems, improving large-scale distributed systems... ...tools such as Terraform or AnsibleScripting...Temporary work$152.5k - $205k
...Circle to power trusted, internet-scale financial innovation. Learn... ...:As a Senior Site Reliability Engineer on Circle’s platform team, you... ...critical digital-assets, AI, and application workloads. You... ...Kubernetes platforms, and using Terraform to make infrastructure repeatable...Flexible hours$170k - $250k
...infrastructure company operating in the AI space, backed by a leading... ...revenue within six months and is scaling rapidly with a small, high-performing... ...This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role...Full timeVisa sponsorshipFlexible hours- ...Cloud, is a leader in AI cloud... .... One person, one GPU.If you'd like to build... ...for building and scaling the internal systems... ...company—Finance, GTM, Engineering, and People—to implement... ..., and methods for large-scale distributed... ...(Chef, Ansible, Terraform, GitHub Actions,...Work at officeLocal areaWork from homeFlexible hours
$250k - $300k
...vertically integrated AI infrastructure... ...who believe in the scale of our ambition... ...looking for a Senior Staff Network Architect... ...Architect and network engineering leadership to... ...Ansible, Nornir, or Terraform, driven by the... ...-team initiatives.Large-Scale Data Center...Temporary workWork at office$250k - $300k
...vertically integrated AI infrastructure... ...believe in the scale of our ambition and... ...Role:As a Senior Staff/Principal Deployment Automation Engineer for the Compute... ...testing automation of large-scale, multi-node GPU clusters. You... ..., Docker, Terraform, and Postgres.CI/...Temporary work$200k
...high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing,... ...deep experience with Ansible, Terraform, or SaltStack. High-Scale Networking: A strong foundation... ...Solver" DNA: A background in large-scale DataCenter environments...Full time$204k - $306k
...Every Identity, from AI to HumanIdentity is... ...Site Reliability Engineering GroupOkta authenticates... ...us continue to scale the service with great... ...Edge networking, K8s platform, CI/CD,... ...scaleExperience running large-scale... ...(Kubernetes), IaC (Terraform), and CI/CD pipelinesStrong...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI. Be the first to apply!
- software engineer staff San Francisco, CA
- assistant engineer San Francisco, CA
- engineering aide San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineering manager San Francisco, CA
- senior staff systems engineer San Francisco, CA
- technology administrator San Francisco, CA
- project engineer assistant project manager San Francisco, CA



