Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI
Magic AI, Inc
Magic AI, Inc. is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and inference. You will implement Terraform-driven IaC across cloud and hybrid environments, manage Kubernetes clusters, and ensure reproducibility and reliability of thousands of GPUs. This role offers visa sponsorship and relocation support to San Francisco. #J-18808-Ljbffr Magic AI, Inc
- Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building...Suggested
- Salesforce, Inc. is seeking a Senior Software Engineer to join the team responsible for the voice infrastructure at scale. The role focuses on deploying, maintaining, and monitoring voice services across multiple regions, ensuring high availability and quality for customers...Suggested
- A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design... ...training systems and optimize GPU utilization while collaborating with...Suggested
- ...innovative company is seeking a talented software engineer to join their dynamic Inference team.... ...and implementing infrastructure for large-scale multimodal models, focusing on high-performance... ...product teams to push the boundaries of AI technology, ensuring reliable production...Suggested
- Crusoe is seeking a Staff TPM to own deployment programs for new sites and capacity expansions in our GPU‑powered AI cloud. You will define targets, gating criteria, and DRI matrices... ...junior TPMs, shaping how we deploy at scale across hyperscaler customers and internal...Suggested
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...
$315k
A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams...- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...
$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed for... ...looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... ...-Code tooling such as Terraform and Ansible Proven...Full timeRemote work- Rox Data Corp is hiring a Foundations Engineer (Deep Infra) to design and operate the systems powering Rox’s agent runtime. You’ll work at... ...will build and maintain streaming, real-time analytics, and large-scale data infrastructure that powers production deployments, ensuring...
$179k - $218k
...only vertically integrated AI infrastructure company... ...urgency, who believe in the scale of our ambition and thrive... ...bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to... ...ROCm). Experience using large datasets or basic ML...Temporary work- Scale AI is hiring for a senior software engineer to design, build, and scale full‑stack systems powering our GenAI data engine. You will work across front-end, back-end, and infrastructure, using React, TypeScript, Node.js, Python and data stores like MongoDB, Elasticsearch...
- Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress. Collaborate with researchers to...
- Imprezia is building the ads infra for the first 1 billion AI users. Besides cars and perhaps cellphones, the only products... ...advertising across the AI ecosystem. We are scaling rapidly and are hiring a generalist engineer who can take ownership of important systems and...
$250k
A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include...- A leading technology firm in San Francisco is seeking a Software Architect to design high-performance infrastructure and foster engineering best practices. The ideal candidate will have experience with AWS Serverless products, including Lambda and Step Functions, and be...Full time
$232k - $319k
...Identity, from AI to HumanIdentity... ...us continue to scale the service... ...Edge networking, K8s platform, Observability... ...doing Lead the Infra platform and... ...and product engineering by developing... ...experience running large-scale... ..., Nginx), IaC (Terraform), Splunk, Grafana...Permanent employmentLocal areaWorldwideFlexible hours$190k - $250k
...Sciforium is an AI infrastructure company developing next-... ...with hands-on support from AMD engineers the team is scaling rapidly to build the full... ...are seeking a highly skilled GPU Kernel Engineer who is passionate... ...that power next-generation large-scale AI systems. You will...Full timeFlexible hours- Doordashusa is seeking an Infrastructure Software Engineer to enhance simulation and machine learning infrastructures. You will work... ...dynamic environment and collaborate with various teams to support large-scale operations. Ideal candidates will hold degrees in relevant...
$152.5k - $205k
...Circle to power trusted, internet-scale financial innovation. Learn... ...:As a Senior Site Reliability Engineer on Circle’s platform team, you... ...critical digital-assets, AI, and application workloads. You... ...Kubernetes platforms, and using Terraform to make infrastructure repeatable...Flexible hours$300 per month
...vertically integrated AI infrastructure... ...who believe in the scale of our ambition and... ...platform — and Production Engineering sits at the heart... ...of Crusoe’s GPU cloud that powers next... ...problems, improving large-scale distributed systems... ...tools such as Terraform or AnsibleScripting...Temporary work$280k - $330k
...Francisco is hiring a Senior Machine Learning Engineer for Trust and Safety. The role is on-site... ...bonus. You will join a world-class ML/AI team building scalable models for content... ...AI-driven solutions. You will develop large-scale production ML models to tackle online safety...Work at office$250k - $300k
...vertically integrated AI infrastructure... ...who believe in the scale of our ambition... ...looking for a Senior Staff Network Architect... ...Architect and network engineering leadership to... ...Ansible, Nornir, or Terraform, driven by the... ...-team initiatives.Large-Scale Data Center...Temporary workWork at office$204k - $306k
...Every Identity, from AI to HumanIdentity is... ...Site Reliability Engineering GroupOkta authenticates... ...us continue to scale the service with great... ...Edge networking, K8s platform, CI/CD,... ...scaleExperience running large-scale... ...(Kubernetes), IaC (Terraform), and CI/CD pipelinesStrong...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$286.2k - $326.7k
...Senior Distinguished Engineer, AI Compute (Remote Eligible) At Capital... ...AI+ML system delivering the high-scale developer and runtime... ...use your experience in building large scale, highly available and high... ...infrastructure on top of CPU and GPU substrates. Your contributions...Full timePart timeLocal areaRemote work$250k - $300k
...vertically integrated AI infrastructure... ...believe in the scale of our ambition and... ...Role:As a Senior Staff/Principal Deployment Automation Engineer for the Compute... ...testing automation of large-scale, multi-node GPU clusters. You... ..., Docker, Terraform, and Postgres.CI/...Temporary work- ...one of the world’s largest AI infrastructure networks.... ...deliver highly available GPU infrastructure for AI... ...Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support... ...Python, Git, REST APIs, Terraform, or similar automation...Permanent employment
$170k - $250k
...infrastructure company operating in the AI space, backed by a leading... ...revenue within six months and is scaling rapidly with a small, high-performing... ...This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role...Full timeVisa sponsorshipFlexible hours$175k - $250k
...Senior Cloud Infrastructure Engineer Location: San Francisco... ...with generative AI. They are the team behind... ...deploying, and maintaining large-scale distributed systems... ...scale Manage and automate GPU compute clusters using... ...as Python, Kubernetes, Terraform, and Ansible Architect...Full timeRemote workRelocationRelocation package$200k
...high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing,... ...deep experience with Ansible, Terraform, or SaltStack. High-Scale Networking: A strong foundation... ...Solver" DNA: A background in large-scale DataCenter environments...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI. Be the first to apply!
- staff engineer San Francisco, CA
- assistant engineer San Francisco, CA
- research assistant engineering San Francisco, CA
- staff design engineer San Francisco, CA
- staff security engineer San Francisco, CA
- engineering aide San Francisco, CA
- senior staff engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- assistant chief engineer San Francisco, CA
- assistant engineering manager San Francisco, CA


