Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer

vCluster Labs

As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engagement scoped to reach production. You will be one of the first team members a neocloud or AI Factory engages with at a technical depth, and the playbooks you develop will scale the motion for the next hire and customer.

vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen.

As an AI Infrastructure Engineer, your role will include:

  • Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment.

  • Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand.

  • Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.

  • Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently.

  • Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start.

  • Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap.

  • Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.

This role could be a fit for you if you bring:

  • Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.

  • GPU Fluency: Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.

  • Networking Fundamentals: Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments.

  • Storage Expertise: Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn.

  • Operational Agility: Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time.

  • Modern Tech Mindset: You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services.

Bonus points for:
  • Automation Skills: Experience writing automation scripts with Bash, Python, or Go.

  • Kubernetes Depth: Relevant certifications such as CKA (Certified Kubernetes Administrator) or experience writing Kubernetes Operators.

  • AI/ML Familiarity: Experience with inference serving, GPU scheduling, and the tooling around LLM deployment.

  • Documentation: Experience building AI Automation in documentation to contribute to a shared knowledge base.

About vCluster Labs

We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.

We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.

We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

Benefits

We offer the following benefits:

  • Competitive Salary : We offer a competitive compensation package, including equity.

  • Platinum-Level Insurance : Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).

  • Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

  • Workplace Flexibility:  We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Culture & Values

At vCluster Labs, we value and stand for:

  1. Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.

  2. Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.

  3. Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.

  4. Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy — the strongest ideas win, no matter who or where they come from.

  5. Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in United States vacancy
  •  ...AI Infrastructure Engineer Redstone Arsenal/Huntsville, AL At IPT Associates (IPTA), we enjoy solving real-world problems with technology. We work closely with our customers, teammates, experts, and partners to come up with practical solutions that actually make... 
    Suggested
    Full time

    Interactive Process Technology Llc

    Remote
    21 days ago
  • $154.39k - $247.02k

     ...ownership and drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter.AI Infrastructure Engineer, Corporate AI TeamTeam & Role OverviewAxon’s Corporate AI Team sits within Business Technology and builds internal-facing... 
    Suggested
    Work experience placement

    Axon

    San Francisco, CA
    1 day ago
  • $184k - $287.5k

    AI Infrastructure Engineers at NVIDIA build the systems, tooling, and data infrastructure that enable operation of our GPU cloud services. We are enabling engineering teams to innovate while proactively identifying, tracking, and mitigating risks across the entire technical... 
    Suggested
    Full time
    Remote work

    Nvidia

    Washington DC
    8 hours ago
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered...  ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Detroit, MI
    8 hours ago
  • $150k - $170k

     ...affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure... 
    Suggested
    Work experience placement
    Local area
    Remote work
    Worldwide

    Analysis Group

    Boston, MA
    3 days ago
  • $215k - $350k

    We are seeking an AI Infrastructure Engineer to build, operate, and continuously enhance the Linux and GPU-based infrastructure that powers our AI platforms and performance testing environments. This is a highly hands-on infrastructure, automation, and performance engineering... 
    Worldwide
    Home office

    Fortinet

    New York, NY
    3 days ago
  •  ...kJob Description A growing fintech company focused on modern financial services and trading technology is looking for an AI Infrastructure Engineer to build and operate the platform supporting its next generation of AI-powered financial workflows. The company combines... 
    Full time
    Work from home

    Motion Recruitment

    Miami, FL
    3 days ago
  • $191k - $253k

     ...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands...  ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems...  ...THE JOB We are looking for a Senior AI Infrastructure Engineer to build, scale, and optimize the... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    6 hours ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded...  ...THE ROLE: We are seeking a DevOps / Platform Engineer to join our team building and operating large-scale GPU compute infrastructure that powers AI and ML workloads. THE PERSON:... 

    AMD

    San Jose, CA
    8 hours ago
  •  ...are poised to disrupt the billion-dollar engineering simulation industry with our fast-...  ...You will help develop and improve core infrastructure and user-facing capabilities across our...  ...generation of engineering software and AI-enabled workflows. What You’ll Work... 
    Full time

    Flexcompute

    Remote
    more than 2 months ago
  • $191k - $315k

     ...Overview:  The Network Growth and Relationship AI team is at the forefront of creating...  ...close collaboration with the product, engineering and data science team and has a very...  ...Prior experience with large scale ML data infrastructure ~ Experience with developing and designing... 
    Full time
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    more than 2 months ago
  •  ...The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to automate the lifecycle of our AI workflow, including model development, evaluation, optimization, deployment, and monitoring across thousands of video... 
    Full time

    Zensors

    San Francisco, CA
    a month ago
  • ​GTSC seeks a   Cloud Architect Artificial Intelligence (AI) Engineer , Level IV, to support our customer in the   Annapolis Junction, MD   area. Location:   Annapolis Junction, Maryland All work is on-site. This is not a hybrid or remote position.   Mission... 
    Full time
    Temporary work

    Gtsc-Talent Solutions

    Remote
    more than 2 months ago
  • $175k - $275k

     ...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in...  ...would be possible. The Role As an AI Engineer - Cloud Infrastructure on Traversal’s Infrastructure team, you’ll design,... 
    Full time
    Work at office
    Flexible hours

    Traversal

    New York, NY
    more than 2 months ago
  •  ...department founder and collaborate with a world-class team of engineers, researchers, and commercial leaders. Competitive...  ...execute core 0-to-1 initiatives in cutting-edge domains such as AI Infrastructure. You will translate high-level strategic vision into concrete... 
    Remote job
    Full time
    Relocation

    Ewor Gmbh

    Huntsville, AL
    12 days ago
  • $220k - $292k

     ...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands...  ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems...  ...We are looking for a founding Staff AI Infrastructure Engineer to architect, build, and scale... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $190k - $260k

     ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...large-scale world models - depends on infrastructure that turns thousands of hours of multimodal...  ...training throughput. We are looking for engineers who make model training fast: streaming... 
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    3 days ago
  • At the National Robotics Engineering Center (NREC), it is our engineers and technicians who drive the breakthroughs that define our...  ...career development.We are seeking a dynamic Senior Applied AI Infrastructure Engineer to lead and contribute to the evaluation, deployment... 
    Full time
    Work experience placement
    Flexible hours

    Carnegie Mellon University

    Pittsburgh, PA
    4 days ago
  •  ...AI Platform Engineer About Tessera Labs Tessera Labs is redefining how enterprises adopt and operationalize Artificial Intelligence...  ...platform team builds and operates the foundational AI agent infrastructure that lets Tessera run reliably, securely, and consistently... 
    Remote work

    Tessera Labs

    United States
    1 day ago
  •  ...Role Overview: Primary hiring focus is an AI Infrastructure Senior Engineer supporting the build-out of the company's Azure-based technology stack. The role will be the first hire on the infrastructure engineering team and will work under the infrastructure... 
    Currently hiring
    Work at office
    Remote work

    ConsultNet

    New York, NY
    5 days ago
  •  ...AI Infrastructure Engineer Percepta's mission is to transform critical institutions with applied AI. We care that industries that power the world (e.g. healthcare, manufacturing, energy) benefit from frontier technology. To make that happen, we embed with industry... 
    Contract work

    Percepta

    New York, NY
    3 days ago
  •  ...Maxonic maintains a close and long-term relationship with our direct client. In support of their needs, we are looking for an AI Infrastructure Engineer. Job Title: AI Infrastructure Engineer Job Type: Contract / Contract to Hire Job Location: San... 
    Contract work

    Maxonic

    San Jose, CA
    3 days ago
  •  ...Senior AI Infrastructure Engineer Austin, Texas, United States; Reston, Virginia, United States Seekr is building the infrastructure that powers the next generation of enterprise AI. As a Senior AI Infrastructure Engineer, you will design, build, and operate the... 
    Permanent employment
    Work experience placement
    Flexible hours

    Seekr

    Austin, TX
    4 days ago
  •  ...Job Title: AI Infrastructure Engineer Location: Remote, USA Job Description This role focuses on managing and optimizing our AI infrastructure, ensuring seamless operations, and providing guidance and training to our team members. The ideal candidate will... 
    Remote work

    United IT Solutions

    Dallas, TX
    13 hours ago
  • $180k - $240k

     ...facilitating effortless integration into customers' logistics operations. About the role We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high-performance AI platform powering our autonomous driving models. While researchers focus... 
    Odd job
    Work at office

    Gatik AI

    Santa Clara, CA
    5 days ago
  • Job Description: RADIOLOGY PARTNERS OVERVIEW Radiology Partners, through its owned and affiliated practices, is a leading radiology practice in the U.S., serving hospitals and other healthcare facilities across the nation. As a physician-led and physician...
    Remote work

    Radiology Partners

    United States
    5 days ago
  •  ...The Role: Spellbrush, the world’s leading generative AI studio behind niji・journey , is looking for an AI Infrastructure Engineer to join us in building out end-to-end ML infrastructure to run our models on all platforms. What you’ll do: Design, implement... 
    Work experience placement
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    2 days ago
  •  ...To build automated, intelligence-driven systems that support NVIDIA's critical AI platforms, the full-time Senior AI Infrastructure Engineer will manage scalable telemetry pipelines, standardize operational workflows, and maintain infrastructure catalogs while working... 
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    1 day ago
  •  ...AI Infrastructure Engineer (Forward Deployed AI Engineer) Location: Fort Mill, SC (Onsite – 5 Days/Week) Tax Term (W2, C2C): W2 / C2C Job Type (Permanent/Contract): Permanent Duration: Full-Time Description We are building a Forward Deployed AI Engineering... 
    Permanent employment
    Full time
    Contract work

    Apolis

    Fort Mill, York County, SC
    5 days ago
  •  ...Design, build, and operate on-prem infrastructure that behaves like a cloud environment for internal teams, including AI/ML workloads Own datacenter and infrastructure...  ...hardware troubleshooting) in addition to infra engineering Benefits ~ Medical Insurance... 
    Flexible hours

    AMAX

    Fremont, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!