Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI/ML Infra Engineer - Hosting

$250k
Full-time

Ready to take the next step in your career?

Join a rapidly growing AI cloud infrastructure provider building high-performance compute platforms for large-scale AI training and inference workloads. With expanding GPU infrastructure across Europe and the United States, the organisation enables AI teams to access scalable compute environments without traditional infrastructure limitations.

As a Senior ML Infrastructure Engineer, the successful candidate will help build and scale Kubernetes-based machine learning platforms supporting large-scale training and inference systems. The role focuses on workload orchestration, GPU scheduling, inference optimisation, and distributed systems reliability, working alongside highly technical teams at the intersection of machine learning, cloud infrastructure, and high-performance computing.

If you would like to learn more about this opportunity, feel free to reach out and apply today!

Responsibilities:

  • Build and scale internal ML infrastructure platforms focused on AI training and inference workloads
  • Develop systems for workload orchestration, job scheduling, and reliable execution across Kubernetes environments
  • Improve and maintain inference infrastructure, including model packaging, deployment, and serving optimisation
  • Collaborate with infrastructure and platform teams to maximise GPU utilisation, hardware performance, and operational reliability
  • Design scalable systems and reusable platform capabilities that improve developer experience and operational efficiency
  • Support CI/CD, GitOps, and infrastructure automation workflows across ML platform environments
  • Troubleshoot GPU performance, distributed systems behaviour, networking, and storage bottlenecks
  • Contribute to platform architecture discussions and long-term infrastructure scalability initiatives

Skills/Must Have:

  • Strong ML engineering background with hands-on experience supporting both training and inference infrastructure
  • Experience with infrastructure engineering, platform engineering, or software engineering environments
  • Strong programming skills in Python (Go experience is a plus)
  • Deep experience with Kubernetes, including operators, CRDs, workload orchestration, and GPU scheduling
  • Comfortable operating in Linux environments and debugging GPU-related issues, including CUDA, drivers, networking, and filesystems
  • Strong systems thinking and ability to design scalable, reliable, distributed infrastructure
  • Experience with CI/CD pipelines, GitOps workflows, and infrastructure automation

Desirable Skills:

  • Familiarity with orchestration and scheduling platforms such as Kueue, Flyte, Ray, or Slurm
  • Experience with PyTorch or JAX environments
  • Hands-on experience deploying inference workloads using vLLM, SGLang, TensorRT-LLM, or Triton
  • Knowledge of GPU networking and performance optimisation, including InfiniBand, NVLink, and NCCL
  • Experience working within HPC or large-scale distributed systems environments

Benefits:

  • Stock options

Salary:

  • $250,000 base salary 
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the AI/ML Infra Engineer - Hosting in San Francisco, CA vacancy
  • $250k - $400k

     ...Define how large-scale AI systems for scientific discovery...  ...in isolation. It's building the engine that research runs on. You'll...  ...Experience building and scaling ML systems in production Strong...  ...Roles available: ML Engineer, ML Infra, Research Engineers & Research... 
    Suggested
    Remote work

    techire ai

    San Francisco, CA
    1 day ago
  •  ...Mithrl is building the world’s first commercially available AI Co-Scientist. It is a discovery engine that transforms messy biological data into insights in...  ...outcomes. ABOUT THE ROLE We are hiring an ML Engineer, Analysis and Simulation to build the core analytical... 
    Suggested
    Full time
    Work at office

    Mithrl

    San Francisco, CA
    11 hours ago
  •  ...for founding Machine Learning Engineers (MLEs) to own and improve our...  ...and user context. Unlike hosted browser solutions that introduce...  ..., or consumer-focused "AI browsers," we run AI directly...  ...This architecture creates unique ML challenges. This is a high-... 
    Suggested
    Full time
    Sleeping nights

    Composite

    San Francisco, CA
    11 hours ago
  •  ...Tilde Research is a moonshot AI lab advancing mechanistic interpretability, new architectures, and pretraining science. We build...  ...advance the frontier of intelligence. About the role: As a ML Engineer, you’ll build and operate the infrastructure that makes cutting... 
    Suggested
    Full time
    Internship

    Tilde Research

    San Francisco, CA
    11 hours ago
  •  ...Mithrl is building the world’s first commercially available AI Co-Scientist. It is a discovery engine that transforms messy biological data into insights in...  ...outcomes. ABOUT THE ROLE We are hiring an ML Engineer, Discovery Applications to build the high level... 
    Suggested
    Full time
    Work at office

    Mithrl

    San Francisco, CA
    11 hours ago
  •  ...Ship models, not slide decks — partner with research and infra to prototype, train, and deploy state-of-the-art voice models...  ...Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code across the stack. Hands... 
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    11 hours ago
  •  ...by a team with deep experience in enterprise infrastructure and AI, with leadership roots at companies including Splunk, WebLogic,...  ...About the Role We are looking for a visionary Senior ML Engineer who will bridge the gap between high-level architecture and hands... 
    Full time
    Shift work

    Palm Venture Studios

    San Francisco, CA
    11 hours ago
  • $200k - $400k

     ...Goodfire Behind our name: Like fire, AI holds the potential for both immense...  ...world’s top interpretability researchers and engineers from organizations like OpenAI and DeepMind...  ...~5+ years of experience in ML infra, research engineering, or systems programming... 
    Full time

    Goodfire

    San Francisco, CA
    11 hours ago
  •  ...anticipate your needs before you ask. We’re a team of AI researchers, designers, growth experts, and engineers rethinking human-computer interaction from the...  ...is just the beginning. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building... 
    Full time

    Wispr Flow

    San Francisco, CA
    11 hours ago
  • $244k - $320k

     ...Attentive® is the AI marketing platform for 1:1 personalization...  ...our AI-powered personalization engine delivers bespoke experiences that...  ...and operating production-grade ML systems that drive real-time personalization...  ...AirFlow, Postgres, and Redis, hosted via AWS Our infrastructure... 
    Full time

    Attentive

    San Francisco, CA
    11 hours ago
  • $129k - $198.4k

    Job DescriptionRole: As an AI/ML Engineer on the Metrics Frameworks team, part of the Simulation, Evaluation, and Data organization, you will...  ...the organization. Collaborate with other frameworks and data infra teams to build and deploy tools to improve productivity. Work... 
    Full time
    Local area
    Work from home

    General Motors

    San Francisco, CA
    3 days ago
  • $128.7k - $261.3k

     ...sustainable, and more accessible mobility. For the AI Kernels & Compilers team, that mission...  ..., kernel development, and performance engineering so that every cycle on our accelerators...  ...that path fast, reliable, and effortless for ML engineers across the AV organization to... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    2 days ago
  • $160k - $240k

     ...Fortune 500 company and a leading AI platform for managing people,...  ...the RoleAs a Machine Learning Engineer on the AI Core team, you will develop...  ...other engineers to deliver ML solutions across Workday’s...  ...experience building services to host machine learning models in production... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    San Francisco, CA
    4 days ago
  •  ...artificial intelligence and advanced ML, deep learning techniques to...  ...looking for a Machine Learning Engineer to help design, build, optimize...  ...partnership with Platform and Infra teams.Write high-quality,...  ...related field.Proficiency in using AI coding tools (e.g., Claude Code... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    4 days ago
  • $220k - $320k

     ...systems that train specialized AI models for the fastest-growing...  ...If you love taking cutting-edge ML techniques and turning them into...  ...Inference.net trains and hosts specialized language models for...  ...well-funded ten-person team of engineers who work in-person in downtown... 
    Full time
    Work at office

    Inference

    San Francisco, CA
    11 hours ago
  •  ...a journey to fix this. We're creating an AI-powered experience that replicates the flow...  ...looking for an experienced Machine Learning Engineer to join our team and help develop cutting-...  ...is an incredibly exciting time to join an ML team designing a personalized learning... 
    Full time
    Live in
    Work at office
    Worldwide

    Speak

    San Francisco, CA
    11 hours ago
  • $137.1k - $201.6k

     ...that delivers lower delivery fees and a host of additional benefits and value to a...  ...forming a new team that will leverage AI and advanced ML to power decision making in real-time –...  ...re looking for a Staff Machine Learning Engineer to drive the design and development of... 
    Hourly pay
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    From Restaurants Near You

    San Francisco, CA
    11 hours ago
  • $170k - $280k

     ...Our mission is to give leaders clarity and engineers time. We help leaders understand how...  ...the role We're looking for an Applied ML Engineer to help build and improve the machine...  ...systems that power Macroscope's core AI capabilities. Your primary focus will be on... 
    Odd job
    Full time

    Macroscope Inc

    San Francisco, CA
    11 hours ago
  • $130k - $250k

     ...What You'll Own   Build custom ML models to classify prompts, predict opportunity...  ...scoring, and evaluation systems for noisy AI commerce outputs   Build incrementality...  ...Prior founding experience or was an early engineer at a Seed, Series A, or Series B company... 
    Full time

    Talent Search Pro

    San Francisco, CA
    11 hours ago
  •  ...’s Possible. At Pinterest, AI isn't just a feature, it's a powerful...  ...architecture across the ads ML stack, and mentoring top-tier...  ...product, applied science, and infra teams to translate business...  ...partners across product, engineering and research. Bachelor’s/Master... 
    Full time
    Work at office
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    11 hours ago
  •  ...The Role We're looking for an Applied ML Engineer who thrives at the intersection of applied research and building real-world products,...  ...machine learning, applied research, or systems-level engineering for AI; 2+ years in a customer-facing technical role ~ Strong... 
    Full time
    Flexible hours

    Adaption

    San Francisco, CA
    11 hours ago
  • $77k - $202k

     ...Bachelor's Degree- At least 4 years of experience in software engineering or data engineeringWhat Sets You Apart- Master's Degree in Computer...  ...CLI tools with argparse and YAML configuration- Designing AI agent systems for development workflows- Familiarity with corporate... 
    Full time
    H1b

    PwC

    San Francisco, CA
    2 days ago
  • $197.3k - $313.7k

     ...EngineeringJob DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive customer...  ...*Slack is looking for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and ship NLP models... 
    Full time

    Salesforce

    San Francisco, CA
    3 days ago
  • $276k - $414k

     ...and other digital services.Snap Engineering teams build fun and...  ...Engineer to join the Content ML team at Snap! We build large-scale...  ...system.Work with cross-team ML, Infra, and Research partners to design...  ...judgmentExperience contributing to AI publicationsIf you have a... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    4 days ago
  • $166k - $210.25k

    RDQ127R59SummaryAs a Senior Applied ML Engineer on the Applied AI team at Databricks, you will use machine learning, scheduling, and optimization algorithms to maximize the efficiency and performance of our infrastructure. Your work will span the entire stack—from cluster... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  •  ...Airbnb was born in 2007 when two hosts welcomed three guests to their...  ...to Marketing we rely on ML to ensure that guests and hosts...  ...initiatives by adopting the Generative AI technologies to enable an...  ...product, design, and other engineering counterparts to design and build... 
    Remote job
    Full time
    Casual work
    Live in
    Work at office

    Airbnb, Inc.

    San Francisco, CA
    11 hours ago
  • $155k - $180k

     ...machine learning open source and hosted tools. That includes...  ...all roles (not only product and engineering), so Roboflow employs developers...  ...increasingly authored with the help of AI agents. That's a great problem...  ...latest computer vision and ML models to our users. Teach... 
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    11 hours ago
  •  ...Monitoring Platform Raindrop is the monitoring platform for AI agents. Engineering teams at Fortune 100s and the fastest-growing AI companies...  ...requests a day + growing. Architect, implement, and scale ML pipelines Quick iteration without compromising on quality... 
    Temporary work

    Raindrop

    San Francisco, CA
    3 days ago
  •  ...expect This role is for a machine learning engineer who wants to work on the models that give...  ...mechanical team at the frontier of embodied AI, where the problems are genuinely open and...  ...machine learning models Integrate ML systems with robotic hardware and embedded... 
    Immediate start

    Human Computer Lab

    San Francisco, CA
    1 day ago
  •  ..., evolves, and documents codebases autonomously. As a Founding ML Engineer , you’ll architect the intelligence powering autonomous pull requests...  ..., and self-improving workflows. You’ll work across product, infra, and full-stack teams to embed real-time decision-making into... 
    Remote work
    Flexible hours

    Kodezi Inc.

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI/ML Infra Engineer - Hosting. Be the first to apply!