Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Infrastructure Engineer

$200k - $300k

Recruiting from Scratch

Who is Recruiting from Scratch:

Recruiting from Scratch is a premier talent firm that focuses on placing the best product managers, software, and hardware talent at innovative companies. Our team is 100% remote and we work with teams across the United States to help them hire.

ML Infrastructure Engineer
Location - San Francisco, CA - (On-site - 5 Days Per Week)
Compensation - $200,000 - $300,000 Base + Competitive Equity
Visa - Visa Sponsorship Available (New H1B, H1B Transfers, TN, Eligible Visa Types)
Company Stage - Series B ($100M+ Raised)
Industry

Artificial Intelligence, Machine Learning Infrastructure, AI Infrastructure, Developer Tools, Enterprise AI, Document AI, Enterprise SaaS
About the Company

The company is building one of the leading AI infrastructure platforms powering document intelligence for modern AI applications.

Its platform transforms complex documents-including PDFs, spreadsheets, presentations, images, and enterprise files-into structured, production-ready data for AI systems. Processing tens of millions of pages every month, the platform serves hundreds of customers ranging from AI-native startups to Fortune 10 enterprises, enabling reliable document ingestion for large language models and AI workflows.

Backed by Andreessen Horowitz, Benchmark, and First Round Capital with over $100M in funding, the company has experienced exceptional growth, increasing revenue more than 8x year-over-year while becoming one of the fastest-growing AI infrastructure companies in the market.

As an ML Infrastructure Engineer, you'll own critical training and inference infrastructure powering production machine learning systems. You'll work closely with ML researchers and product engineers to ensure models are trained, deployed, monitored, and served efficiently while continuously improving latency, reliability, scalability, and cost across production AI workloads.

This is an exceptional opportunity to join a fast-growing AI-native company where ML Infrastructure Engineers build the foundational systems enabling next-generation AI products at massive production scale.
What You'll Do
  • Build and maintain production ML model serving infrastructure powering millions of document processing requests
  • Design and improve training infrastructure supporting models ranging from hundreds of millions to tens of billions of parameters
  • Optimize inference latency, throughput, reliability, and GPU utilization across production systems
  • Develop observability, monitoring, logging, and alerting across the ML infrastructure stack
  • Build internal tooling and data pipelines enabling ML researchers to rapidly deploy models into production
  • Architect infrastructure that intelligently routes inference workloads across multiple cloud providers
  • Optimize infrastructure for model accuracy, latency, reliability, and operational cost
  • Collaborate closely with ML researchers, software engineers, and product teams to accelerate experimentation
  • Improve Kubernetes-based infrastructure supporting large-scale ML workloads
  • Debug complex production infrastructure issues involving GPUs, distributed systems, and model serving
  • Design scalable systems supporting rapid AI model deployment and iteration
  • Champion engineering excellence across ML infrastructure, automation, and production reliability
Ideal Candidate Background
Experience Requirements
  • 3+ years of ML Infrastructure Engineering experience
  • Experience building production ML model serving and inference infrastructure
  • Experience working at AI-native startups or organizations training production ML models
  • Experience deploying and maintaining large-scale production ML systems
  • Experience supporting ML researchers with production infrastructure
  • Strong startup ownership mentality with demonstrated engineering execution
  • Experience building highly scalable infrastructure supporting AI products
  • Experience collaborating closely with ML, Product, and Infrastructure teams
  • Experience optimizing production inference systems
  • Previous experience working with GPU-intensive workloads strongly preferred
Technical Requirements
  • Strong software engineering fundamentals across distributed systems and infrastructure
  • Expert-level Python programming skills
  • Strong Kubernetes, Docker, Helm, and cloud infrastructure experience
  • Deep understanding of GPU optimization, model serving, and inference infrastructure
  • Experience with PyTorch and modern ML deployment workflows
  • Experience building observability, monitoring, and logging systems for ML infrastructure
  • Strong understanding of model deployment, training pipelines, and production inference
  • Experience working with 1-3 node model training and single/double-node serving
  • Familiarity with cloud platforms including AWS, Azure, or GCP
  • Ability to rapidly debug production infrastructure and performance bottlenecks
Education
  • Bachelor's degree in Computer Science, Machine Learning, Engineering, Mathematics, or another technical discipline preferred
  • Strong engineering background with demonstrated infrastructure expertise
  • Top technical universities, competitive programming backgrounds, or exceptional engineering accomplishments strongly preferred
  • Publications at top ML conferences (NeurIPS, ICML, ICLR, CVPR) are a strong plus
Soft Skills
  • Outstanding technical problem-solving
  • Strong engineering ownership
  • Excellent collaboration skills
  • High execution mentality
  • Comfortable operating with ambiguity
  • Strong cross-functional communication
  • Startup mindset with high execution velocity
  • Structured systems thinker
  • Passion for AI infrastructure and production ML systems
  • Continuous learner with deep technical curiosity
Compensation & Benefits
  • Base Salary: $200,000 - $300,000
  • Competitive Equity Package
  • Visa Sponsorship Available
  • On-site work model (5 days/week in San Francisco)
  • Opportunity to build core infrastructure powering one of the fastest-growing AI platforms
  • Direct collaboration with ML researchers and engineering leadership
  • Significant ownership across production AI infrastructure
  • High-impact engineering culture with exceptional career growth
  • Well-funded Series B company backed by leading venture firms
  • Opportunity to build foundational infrastructure enabling next-generation AI systems
Why Join

This is an opportunity to build the infrastructure powering one of the fastest-growing AI platforms transforming how enterprises process information.

You'll work at the intersection of machine learning infrastructure, distributed systems, GPU optimization, model serving, production engineering, and AI infrastructure while solving some of the hardest scalability challenges in modern AI.

As an ML Infrastructure Engineer, you'll own foundational systems that enable production AI models to operate reliably, efficiently, and at massive scale while helping define the future of enterprise AI infrastructure.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the ML Infrastructure Engineer in San Francisco, CA vacancy
  •  ...oprecruiting.comPhone: (***) ***-****Job Title: Machine Learning Infrastructure EngineerLocation: San Francisco, CA Metro Area (100% On-Site)...  ...company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    4 days ago
  • $180k - $230k

     ...ML Infrastructure Engineer San Francisco Company Overview Echo Neurotechnologies is an exciting new startup in the Brain-Computer Interface (BCI) space, driving innovation through advanced hardware engineering and AI solutions. Our mission is to deliver cutting... 
    Suggested
    Flexible hours

    Echo Neurotechnologies

    San Francisco, CA
    5 days ago
  • $190k - $250k

     ...Senior ML Infrastructure Engineer Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid. We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the... 
    Suggested
    Local area

    Gridware

    San Francisco, CA
    4 days ago
  • $180k - $250k

     ...generation of AI. We are developing the context engine layer that solves a fundamental...  ...wave of AI progress will come from better infrastructure around models: Better Memory &...  ...Khan, CEO: ex-Amazon; PhD in Robotics and ML. Clark Zhang, CTO: ex-Meta; PhD in... 
    Suggested

    GraphOn

    San Francisco, CA
    10 hours ago
  •  ...Role Description We are looking to recruit an exceptional Infrastructure Engineer to own and build the backend systems that power machine...  ...Lead technical and commercial discussions with cloud and ML compute providers, including capacity planning, performance,... 
    Suggested

    Maven Robotics

    San Francisco, CA
    3 days ago
  • $250k

     ...ML Infrastructure Engineer – Open Source ML Infra – Up to $250K Total Comp: $500K – Hybrid We are working with one of the leading ML Infra companies supporting hundreds of custom LLMs. You will join a talented Infrastructure engineering team to drive performance... 
    Flexible hours

    well-funded deeptech startup

    San Francisco, CA
    4 days ago
  •  ...Job Description Job Description We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and... 
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    14 days ago
  • $160k - $250k

     ...enjoy facing adversity, and can do the impossible at record breaking speeds. About You and The Role  As an ML Training & Inference Infrastructure Engineer on the Data Platform team you will be building and scaling the systems powering our data flywheel.  This person... 
    Local area

    Zipline

    San Francisco, CA
    17 days ago
  •  ...The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to automate the lifecycle of our AI workflow...  ...of video streams. As a Machine Learning Engineer in ML Runtime & Optimization , you will develop... 
    Full time

    Zensors

    San Francisco, CA
    a month ago
  • $130.2k - $195.3k

     ...Senior ML Platform Engineer We're on a mission to unleash the power of content… you in? We've got the brands, we've got the stars, we've got the power to achieve our mission to entertain the planet – now all we're missing is… YOU! Becoming a part of Paramount means... 

    Paramount Global Services

    San Francisco, CA
    5 days ago
  • $197.3k - $225.1k

     ...Lead AI/ML Engineer (Platform, kubeflow) Overview At Capital One, we are creating responsible and reliable AI systems, changing...  ...customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine... 
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    1 day ago
  • $292k - $417.2k

     ...ads optimization that shape the future of streaming.We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    3 days ago
  •  ...still be proud of in ten years from now. Machine Learning Engineer, Platform About the Role So much of the work society depends...  ...a Machine Learning Engineer on Platform, you'll build the core ML capabilities every product we ship stands on — built once,... 
    Full time
    Shift work

    Brain Corp

    San Francisco, CA
    a month ago
  • $200k - $400k

     ...Machine Learning Engineer - Infrastructure Company: Causal Labs Location: San Francisco, CA (South Park office, in person 5 days per week...  ...performance for large models. If you have built large-scale ML infrastructure for language, vision, robotics or biology... 
    Full time
    Work at office
    Relocation
    Visa sponsorship

    Transparent Search Group

    San Francisco, CA
    2 days ago
  •  ...in the enterprise through purpose-built ML systems that learn from Rippling's proprietary...  ...at Rippling.As a Staff Machine Learning Engineer, you will own the end-to-end ML...  ...Build robust evaluation and experimentation infrastructure: offline benchmarks, A/B testing, and... 
    Work at office
    3 days per week

    Rippling

    San Francisco, CA
    1 day ago
  • $295k - $405.5k

     ...roleAs the Senior Staff Machine Learning Platform Engineer, you will own the technical vision and evolution of Faire’s ML platform. You will set standards, influence...  ...background in distributed systems, ML infrastructure, and cloud architecture.Demonstrated technical... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    2 days ago
  • $246.5k - $339k

     ...this roleAs a Staff Machine Learning Platform Engineer, you will help design, improve, and operate a scalable ML platform to accelerate model training, deployment...  ....What You Will DoDesign and operate ML infrastructure, including workspaces, clusters, jobs, and workflowsProductionize... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    3 days ago
  • $212k - $318k

     ...creator economy and are looking for a Senior Machine Learning Engineer, Infrastructure to support our mission.This role is based in San Francisco...  ...TeamYou'll join the Relevance team, whose mission is to build the ML systems that power how fans discover creators and how... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours
    2 days per week
    3 days per week

    Patreon

    San Francisco, CA
    4 days ago
  •  ...leave a clean trail of tests and dashboards behind. Required Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code across the stack. Hands-on experience training or fine-tuning large language or other... 
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    more than 2 months ago
  • $200k - $300k

     ...building that model: the Food Foundation Model. As a Senior ML Engineer, Foundation Models, you will work at the frontier of large-...  ...fine-tuning the Food Foundation Model, building the data infrastructure that feeds it, and deploying it onto physical robots in production... 
    Full time
    Flexible hours

    Chef Robotics

    San Francisco, CA
    more than 2 months ago
  •  ...Gentoro was founded by a team with deep experience in enterprise infrastructure and AI, with leadership roots at companies including Splunk,...  ...About the Role We are looking for a visionary Senior ML Engineer who will bridge the gap between high-level architecture and... 
    Full time
    Shift work

    Palm Venture Studios

    San Francisco, CA
    a month ago
  •  ...their patients. Position Overview We are hiring two ML Engineers / Researchers to help build the next generation of Knowtex's...  ...clinical outcomes Build datasets, benchmarks, and evaluation infrastructure that make model improvements measurable and reproducible... 
    Full time

    Knowtex

    San Francisco, CA
    a month ago
  • $245k - $345k

     ...out the latest Whatnot updates on our news and engineering blogs and join us as we enable anyone to turn their...  ...engineers eager to shape the future of AI and ML at Whatnot. You'll design and scale the core infrastructure that powers machine learning and self-hosted large... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Work from home
    Home office
    Flexible hours

    Whatnot

    San Francisco, CA
    3 days ago
  •  ...-scale training running on top of that: infrastructure orchestration, distributed compute, and...  ...tolerant infrastructure for distributed ML. GPU clusters, NVIDIA runtime, S3 checkpointing...  ...For Infrastructure and platform engineering (required) : Production experience with... 
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Pluralis

    San Francisco, CA
    4 days ago
  •  ...cost-efficient, enterprise grade search infrastructure. Powering search and discovery in the agentic...  ..., and Product to determine where new ML capabilities can meaningfully change...  ...simplify fragmented systems, mentor senior engineers, and align technical and product leaders... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    5 hours ago
  •  ...000s of developers and enterprise users.ML performance, quality, and systems acumen...  ...Qualifications 7+ years of software engineering experience building and operating enterprise...  ...building the next layer of enterprise AI infrastructure.Job SummaryCategory: Engineering
    Work at office

    Atlassian

    San Francisco, CA
    2 days ago
  • Title: ML Engineer Location: San Francisco, CA (Onsite) Direct HireCompany Mission Our client’s mission is to scale medical knowledge...  ...influence on the development of next-generation predictive infrastructure for healthcare.Opportunities to publish, co-author patents,... 

    Spectraforce Technologies

    San Francisco, CA
    1 day ago
  • $170.1k - $258.3k

     ...export, kernel development, and performance engineering so that every cycle on our accelerators...  ...that sit at the heart of our on‑vehicle ML inference for ADAS and autonomous...  ...workloads.Build and improve tooling and infrastructure that make it easier to profile, debug, and... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    4 days ago
  • $128.7k - $261.3k

     ...approaches to model export, kernel development, and performance engineering so that every cycle on our accelerators translates into better...  ...tooling that makes that path fast, reliable, and effortless for ML engineers across the AV organization to compile their models.... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    1 day ago
  • $92k - $138k

     ...The opportunity Unity Vector builds an offline ML platform that powers insight, experimentation, attribution...  ...production ML systems. We’re looking for a Machine Learning Engineer to join our Offline Infrastructure team. This is an ideal role for a recent university... 
    Work at office
    Worldwide
    Relocation package

    UNITY

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Infrastructure Engineer. Be the first to apply!