ML Infrastructure Engineer
$200k - $300kRecruiting from Scratch
Who is Recruiting from Scratch: Recruiting from Scratch is a premier talent firm that focuses on placing the best product managers, software, and hardware talent at innovative companies. Our team is 100% remote and we work with teams across the United States to help them hire. ML Infrastructure Engineer
Location - San Francisco, CA - (On-site - 5 Days Per Week)
Compensation - $200,000 - $300,000 Base + Competitive Equity
Visa - Visa Sponsorship Available (New H1B, H1B Transfers, TN, Eligible Visa Types)
Company Stage - Series B ($100M+ Raised)
Industry Artificial Intelligence, Machine Learning Infrastructure, AI Infrastructure, Developer Tools, Enterprise AI, Document AI, Enterprise SaaS
About the Company The company is building one of the leading AI infrastructure platforms powering document intelligence for modern AI applications. Its platform transforms complex documents-including PDFs, spreadsheets, presentations, images, and enterprise files-into structured, production-ready data for AI systems. Processing tens of millions of pages every month, the platform serves hundreds of customers ranging from AI-native startups to Fortune 10 enterprises, enabling reliable document ingestion for large language models and AI workflows. Backed by Andreessen Horowitz, Benchmark, and First Round Capital with over $100M in funding, the company has experienced exceptional growth, increasing revenue more than 8x year-over-year while becoming one of the fastest-growing AI infrastructure companies in the market. As an ML Infrastructure Engineer, you'll own critical training and inference infrastructure powering production machine learning systems. You'll work closely with ML researchers and product engineers to ensure models are trained, deployed, monitored, and served efficiently while continuously improving latency, reliability, scalability, and cost across production AI workloads. This is an exceptional opportunity to join a fast-growing AI-native company where ML Infrastructure Engineers build the foundational systems enabling next-generation AI products at massive production scale.
What You'll Do
Experience Requirements
Location - San Francisco, CA - (On-site - 5 Days Per Week)
Compensation - $200,000 - $300,000 Base + Competitive Equity
Visa - Visa Sponsorship Available (New H1B, H1B Transfers, TN, Eligible Visa Types)
Company Stage - Series B ($100M+ Raised)
Industry Artificial Intelligence, Machine Learning Infrastructure, AI Infrastructure, Developer Tools, Enterprise AI, Document AI, Enterprise SaaS
About the Company The company is building one of the leading AI infrastructure platforms powering document intelligence for modern AI applications. Its platform transforms complex documents-including PDFs, spreadsheets, presentations, images, and enterprise files-into structured, production-ready data for AI systems. Processing tens of millions of pages every month, the platform serves hundreds of customers ranging from AI-native startups to Fortune 10 enterprises, enabling reliable document ingestion for large language models and AI workflows. Backed by Andreessen Horowitz, Benchmark, and First Round Capital with over $100M in funding, the company has experienced exceptional growth, increasing revenue more than 8x year-over-year while becoming one of the fastest-growing AI infrastructure companies in the market. As an ML Infrastructure Engineer, you'll own critical training and inference infrastructure powering production machine learning systems. You'll work closely with ML researchers and product engineers to ensure models are trained, deployed, monitored, and served efficiently while continuously improving latency, reliability, scalability, and cost across production AI workloads. This is an exceptional opportunity to join a fast-growing AI-native company where ML Infrastructure Engineers build the foundational systems enabling next-generation AI products at massive production scale.
What You'll Do
- Build and maintain production ML model serving infrastructure powering millions of document processing requests
- Design and improve training infrastructure supporting models ranging from hundreds of millions to tens of billions of parameters
- Optimize inference latency, throughput, reliability, and GPU utilization across production systems
- Develop observability, monitoring, logging, and alerting across the ML infrastructure stack
- Build internal tooling and data pipelines enabling ML researchers to rapidly deploy models into production
- Architect infrastructure that intelligently routes inference workloads across multiple cloud providers
- Optimize infrastructure for model accuracy, latency, reliability, and operational cost
- Collaborate closely with ML researchers, software engineers, and product teams to accelerate experimentation
- Improve Kubernetes-based infrastructure supporting large-scale ML workloads
- Debug complex production infrastructure issues involving GPUs, distributed systems, and model serving
- Design scalable systems supporting rapid AI model deployment and iteration
- Champion engineering excellence across ML infrastructure, automation, and production reliability
Experience Requirements
- 3+ years of ML Infrastructure Engineering experience
- Experience building production ML model serving and inference infrastructure
- Experience working at AI-native startups or organizations training production ML models
- Experience deploying and maintaining large-scale production ML systems
- Experience supporting ML researchers with production infrastructure
- Strong startup ownership mentality with demonstrated engineering execution
- Experience building highly scalable infrastructure supporting AI products
- Experience collaborating closely with ML, Product, and Infrastructure teams
- Experience optimizing production inference systems
- Previous experience working with GPU-intensive workloads strongly preferred
- Strong software engineering fundamentals across distributed systems and infrastructure
- Expert-level Python programming skills
- Strong Kubernetes, Docker, Helm, and cloud infrastructure experience
- Deep understanding of GPU optimization, model serving, and inference infrastructure
- Experience with PyTorch and modern ML deployment workflows
- Experience building observability, monitoring, and logging systems for ML infrastructure
- Strong understanding of model deployment, training pipelines, and production inference
- Experience working with 1-3 node model training and single/double-node serving
- Familiarity with cloud platforms including AWS, Azure, or GCP
- Ability to rapidly debug production infrastructure and performance bottlenecks
- Bachelor's degree in Computer Science, Machine Learning, Engineering, Mathematics, or another technical discipline preferred
- Strong engineering background with demonstrated infrastructure expertise
- Top technical universities, competitive programming backgrounds, or exceptional engineering accomplishments strongly preferred
- Publications at top ML conferences (NeurIPS, ICML, ICLR, CVPR) are a strong plus
- Outstanding technical problem-solving
- Strong engineering ownership
- Excellent collaboration skills
- High execution mentality
- Comfortable operating with ambiguity
- Strong cross-functional communication
- Startup mindset with high execution velocity
- Structured systems thinker
- Passion for AI infrastructure and production ML systems
- Continuous learner with deep technical curiosity
- Base Salary: $200,000 - $300,000
- Competitive Equity Package
- Visa Sponsorship Available
- On-site work model (5 days/week in San Francisco)
- Opportunity to build core infrastructure powering one of the fastest-growing AI platforms
- Direct collaboration with ML researchers and engineering leadership
- Significant ownership across production AI infrastructure
- High-impact engineering culture with exceptional career growth
- Well-funded Series B company backed by leading venture firms
- Opportunity to build foundational infrastructure enabling next-generation AI systems
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the ML Infrastructure Engineer in San Francisco, CA vacancy
- ...oprecruiting.comPhone: (***) ***-****Job Title: Machine Learning Infrastructure EngineerLocation: San Francisco, CA Metro Area (100% On-Site)... ...company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks...SuggestedFull timeWork at officeFlexible hours
- ...The RoleAt Mach9, ML infrastructure engineers build and maintain the systems that power production AI models for civil engineering and surveying. Our ML pipeline spans 10,000+ miles of labeled survey data, image segmentation networks, and 3D prediction models serving real...SuggestedWork experience placement
$200k - $300k
...ML Infrastructure EngineerThe company is building one of the leading AI infrastructure platforms powering document intelligence for modern AI applications.As an ML Infrastructure Engineer, you'll own critical training and inference infrastructure powering production machine...Suggested- ...ML Infrastructure EngineerSan FranciscoCompany OverviewEcho Neurotechnologies is an exciting new startup in the Brain-Computer Interface (BCI) space, driving innovation through advanced hardware engineering and AI solutions. Our mission is to deliver cutting-edge technologies...SuggestedFlexible hours
$190k - $250k
...Senior ML Infrastructure EngineerGridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical... ...and Silicon Valley investors.As a Senior ML Infrastructure Engineer, you will work directly in the Automation org with the core...SuggestedLocal area- ...Role Description We are looking to recruit an exceptional Infrastructure Engineer to own and build the backend systems that power machine... ...Lead technical and commercial discussions with cloud and ML compute providers, including capacity planning, performance,...
$180k - $250k
...generation of AI. We are developing the context engine layer that solves a fundamental... ...wave of AI progress will come from better infrastructure around models: Better Memory &... ...Khan, CEO: ex-Amazon; PhD in Robotics and ML. Clark Zhang, CTO: ex-Meta; PhD in...- Anthropic is seeking an engineer to own the infrastructure behind safeguards research. You will build the tooling researchers rely on to run experiments, train detection methods, and select detections for launch. This role sits between research and production, ensuring...
- Anthropic is seeking an engineer for its Safeguards team to own the infrastructure behind cutting-edge AI research. You will build the tooling, pipelines, and production workflows that enable researchers to run experiments, train detection methods, and deploy results at...
$166k - $244k
A leading technology company is seeking a Research Software Engineer to develop next-generation technologies that enhance communication through advanced software and hardware. This role involves collaborating with researchers, optimizing machine learning performance, and...- ...moonshot AI lab focused on mechanistic interpretability, new architectures, and pretraining science. As an ML Engineer, you will build and operate the infrastructure enabling cutting-edge research in training and evaluating large models. You will optimize inference and...
$150k - $250k
Garuda Ventures is looking for a Senior Software Engineer (IC) to develop cloud and on-prem systems that will power future factories... ...software and a strong background in backend systems and infrastructure. The position offers a competitive salary range of $150,000—$...- ...Inc. seeks a Staff / Sr. Machine Learning Engineer for the AI, Search & Knowledge Platforms... ...Safari. The role emphasizes building ML pipelines, applying NLP and entity linking... ...on open web content, and advancing ML infrastructure for scalable data processing across billions...
- Parafin, Inc. is seeking a Software Engineer to lead the evolution of its ML Platform within the Infrastructure team. You will build scalable, reliable systems for model experimentation, training, evaluation, inference, and retraining powering underwriting and ML-driven...
- ...environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. We're hiring Machine Learning Engineers to build and improve the AI systems that help strategic partners adapt OpenAI models to #J-18808-Ljbffr JobCubby
- ...Francisco is seeking a Machine Learning Engineer to develop tools for interpretable AI systems. The role involves optimizing ML pipelines, implementing research into production... ...have over 5 years of experience in ML infrastructure, with expertise in Python and distributed...
- Brain Co. seeks a Machine Learning Engineer to build platform-level ML capabilities that power multiple products across industries. You will work on agent-native systems, foundation models, and production-grade pipelines that improve document extraction, reasoning, and...
- Lantern is seeking a Machine Learning Engineer in San Francisco to scale the core ML platform behind our product. You will own the data ingestion, training pipelines, and the end-to-end ML workflow that generates customer recommendations used across all accounts. We expect...Work at office
- ...automate complex construction documents and other workflows using agents and foundation models. As a Machine Learning Engineer on Platform, you will own core ML capabilities end-to-end, push frontier models into production, and enable dozens of institutional workflows with...
- Faire is seeking a Staff Machine Learning Platform Engineer to design, improve, and operate a scalable ML platform that accelerates model training,... ...Spark, Delta Lake, MLflow, Python and SQL, cloud/infrastructure-as-code, and strong MLOps practices. #J-18808-Ljbffr...Remote jobLocal area
- A leading technology company in San Francisco is seeking an experienced Software Engineer for its Machine Learning Platform team. You will design and build services that support Apple’s machine learning and computer vision efforts. Responsibilities include leading systems...
- CoreWeave, the AI hyperscaler, is seeking a Senior Software Engineer to own the architecture and evolution of core systems in the Weights & Biases platform. You will drive ML Workflows, Artifacts, Registry, Automations, and Launch, collaborating with product, design, and...
- ...generative AI company, is hiring a Machine Learning Engineer focused on Ads in the San Francisco Bay Area.... ...will work at the intersection of large-scale ML, generative AI, and advertising systems, building models and infrastructure that determine how creative is generated,...
- Google Cloud is seeking a Customer Engineer in San Francisco to partner with the technical Sales team and serve as the AI/ML subject matter expert to differentiate Google Cloud for customers. You will help prospects and partners understand Google Cloud, develop cloud architectures...
- Sardine is seeking a Senior Data/ML Engineer to own the data and machine learning foundation for compliance decisions. You will design and maintain end-to-end pipelines that transform raw telemetry and signals into features used by models across fraud and KYC on a remote...Remote job
- Superhuman is building a platform for AI Agents, and as a Machine Learning Engineer you will help accelerate the transformation by integrating advanced ML models into product experiences and shaping an AI-native productivity suite for enterprises. You will own core ML systems...
$350k
...Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to... ...Policy commitments. We're looking for an engineer to own the infrastructure behind that research. This is the tooling our researchers rely...Work at officeVisa sponsorshipFlexible hoursShift work$292k - $417.2k
...ads optimization that shape the future of streaming.We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction...Full timeTemporary workLocal areaFlexible hours- Radical Numerics in San Francisco is seeking a Member of Technical Staff, ML Product Engineer to develop APIs and systems that support genome-scale workloads. This critical role involves building reliable batch systems and ensuring efficient inference for scientific applications...
- Brain Co. is building AI-native operating systems for large, regulated institutions. As a Machine Learning Engineer on the Platform, you will develop core ML capabilities—document extraction, model routing, and a unified evaluation system—that support production-grade...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Infrastructure Engineer. Be the first to apply!
Related searches
- data scientist machine learning engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- graduate machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- data infrastructure engineer San Francisco, CA

