Machine Learning Training Infrastructure Engineer
$1,000 per monthJobleads-US
You let people adapt models and run appropriately sized training jobs on resources they control, and you make those jobs reproducible, interruptible and accountable. Interruptible matters more than it sounds: this is somebody's own machine, and they get to want it back.
The work: Build training pipelines, checkpoints, optimizer-state handling, resource scheduling and distributed execution where the network supports it. Coordinate local inference with background training. Enforce dataset permissions and explicit limits before a job moves to remote compute.
What good looks like: In your first 90 days, deliver a recoverable local fine-tuning workflow and a documented boundary between supported local jobs and remote workloads.
Evidence we look for: Bring practical ML training systems experience and distributed-computing fundamentals. Understand memory accounting, numerical stability and the difference between fine-tuning and large-scale pretraining.
Required:
- Practical ML training systems experience plus distributed-computing fundamentals
- Memory accounting and numerical stability in practice
- You understand the difference between fine-tuning and large-scale training, and design for the first honestly
- You make runs reproducible, including the parts people usually leave out
Nice to have:
- Parameter-efficient fine-tuning methods
- Checkpointing and preemption handling
- Federated or on-device training
The exercise: Plan a training run that must survive a power interruption while preserving the ability to compare results with the original baseline.
How we work: We work together in the office, five days a week, and you can be based at any of our garages: Kirkland Garage (Kirkland, WA); UAE Garage (Dubai, Dubai). We hire across the United States, India and the UAE. We are remote-friendly around family: if you need to work from home some days to look after the people you love, we arrange that with you one person at a time, and we encourage people to use it rather than tough it out.
Pay: We publish what we pay. Indicative ranges by market and level are on our compensation page, and they are realistic going market rates rather than headline numbers. Wherever we hire we pay at least the local market rate, and for full-time roles our floor is a living wage, never the statutory minimum. Compensation is reviewed every year and on promotion.
Equity: Every full-time teammate gets stock options. Four-year vesting with a one-year cliff, sized to role, level and impact, and confirmed in writing at offer. High performers earn refresh grants.
Bonus and commission: An annual performance bonus tied to clear company and personal goals, indicatively 10 to 20 percent of base for non-sales roles. Customer-facing roles carry on-target earnings, typically a 50/50 split of base and variable, with uncapped commission and accelerators above quota.
Health and peace of mind: Medical, dental and vision for you and your family, plus life and disability cover. We have chosen the highest plan tier available to us rather than the cheapest one that clears the bar, because the point of this is that you never have to think about it. The specific plan numbers are confirmed in your offer letter.
401(k) with company matching: A 401(k) with a company match, so the years you spend here compound into something that is yours whatever happens next. The match formula is confirmed in your offer letter.
AI tokens: A budget of AI tokens of your own, because a company that says you should own your AI cannot be the company that rations it. Use them on the work and on whatever you are curious about.
Gym, and the everyday things: A gym membership, and a corporate benefits programme with its own app, where you redeem real discounts with a long list of retailers on the ordinary purchases of a life.
Family: We are remote-friendly around your family, arranged one person at a time, and we would rather you took it than toughed it out.
Referrals: Refer someone we hire full-time who stays a year and you get $1,000 plus $10,000 in referral stock-based equity, on top of your own package.
#J-18808-Ljbffr Jobleads-US- ...HushOne, Inc. seeks a senior engineer to build robust ML training systems, enabling distributed execution and reproducible experiments on local resources. You will ensure memory accounting, numerical stability, and a clear boundary between local and remote workloads....TrainingRemote jobLocal area
$172.5k - $260.1k
...Salesforce.As a Lead Network Engineer at Salesforce, you will be part of the Infrastructure Strategy and Data Center Operations... ...of the current state of AI, machine learning, LLMs, MCP, etc.A related... ...compensation, promotion, benefits, training, assessment of job performance...TrainingFull timeShift work$148.5k - $223.9k
...experienced and hands-on software engineer to join our team to build... ...the next generation of infrastructure management service and... ...operating systems.Apply AI and Machine Learning: Leverage AI/ML models for... ...compensation, promotion, benefits, training, assessment of job...TrainingFull time$143.7k - $194.4k
...and motivated software engineer to help us in our... ...builds and operates the infrastructure for collecting, ingesting... ...and RDS databases.To learn more about Amazon RDS... ...experience- Knowledge of Machine Learning and LLM... ...transformer architecture, training/inference lifecycles,...TrainingInternshipFlexible hours- ...Roles & Responsibilities An ML Engineer designs, builds, deploys, and maintains end to end machine learning solutions across data ingestion, model training, and production deployment... ...experience in creating secure ML infrastructure aligned with enterprise security...TrainingFull time
$20k
...enabling human life on Mars.NETWORK ENGINEER, AI INFRASTRUCTURE (STARSHIELD) Starshield leverages SpaceX... ...(100k+ GPU scale)Collaborate with ML training teams to translate workload... ...authorizations from the U.S. Department of State. Learn more about the ITAR here. SpaceX is...TrainingPermanent employmentTemporary workInternshipWork at officeImmediate startMonday to FridayWeekend work- ...Responsibilities Build and scale distributed training infrastructure for large AI models across GPU... ..., orchestration, and performance engineering teams. Diagnose issues affecting... ...training systems or large-scale machine learning infrastructure. Experience supporting...TrainingFull timeWork at officeRelocation3 days per week
- ...AI Training Infrastructure Engineer Location: Hybrid | Bellevue, WA (downtown) Titles: Senior and Staff (multiple roles available) Build... ...with infrastructure, orchestration, performance, and machine learning teams to solve complex challenges around distributed computing...TrainingWork at officeRelocation3 days per week
- ...Job Description Job Description Infrastructure Engineer, AI for Chip Design Full-time · On-site · San Jose, CA · Austin, TX or Taiwan... ...infrastructure — GPU clusters and scheduling, distributed training, model serving and inference, and model/environment management...TrainingFull time
- ...Cloud and Customer Solutions Engineer on AMD's Applied AI team,... ...things break at scale. What you learn in the field, you convert... ...You can debug a distributed training hang at 2am, explain the root... ...years of production software or infrastructure engineering, including...TrainingPermanent employmentFlexible hours
- Senior Machine Learning Engineer, Data InfrastructureUnity Vector builds an Data platform that powers... ...also supports large-scale model training, feature generation, and experimentation... ...role focuses on building reliable infrastructure for generating data infrastructure,...TrainingFull timeWork at officeWorldwide
$215k - $240k
Team Charter:The Cloud Engineering team is dedicated to maintaining and evolving cloud infrastructure, CI/CD pipelines, Application Security... ...a dynamic team to do it. To learn more about working at OfferUp... ...absence, compensation, and training.OfferUp expressly prohibits any...TrainingFull timeLocal areaRemote workFlexible hours$200k - $287.5k
...of how work gets done.Senior Software Engineer — Cortex TrainingThe Snowflake ML Platform... .../AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it... ...for an engineer who thrives in the ML infrastructure layer and brings a solid understanding...Training$160k - $225k
...human life on Mars.SR. SOFTWARE ENGINEER (PLATFORM TEAM) The Platform... ...tooling and security infrastructure that empowers every team at SpaceX... ..., deployed applications, and trained models on managed, reliable computeDevelop... ...U.S. Department of State. Learn more about the ITAR here....TrainingPermanent employmentTemporary work$112.8k - $153.7k
...listen to your ideas. This is a senior engineering role. You will architect, operate, and... ...Snowflake, Databricks) and own BI delivery infrastructure (e.g., Power BI, Tableau). You set the... ..., leaves of absence, compensation and training. Armanino expressly prohibits any form...TrainingFull timeContract workLocal areaFlexible hours$147k - $210k
...Software Engineer III, AI/ML, Platforms and Devices Share Software... ...the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML... ...experience, and relevant education or training. US: $147000 - $210000 (USD)...TrainingTemporary work- Software Engineer, AI Infrastructure About the Company AI is transforming digital industries, but progress in the physical sciences — including drug... ...organizations to securely share proprietary datasets, train models on high-performance GPU infrastructure, and deploy those...Training
$117.2k - $313.7k
...Salesforce.Distributed Systems Software Engineer - Public Cloud (Senior/Lead/Principal)... ...production-ready.Your Impact:Deliver cloud infrastructure automation tools, frameworks, workflows... ..., compensation, promotion, benefits, training, assessment of job performance,...TrainingFull time$1,000 per month
...to computers they own or are authorised to use, wherever those machines are. Much of this happens across home networks, which are hostile... ...Separate remote task dispatch from tightly coupled distributed training, and enforce limits on data leaving the user's device. What...TrainingRemote jobFull timeWork at officeLocal areaWork from home$140k - $200k
Water / Wastewater Infrastructure Senior Project EngineerWater and Environment Jobs with David... ...Practice is seeking a Project Engineer in Bellevue, Olympia, Tacoma, Seattle,... ...development:Support for continuing education and training opportunities.Work-life balance: Paid...TrainingWork at officeLocal areaFlexible hours- ...Amazon’s Artificial General Intelligence team in Bellevue, WA, seeks a Software Development Engineer to build scalable data platforms for evaluation and training of large language models and generative AI systems. You will partner with scientists and ML engineers to...Training
$157k - $213.8k
...the world's best data and AI infrastructure platform so our customers... ...experienced Senior Software Engineers with large-scale distributed... ...relevant certifications and training, and specific work location.... ...Lakehouse, and Unity Catalog. To learn more, follow Databricks on...TrainingLocal areaWorldwide$117.2k - $313.7k
...efforts. Job Category Software Engineering Job Details About Salesforce Salesforce... ...Your Impact: Deliver cloud infrastructure automation tools, frameworks, workflows... ..., compensation, promotion, benefits, training, assessment of job performance,...Training$224k - $356.5k
...seeking a Senior Performance Engineer to characterize workloads,... ...technically diverse team of infrastructure experts to unlock more efficient... ...plans.Partner with deep learning engineers, platform teams, and... ...AI clusters or distributed training and inference workloadsExperience...TrainingFull timeRemote work$217.2k - $288.4k
...the world's best data and AI infrastructure platform, so our customers... ...central to their missions.Our engineering teams build highly technical... ...relevant certifications and training, and specific work location.... ..., and Unity Catalog. To learn more, follow Databricks on LinkedIn...TrainingLocal areaWorldwide$184k - $287.5k
...looking for a Senior Software Engineer to lead the bring-up, triage... ...optimization of distributed training and inference workloads... ...at the intersection of deep learning systems, GPU performance, distributed... ...of large-scale AI clusters, infrastructure, and end-to-end workloads,...TrainingFull timeRemote work$95k - $120k
...enabling human life on Mars.DATA INFRASTRUCTURE OPERATIONS SPECIALIST II (... ...datasets that power the training, validation, and continuous... ...Align sourcing strategy with engineering, quality, and performance requirements... ...U.S. Department of State. Learn more about the ITAR here....TrainingPermanent employmentTemporary workWork at officeMonday to FridayFlexible hours$119.6k - $161.7k
...hardware and applying robotics, autonomy, supply chain optimization, machine learning, manipulation, image processing, and real-time data processing using distributed systems.Sub Same Day (SSD) Engineering is looking for a Senior Process Engineer with a strong delivery...Flexible hoursNight shift$182k - $242k
...enterprises, CoreWeave combines superior infrastructure performance with deep technical... ...(Nasdaq: CRWV) in March 2025. Learn more at What You'll Do: IT Engineering at CoreWeave designs, builds,... ...entitlement, data-retention and training opt-out settings, browser-...TrainingPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...OpportunityA well-funded, rapidly growing AI infrastructure company is building a next-generation... ..., including large-scale compute, model training, fine-tuning, inference, and emerging... ...of an established parent organization. Engineering teams are intentionally lean, highly...TrainingWork at officeRelocation3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Training Infrastructure Engineer. Be the first to apply!
- remote infrastructure engineer Kirkland, WA
- infrastructure engineer Kirkland, WA
- senior infrastructure engineer Kirkland, WA
- infrastructure developer Kirkland, WA
- artificial intelligence - machine learning intern Kirkland, WA
- machine learning Kirkland, WA
- machine learning research scientist Kirkland, WA
- machine learning scientist Kirkland, WA
- computer vision machine learning engineer
- ai ml developer





