ML Systems Engineer
Basis
Basis is a nonprofit applied AI research organization with two mutually reinforcing goals. The first is to understand and build intelligence. This entails establishing the mathematical principles of reasoning, learning, decision-making, understanding, and explaining, and constructing software that embodies these principles. The second is to advance society’s ability to solve intractable problems. This involves expanding the scale, complexity, and breadth of problems we can solve today and, more importantly, accelerating our ability to solve problems in the future. To achieve these goals, we are building both a new technological foundation inspired by human reasoning, and a new type of collaborative organization that prioritizes human value. About the Role ML Systems Engineers at Basis ensure training and evaluation infrastructure is fast, reliable, and scalable. You will own the full stack from distributed training frameworks through cloud administration, making it possible for researchers to iterate quickly on complex models while managing computational resources efficiently. We are looking for engineers who combine deep understanding of ML systems with operational excellence. The ideal ML Systems Engineer has experience with distributed training at scale, understands the intricacies of debugging numerical instabilities, and can manage cloud infrastructure that scales from experiments to production. You will be the guardian of training stability, the optimizer of compute costs, and the enabler of reproducible research. This role spans traditional ML engineering and cloud/DevOps responsibilities. You will manage GPU clusters, optimize cloud spending, ensure security and compliance, and build the infrastructure that lets researchers focus on algorithms rather than operations. We seek individuals who aspire to build robust ML infrastructure, maintain “logbook culture” for documenting issues and solutions, and treat operational excellence as a first-class concern. Have demonstrated expertise in ML systems engineering . Examples include: Managing distributed training jobs across hundreds of GPUs Debugging and fixing numerical instabilities in large-scale training Building infrastructure for reproducible ML experiments Optimizing training throughput and resource utilization Possess deep knowledge of distributed training frameworks including PyTorch/JAX distributed strategies (DDP, FSDP, ZeRO), gradient accumulation, mixed precision training, and checkpoint/recovery systems. Have strong cloud administration skills including AWS/GCP/Azure services, infrastructure as code (Terraform), Kubernetes orchestration, cost optimization, security best practices, and compliance requirements. Understand the full ML stack from hardware (GPUs, interconnects, storage) through frameworks (PyTorch, JAX) to high-level training loops and evaluation pipelines. Be skilled at debugging complex failures across the stack—GPU/NCCL issues, data loading bottlenecks, memory leaks, gradient explosions, and convergence problems. Value documentation and knowledge sharing . You maintain comprehensive logs of issues encountered, solutions found, and lessons learned, building institutional knowledge. Progress with autonomy while coordinating closely with researchers. You can anticipate infrastructure needs, prevent problems before they occur, and respond quickly when issues arise. In addition, the following would be an advantage: Experience at organizations training large models (OpenAI, Anthropic, Google, Meta). Background in both ML research and production systems. Contributions to ML frameworks or distributed training libraries. Experience with on-premise GPU cluster management. Knowledge of optimization theory and numerical methods. Understanding of robotics-specific infrastructure requirements. Responsibilities Own distributed training infrastructure including job launchers, checkpointing systems, recovery mechanisms, and monitoring that ensures experiments run reliably at scale. Debug and resolve training failures by diagnosing issues across GPUs, networking, numerics, and data pipelines, maintaining detailed logs of problems and solutions. Profile and optimize training performance by identifying bottlenecks in data loading, gradient computation, communication overhead, and implementing solutions that improve step time. Manage cloud infrastructure and costs including capacity planning, spot instance strategies, storage optimization, and building tools that give researchers visibility into resource usage. Implement security and compliance measures including access controls, data encryption, audit logging, and ensuring infrastructure meets requirements for handling sensitive data. Build evaluation and benchmarking infrastructure that enables consistent, reproducible measurement of model performance across different conditions and datasets. Develop monitoring and alerting systems that detect anomalies in training metrics, resource utilization, or system health, enabling rapid response to issues. Maintain development environments including containerization, dependency management, and tools that ensure researchers can reproduce results across different systems. Document and share knowledge through runbooks, post-mortems, and training materials that help the team understand and operate ML infrastructure effectively. Collaborate with researchers to understand requirements, suggest infrastructure solutions, and ensure systems support rather than constrain research goals. FT/PT: Full-time In-person Policy: We are in the office four days a week. Be prepared to attend multi-day Basis-wide in-person events. Location: New York City or Cambridge, MA. Salary range: Competitive salary. Non-Discrimination Notice Basis Research Institute provides equal employment opportunities without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, or genetics and prohibits discrimination based on all protected characteristics. Privacy Notice By submitting your application, you grant Basis permission to use your materials for both hiring evaluation and recruitment-related research and development purposes. Your information may be processed in different countries, including the US. You retain copyright while providing Basis a license to use these materials for the stated purposes. Read our full Global Data Privacy Notice here . #J-18808-Ljbffr
$110 per hour
...Role Responsibilities Guide research and engineering teams to close knowledge gaps and improve... ...MLOps , training infrastructure, and ML framework-level topics . Design... ...-structured solutions to MLOps and ML systems problems . Evaluate MLOps tasks and...SuggestedRemote jobContract workSummer workWeekday work$250k - $350k
...Applied ML Systems Engineer - Finance - NEW YORK - UNITED STATES Salary: $250,000 - $350,000 base + guaranteed Year 1 bonus (could be 100% to 500%) + sign-on. Total comp is elite tier. Primary Skills - pytorch, programming...SuggestedPermanent employmentFull timeWork experience placementInternshipImmediate startRemote workRelocationRelocation package$264.8k - $331k
...enterprises around the world. The Enterprise ML Research Lab works on the front lines of... ...clients. As an ML Sys Research Engineer, you'll work on building out the algorithms... ...-the-art technologies to optimize our ML system. Your customer will be other MLREs and AAIs...SuggestedFull time$212.5k - $270k
...navigate complex environments safely and efficiently. The system architecture team handles the onboard contract of the model with... ...You will: Tackle challenging real-world problems with ML and engineering solutions. Use state of the art techniques to design and...SuggestedFull timeContract workInternshipRemote work$100k - $250k
...science. This innovative field blends AI, engineering, and materials science, revolutionizing... ...sustainability. The opportunity As an ML Engineer at Radical AI, you will be responsible for developing machine learning systems, data pipelines, and large-scale training...SuggestedFull time$150k - $300k
...construct our built environment. Backed by Lightspeed Venture Partners, Eagle acquires and transforms civil, structural, and MEP engineering firms with applied AI. We’re an AI laboratory dedicated to providing engineers with the tools they need to solve the world’s...Full timeWork at office$281.2k - $401.71k
...features into a unified, agent-powered platform. Generative AI is reshaping how we build products and systems. As part of this shift, we’re creating a shared Agent Engine that powers agent-based experiences across Spotify. You’ll join a cross-functional group of...Remote jobFull timeFlexible hoursShift work- ...Hyperscience AI, building a workflow orchestration layer for computer vision models. Job Description We are seeking a Staff ML Engineer with a passion for building application-layer AI for high consequence, complex user interactions. In this role, you will:...Full timeWork at officeWork from homeFlexible hours3 days per week
- ...Overview We are seeking a ML Engineer who will assist in constructing a system to validate data and ML models. The ideal candidate should be intelligent, contemplative, and composed. They should not rush through tasks, instead diving deeply into the code and being attentive...Remote work
- ...audience science into scalable production systems that help advertisers decide whom to... ...partners. You bridge data science and platform engineering—converting prototypes into production-... ...fusion) into reproducible, production-quality ML services. ~Build and operate large-...Full timeWorldwide
$140k - $155k
...intelligence platform includes robust sensing systems across imaging modes (cameras, X-rays,... ...intelligence, and software-driven Field Engineering that drives real transformation on-site... ...across customer sites Work closely with ML engineers to accelerate the feedback loop...Work experience placementFlexible hours- ...full-time on-site role for a Founding Machine Learning Engineer located in NYC. The Founding ML Engineer will be responsible for the design and implementation... ...Develop models for generating and optimizing optical system layouts based on design goals and constraints. Build...Full time
- ...trajectory in the evolving world of intelligent systems. Location : New York, NY Work type :... ..., build, and deploy production‑grade ML systems with end‑to‑end ownership of the... ...6 years of professional experience in ML engineering. Strong programming skills in Python (TypeScript...Full time
$150k - $300k
...Hudson River Trading (HRT) is looking for GPU Systems Engineers to help scale and evolve our exceptionally sophisticated HPC/AI research environment. Joining our Research and Development team, you will collaborate with experts responsible for the compute, storage...Full timeWork at officeLocal areaImmediate start$180k - $230k
...around it: middlemen taking a cut, fraud nobody stops, and billing systems designed to fight over payment instead of deliver care. The... ...it runs on machine learning at serious scale. We're hiring ML Engineer to build and own the infrastructure that powers it — from training...Full time- We are looking for an engineer with experience in low-level systems programming and optimization to join our growing ML team. Machine learning is a critical pillar of Jane Street's global business. Our ever-evolving trading environment serves as a unique, rapid-feedback...Full time
- ...Electronics in 2019, 2022, and 2023. We are looking to bring on an experienced Data Engineer to join Eight Sleep’s machine learning team. This role involves monitoring production ML systems, building tools and pipelines for data analysis, dataset curation, training and...Full timeSleeping nightsFlexible hoursNight shift
- ...investors. I'll share more once we meet. About the Role As an ML Research Engineer at Maple, you'll be a part of our core product team... ...MIT, Columbia, and IBM, rapidly deploying advanced models and systems that directly impact small businesses. We work in person,...Work at officeLocal area
$180k - $250k
...exceptionally strong team includes software engineers, AI researchers, security engineers, and... ...– and we’re looking for exceptional AI/ML Engineers to help shape and build it. You... ...some of the world’s biggest distributed systems. A core part of this effort is using...Full time$240k - $270k
...smarter decisions—fast. That’s where you come in. As an AI/ML Engineer, you’ll join a growing team focused on building the AI foundation... ...-impact AI/ML opportunities Prototype and productionize AI systems that feel intuitive but do a lot under the hood—...Full timeWork at officeFlexible hours- ...funding, built our team around a strong engineering culture, and are launching publicly after... ...Engineer, you’ll design and build the core systems that power Merciv’s agentic... ...architecture to deployment, working closely with ML engineers, product, and our enterprise customers...Full timeImmediate startHome officeFlexible hours
- ...problem-solving, and obsessive attention to detail. Our engineering team is elite, our lawyers ship code every day, and... ...Sandstone is building the AI-native operating system for in-house legal teams. As an AI & ML Engineer, you will build the systems that make that...Full timeContract workWork at officeFlexible hoursDay shift
- ...security, industrial infrastructure, and enterprise AI. The Role As AI/ML Engineer, you will build and own the intelligence layer at the core of our product: models, pipelines, and agent systems that allow Linero to reason over complex, sensitive enterprise data,...Permanent employmentFull timeImmediate startRemote workRelocationRelocation package
- ...$15M Series A led by Footwork Ventures and Y Combinator to accelerate our momentum. As a Senior AI / ML Engineer , you will help build the core AI systems powering the Confido platform—from LLM-powered document understanding to predictive models that help brands...Full timeLocal areaRelocation packageNight shift
- ...organization's entire data landscape — internal systems, social signals, industry reports,... ...What You'll Do Design, build, and deploy ML models for demand forecasting, time... ...ML pipelines: data preprocessing, feature engineering, model training, evaluation, and production...Full timeImmediate start
$130k - $180k
...foundation for innovation. The Opportunity: We’re hiring an AI/ML Engineer to help build and scale the infrastructure behind Savant’s... ...prototyping model workflows to deploying scalable inference systems. You’ll work closely with our product, clinical, and engineering...Full timeLocal area$140k - $200k
...founders through one of the most consequential moments in a company’s life. Job Overview SimpleClosure is seeking an AI/ML Engineer to turn Asset Hub’s one-of-a-kind inventory, and existing AI buyer relationships, into high-value AI-training products. When...Full timeFor contractorsWork at officeLocal areaImmediate startFlexible hours2 days per week- ...machine learning, generative AI, agent-based systems, and graph technologies to get our... ...the development and deployment of complex ML systems reporting to Dan Wald, Co-Founder... ...both a Data Scientist and Machine Learning Engineer, you will play a pivotal role in designing...Full timeWork at officeRemote work
$170k - $212k
...connections with their fans. We’re looking for a Machine Learning Engineer to help us build systems that more accurately understand the performance that... ...it’s a DIY artist or an industry-facing partner. As an ML Engineer, you will help execute on strategies for...Remote jobFull timeFlexible hours$190k - $260k
*Machine Learning Engineer – Search, Ranking & Personalization* *Stage:* Seed *Founded... ..., personalized search and ranking systems to help users discover and trust products... ...Engineer at Client's company, you will join the ML team to design, build, and scale machine...Full timeH1bRemote workRelocationVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Systems Engineer. Be the first to apply!
- computer vision machine learning engineer New York, NY
- machine learning software engineer New York, NY
- machine learning ai engineer New York, NY
- machine learning engineer New York, NY
- graduate machine learning engineer New York, NY
- ai ml engineer New York, NY
- senior ml engineer New York, NY
- operating system engineer New York, NY
- wireless systems engineer New York, NY
- space systems engineer New York, NY


