Staff ML Performance Engineer Scalable Inference & CUDA
Modal
A leading AI infrastructure company based in New York is seeking experienced engineers to enhance the performance of ML systems and contribute to open-source projects. Ideal candidates will have over 5 years of experience in writing high-quality code and familiarity with Nvidia GPU architecture and ML frameworks. This role offers opportunities for significant growth within a fast-growing team and requires in-person collaboration in NYC, San Francisco, or Stockholm. #J-18808-Ljbffr Modal
- ...jobs, and serve low-latency inference. We have thousands of... ...medalists, and experienced engineering and product leaders with decades... ...with experience in making ML systems performant at scale. If you are interested... ...GPU architecture and CUDA. Experience with ML performance...Performance
- ...About the Role As an ML Research Engineer at Maple, you'll be a... ...systems to monitor performance, detect anomalies,... ...optimized production inference. Lead evaluations... ...maintain robustness and scalability. Balance... ...optimization experience with CUDA/Triton preferred....PerformanceWork at officeLocal area
- Jobzhr, a leading global trading firm, is expanding its ML infrastructure in New York. We are hiring engineers to build distributed training and low-latency inference systems that move models from research into production. You will work closely with researchers and traders...Suggested
$200k
...seeking a Machine Learning Performance Engineer to join our team, focusing on... ...infrastructure, training, and inference challenges to advance our... ...responsibilities include:Building scalable and robust training and... ...-level GPU programming with CUDA, including Tensor Cores, cooperative...PerformanceWork at office$251k - $310k
...generative modeling, Bayesian inference, hierarchical learning... ...and World models to perform 3D Perception using... ...effectively with engineering and research teams across... ...Develop and maintain scalable data pipelines for Training... ...), and debugging of ML models. You have:...PerformanceFull timeTemporary workRemote work- BlackLine seeks a Senior AI/ML Engineer to design, build, and optimize large-scale data pipelines... ...accounting agents. You will lead scalable data infrastructure and collaborate... ...strong focus on security, reliability, and performance. #J-18808-Ljbffr Blackline Systems IncPerformance
- ...Machine Learning Engineer - Inference / Serving Join to apply for the Machine Learning Engineer... ...Today, we are focused on bringing the performance of closed‑web user acquisition to the open... ...CTV products. This is an applied ML systems role—equal parts engineering depth...PerformanceFull timeRemote work
- ...a team of researchers, engineers, designers, and more, who... ...fast, reliable, and scalable model training — and build... ...the full stack of ML systems, this role gives... ...configurations support high-performance training.Investigate... ...performance issues across CUDA/NCCL, networking, IO,...PerformanceFull timeWork at officeLocal areaRemote workHome office
- AI/ML Ops EngineerLocation: Remote / Hybrid (Client... ...an experienced AI/ML Engineer to design, deploy, and... ...learning models into scalable, governed, and... ...real-time APIs, batch inference pipelines, and feature... ...management platforms meet performance and quality requirements...PerformanceFull timeContract workLocal areaRemote workFlexible hours
- ...Information Job Title ML Staff Engineer - LLM & Production Systems... ...of ML platforms that enable scalable, reliable, and impactful... ...deployment strategies and optimize inference workloads through capacity... ..., annual discretionary performance bonus, 401(k) plan with an...PerformancePermanent employmentFull timeWork at officeLocal area1 day per week
- Machine Learning Engineer — AI Investment Research... ...and deploy production ML models end-to-end — from... ...Build and maintain scalable ML pipelines for... ...(A/B testing, causal inference) to validate model improvements... ...optimization, CUDA, or performance profiling. We value engineers...PerformanceFull time
- ...and deploy production‑grade ML systems with end‑to‑end... ...model training, deployment, inference, and monitoring in production... ...infrastructure and processes for scalability and performance. Qualifications Bachelor’... ...experience in ML engineering. Strong programming skills...PerformanceFull time
- ...help healthcare professionals perform at their best. At Solventum,... ....**Job Description:****ML Engineer****3M Health Care is now Solventum... ...AI services are secure and scalable.**Key Responsibilities****1.... ...for model training and inference.* **Feature Management:** Help...PerformanceH1bRemote work
$130k - $180k
...The Opportunity: We’re hiring an AI/ML Engineer to help build and scale the infrastructure... ...model workflows to deploying scalable inference systems. You’ll work closely with our... ...teams to continuously improve model performance and delivery quality. If you’re excited...PerformanceFull timeLocal area$189.6k - $237k
Scale’s ML platform (RLXF) team builds our internal... ...model training and inference. The platform has been... ...software engineering skills, proficient in... ...frameworks and tools such as CUDA, Pytorch, transformers... ...qualifications, interview performance, and relevant education...PerformanceFull time$175k - $280k
...layer, integrating LLM, speech, and vision models. The ideal candidate has significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving. Join a team dedicated to pioneering advancements in voice agents...Performance- ...Talent Network for Principal AI/ML Engineer Roles at VC-Backed Startups... ...in architecting scalable AI systems and leading technical... ...Develop scalable training and inference pipelines for AI-powered applications... ..., and optimization for performance ~ Strong knowledge of...PerformanceFull time
$200k - $300k
...Your Role As a founding AI Engineer at Overtone, you will... ...system behavior and model performance. Matching & ML systems: Apply data science... ...Model serving & production inference: Own the model serving layer... ...user-facing features with scalability and reliability in mind....PerformanceFull timeWork at office$300k - $400k
...Role Description As a Principal AI/ML Engineer in our AdTech team, you will be a key... ...to ensure our ML systems are highly performant, scalable, and reliable. You will also... ...ingestion and training to real-time inference , for our real-time bidding, targeting...PerformanceFull time$209k - $313k
...other digital services.Snap Engineering teams build fun and technically... ...complexity, bias/variance, scalability, and interpretabilityConduct... ...Strong understanding of causal inference and modern approaches to estimating... ...tests) and leveraging causal ML in production...Full timeLive inWork at officeLocal area- ...building the future of sports performance technology, with a mission... ...It is the highest-leverage engineering position in Phase 1 of a platform... ...and implementing impactful, scalable people solutions. Drive a... ...Experience with causal inference or counterfactual modeling over...Performance
$180k - $230k
...serious scale. We're hiring ML Engineer to build and own the infrastructure... ...quotes in seconds against inference-time datasets that run into... ...reliable, reproducible, and scalable as data volume and model... ...validation infrastructure so model performance can be measured quickly and...Performance$190k - $260k
...Description *Machine Learning Engineer – Search, Ranking &... ..., you will join the ML team to design, build,... ...items daily with high performance and reliability. -... ...accuracy and system scalability. - Contribute to... ...processing for real-time inference. - Strong backend integration...PerformanceRemote jobFull timeH1bRelocationVisa sponsorship$100k - $250k
...innovative field blends AI, engineering, and materials science,... ...The opportunity As an ML Engineer at Radical AI, you... ...levels of seniority: Senior, Staff, and Principal. \n Mission... ...and techniques to improve performance and scalability. Collaborate with researchers...PerformanceFull time- ...machine learning function and hiring engineers to develop the distributed training and inference systems that take models from... ...floor than a traditional ML function. What you'll do Build and... ...Deep PyTorch fluency and strong performance-engineering instincts Distributed...PerformanceFlexible hours
- ...+ years in systems or ML systems, with real depth... .../bringup, or high-performance kernels. You've taken... ...internals of modern LLM inference: transformers, attention... ...-level experience in CUDA, Triton, or a vendor kernel... ...tooling that did real engineering work, not demos. #J-1...PerformanceLive in
- VP - AI/ML Engineer - Compliance EngineeringYOUR IMPACTAre you passionate... ...you will:Design and architect scalable and reliable end-to-end AI/ML solutions... ...of scalability and performance optimization techniques for real-time inference such as quantization, pruning,...Performance
$235k - $260k
...organization.The Principal AI/ML Engineer, Semantic Data will design... ...prem GPU environmentsOptimize inference workflows for latency, cost,... ...& InfrastructureBuild scalable, production-grade services and... ...pipelinesEnsure systems meet performance and reliability requirementsGovernance...PerformanceWork at officeLocal areaRemote work1 day per week- ...Senior MLOps Engineer We are looking for... ...build and deploy ML models on a modern... ...clusters to support scalable machine learning workflows... ..., real-time inference as well as batch... ...ensuring optimal performance and reliability.... ...fundamentals: caching, CUDA, autoscaling, high...Performance
- ...for a Senior MLOps engineer to work closely with... ...build and deploy ML models on a modern... ...clusters to support scalable machine learning workflows... ...-time and batch inference systems, ensuring optimal performance and reliability.... ...fundamentals: caching, CUDA, autoscaling, high...Performance
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff ML Performance Engineer Scalable Inference & CUDA. Be the first to apply!
- software engineer staff New York, NY
- assistant engineer New York, NY
- engineering aide New York, NY
- staff engineer New York, NY
- staff security engineer New York, NY
- assistant engineering manager New York, NY
- senior staff systems engineer New York, NY
- technology administrator New York, NY
- project engineer assistant project manager New York, NY
- assistant civil engineer New York, NY



