Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Performance Engineer Scalable Inference & CUDA

Modal

A leading AI infrastructure company based in New York is seeking experienced engineers to enhance the performance of ML systems and contribute to open-source projects. Ideal candidates will have over 5 years of experience in writing high-quality code and familiarity with Nvidia GPU architecture and ML frameworks. This role offers opportunities for significant growth within a fast-growing team and requires in-person collaboration in NYC, San Francisco, or Stockholm. #J-18808-Ljbffr Modal

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff ML Performance Engineer Scalable Inference & CUDA in New York, NY vacancy
  •  ...jobs, and serve low-latency inference. We have thousands of...  ...medalists, and experienced engineering and product leaders with decades...  ...with experience in making ML systems performant at scale. If you are interested...  ...GPU architecture and CUDA. Experience with ML performance... 
    Performance

    Modal Labs

    New York, NY
    2 days ago
  •  ...About the Role As an ML Research Engineer at Maple, you'll be a...  ...systems to monitor performance, detect anomalies,...  ...optimized production inference. Lead evaluations...  ...maintain robustness and scalability. Balance...  ...optimization experience with CUDA/Triton preferred.... 
    Performance
    Work at office
    Local area

    Maple AI, Inc

    New York, NY
    3 days ago
  • Jobzhr, a leading global trading firm, is expanding its ML infrastructure in New York. We are hiring engineers to build distributed training and low-latency inference systems that move models from research into production. You will work closely with researchers and traders... 
    Suggested

    Jobzhr

    New York, NY
    2 days ago
  • $200k

     ...seeking a Machine Learning Performance Engineer to join our team, focusing on...  ...infrastructure, training, and inference challenges to advance our...  ...responsibilities include:Building scalable and robust training and...  ...-level GPU programming with CUDA, including Tensor Cores, cooperative... 
    Performance
    Work at office

    Optiver

    New York, NY
    2 days ago
  • $251k - $310k

     ...generative modeling, Bayesian inference, hierarchical learning...  ...and World models to perform 3D Perception using...  ...effectively with engineering and research teams across...  ...Develop and maintain scalable data pipelines for Training...  ...), and debugging of ML models. You have:... 
    Performance
    Full time
    Temporary work
    Remote work

    Waymo

    New York, NY
    14 hours ago
  • BlackLine seeks a Senior AI/ML Engineer to design, build, and optimize large-scale data pipelines...  ...accounting agents. You will lead scalable data infrastructure and collaborate...  ...strong focus on security, reliability, and performance. #J-18808-Ljbffr Blackline Systems Inc
    Performance

    Blackline Systems Inc

    New York, NY
    3 days ago
  •  ...Machine Learning Engineer - Inference / Serving Join to apply for the Machine Learning Engineer...  ...Today, we are focused on bringing the performance of closed‑web user acquisition to the open...  ...CTV products. This is an applied ML systems role—equal parts engineering depth... 
    Performance
    Full time
    Remote work

    Yobi AI

    New York, NY
    1 day ago
  •  ...a team of researchers, engineers, designers, and more, who...  ...fast, reliable, and scalable model training — and build...  ...the full stack of ML systems, this role gives...  ...configurations support high-performance training.Investigate...  ...performance issues across CUDA/NCCL, networking, IO,... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  • AI/ML Ops EngineerLocation: Remote / Hybrid (Client...  ...an experienced AI/ML Engineer to design, deploy, and...  ...learning models into scalable, governed, and...  ...real-time APIs, batch inference pipelines, and feature...  ...management platforms meet performance and quality requirements... 
    Performance
    Full time
    Contract work
    Local area
    Remote work
    Flexible hours

    Slalom

    New York, NY
    2 days ago
  •  ...Information Job Title ML Staff Engineer - LLM & Production Systems...  ...of ML platforms that enable scalable, reliable, and impactful...  ...deployment strategies and optimize inference workloads through capacity...  ..., annual discretionary performance bonus, 401(k) plan with an... 
    Performance
    Permanent employment
    Full time
    Work at office
    Local area
    1 day per week

    Bain & Company

    New York, NY
    3 days ago
  • Machine Learning Engineer — AI Investment Research...  ...and deploy production ML models end-to-end — from...  ...Build and maintain scalable ML pipelines for...  ...(A/B testing, causal inference) to validate model improvements...  ...optimization, CUDA, or performance profiling. We value engineers... 
    Performance
    Full time

    Riviera Partners

    New York, NY
    1 day ago
  •  ...and deploy production‑grade ML systems with end‑to‑end...  ...model training, deployment, inference, and monitoring in production...  ...infrastructure and processes for scalability and performance. Qualifications Bachelor’...  ...experience in ML engineering. Strong programming skills... 
    Performance
    Full time

    Catalyst Labs

    New York, NY
    1 day ago
  •  ...help healthcare professionals perform at their best. At Solventum,...  ....**Job Description:****ML Engineer****3M Health Care is now Solventum...  ...AI services are secure and scalable.**Key Responsibilities****1....  ...for model training and inference.* **Feature Management:** Help... 
    Performance
    H1b
    Remote work

    Solventum

    New York, NY
    14 hours ago
  • $130k - $180k

     ...The Opportunity: We’re hiring an AI/ML Engineer to help build and scale the infrastructure...  ...model workflows to deploying scalable inference systems. You’ll work closely with our...  ...teams to continuously improve model performance and delivery quality. If you’re excited... 
    Performance
    Full time
    Local area

    Savant Bio

    New York, NY
    14 hours ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal...  ...model training and inference. The platform has been...  ...software engineering skills, proficient in...  ...frameworks and tools such as CUDA, Pytorch, transformers...  ...qualifications, interview performance, and relevant education... 
    Performance
    Full time

    Scale AI

    New York, NY
    2 days ago
  • $175k - $280k

     ...layer, integrating LLM, speech, and vision models. The ideal candidate has significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving. Join a team dedicated to pioneering advancements in voice agents... 
    Performance

    SESAME

    New York, NY
    2 days ago
  •  ...Talent Network for Principal AI/ML Engineer Roles at VC-Backed Startups...  ...in architecting scalable AI systems and leading technical...  ...Develop scalable training and inference pipelines for AI-powered applications...  ..., and optimization for performance ~ Strong knowledge of... 
    Performance
    Full time

    Signal Fire Inc

    New York, NY
    14 hours ago
  • $200k - $300k

     ...Your Role As a founding AI Engineer at Overtone, you will...  ...system behavior and model performance. Matching & ML systems: Apply data science...  ...Model serving & production inference: Own the model serving layer...  ...user-facing features with scalability and reliability in mind.... 
    Performance
    Full time
    Work at office

    Overtone, Inc

    New York, NY
    14 hours ago
  • $300k - $400k

     ...Role Description   As a Principal AI/ML Engineer in our AdTech team, you will be a key...  ...to ensure our ML systems are highly performant, scalable, and reliable. You will also...  ...ingestion and training to real-time inference , for our real-time bidding, targeting... 
    Performance
    Full time

    Zeta Global

    New York, NY
    14 hours ago
  • $209k - $313k

     ...other digital services.Snap Engineering teams build fun and technically...  ...complexity, bias/variance, scalability, and interpretabilityConduct...  ...Strong understanding of causal inference and modern approaches to estimating...  ...tests) and leveraging causal ML in production... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    New York, NY
    14 hours ago
  •  ...building the future of sports performance technology, with a mission...  ...It is the highest-leverage engineering position in Phase 1 of a platform...  ...and implementing impactful, scalable people solutions. Drive a...  ...Experience with causal inference or counterfactual modeling over... 
    Performance

    PVH (Tommy Hilfiger/Calvin Klein)

    New York, NY
    1 day ago
  • $180k - $230k

     ...serious scale. We're hiring ML Engineer to build and own the infrastructure...  ...quotes in seconds against inference-time datasets that run into...  ...reliable, reproducible, and scalable as data volume and model...  ...validation infrastructure so model performance can be measured quickly and... 
    Performance

    Arlo Corporation

    New York, NY
    14 hours ago
  • $190k - $260k

     ...Description *Machine Learning Engineer – Search, Ranking &...  ..., you will join the ML team to design, build,...  ...items daily with high performance and reliability. -...  ...accuracy and system scalability. - Contribute to...  ...processing for real-time inference. - Strong backend integration... 
    Performance
    Remote job
    Full time
    H1b
    Relocation
    Visa sponsorship

    Fuku

    New York, NY
    14 hours ago
  • $100k - $250k

     ...innovative field blends AI, engineering, and materials science,...  ...The opportunity As an ML Engineer at Radical AI, you...  ...levels of seniority: Senior, Staff, and Principal. \n Mission...  ...and techniques to improve performance and scalability. Collaborate with researchers... 
    Performance
    Full time

    Radical Ai

    New York, NY
    14 hours ago
  •  ...machine learning function and hiring engineers to develop the distributed training and inference systems that take models from...  ...floor than a traditional ML function. What you'll do Build and...  ...Deep PyTorch fluency and strong performance-engineering instincts Distributed... 
    Performance
    Flexible hours

    Jobzhr

    New York, NY
    2 days ago
  •  ...+ years in systems or ML systems, with real depth...  .../bringup, or high-performance kernels. You've taken...  ...internals of modern LLM inference: transformers, attention...  ...-level experience in CUDA, Triton, or a vendor kernel...  ...tooling that did real engineering work, not demos. #J-1... 
    Performance
    Live in

    General Compute Inc.

    New York, NY
    1 day ago
  • VP - AI/ML Engineer - Compliance EngineeringYOUR IMPACTAre you passionate...  ...you will:Design and architect scalable and reliable end-to-end AI/ML solutions...  ...of scalability and performance optimization techniques for real-time inference such as quantization, pruning,... 
    Performance

    Goldman Sachs

    New York, NY
    2 days ago
  • $235k - $260k

     ...organization.The Principal AI/ML Engineer, Semantic Data will design...  ...prem GPU environmentsOptimize inference workflows for latency, cost,...  ...& InfrastructureBuild scalable, production-grade services and...  ...pipelinesEnsure systems meet performance and reliability requirementsGovernance... 
    Performance
    Work at office
    Local area
    Remote work
    1 day per week

    Major League Soccer

    New York, NY
    4 days ago
  •  ...Senior MLOps Engineer We are looking for...  ...build and deploy ML models on a modern...  ...clusters to support scalable machine learning workflows...  ..., real-time inference as well as batch...  ...ensuring optimal performance and reliability....  ...fundamentals: caching, CUDA, autoscaling, high... 
    Performance

    Chase

    New York, NY
    1 day ago
  •  ...for a Senior MLOps engineer to work closely with...  ...build and deploy ML models on a modern...  ...clusters to support scalable machine learning workflows...  ...-time and batch inference systems, ensuring optimal performance and reliability....  ...fundamentals: caching, CUDA, autoscaling, high... 
    Performance

    J.P. Morgan

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Performance Engineer Scalable Inference & CUDA. Be the first to apply!