Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer — Accelerate GPU Inference & Training

Jobleads-US

Aionia Group in Menlo Park, CA is hiring a Member of Technical Staff, ML Systems to accelerate model training and inference across image, video, and world-model workloads. You will work with a founding team on kernels, runtimes, and distributed engines that power production-scale ML stacks.

You’ll optimize GPU performance, profile bottlenecks with Nsight, and implement low-level CUDA and Triton improvements.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer — Accelerate GPU Inference & Training in Menlo Park, CA vacancy
  • $250k - $350k

     ...seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's...  ...edge inference acceleration, GPU parallelism, advanced...  ...ensuring our creative AI systems deliver industry-leading user...  ...models into production.Improve Training Efficiency: (Bonus)... 
    Training
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    5 days ago
  • $250k - $350k

     ...state-of-the-art models to accelerate breakthroughs across...  ...some of the world's leading ML systems engineers, including leaders behind...  ...powering our large-scale training, inference, and reinforcement learning...  ...environments Optimize memory, GPU kernels and communication... 
    Training
    Visa sponsorship

    Periodic Labs

    Menlo Park, CA
    5 days ago
  • $300k - $400k

     ...the-art models to accelerate breakthroughs...  ...You will own the systems layer that makes...  ...our frontier model training and inference fast, efficient,...  ...communication and GPU kernels to extract...  ...benchmarking distributed ML systems to...  ...the scientists, engineers, and problem-solvers... 
    Training
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    6 days ago
  •  ...take your software engineering career to the next...  ...hygiene and system architectureRequired...  ...and skillsFormal training or certification on...  ...with emphasis on ML systems.Hands-on experience...  ...or operating GPU workloads in Kubernetes...  ...Serving, Triton Inference Server)Familiarity... 
    Training

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  •  ...innovations that accelerate how teams work, discover...  ...Machine Learning Systems Engineer (P60) to lead...  ...Design and Build ML SystemsArchitect and...  ...scalable systems for training, fine-tuning, and...  ...model training, inference pipelines, or search...  ...computing, or GPU optimization.Familiarity... 
    Training
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    2 days ago
  • $220k - $350k

     ...Job Description ML Infrastructure Engineer Company: Dyna...  ...Engineer to own training infrastructure end...  ...a multi-cloud GPU fleet into a world...  ...low-latency inference pipelines for real...  ...hire Multimodal systems (video, audio,...  ...DeepSpeed or Accelerate; model serving optimization... 
    Training
    Full time
    H1b
    Work at office
    Visa sponsorship

    Transparent Search Group

    Redwood City, CA
    14 days ago
  • $188.5k - $282.7k

     ...Semantic AI Governance Engine, which is the first system designed to...  ....As an Applied ML Engineer on the...  ...curating data, training small models, serving...  ...Serving and Inference Infrastructure (...  ...through shared GPU pools, KV-cache-...  ...in Securing and Accelerating the World's AI TransformationRubrik... 
    Training
    Permanent employment

    Rubrik

    Palo Alto, CA
    3 days ago
  • $171.7k - $303.9k

     ...world!The Data Labeling Engineering team designs, builds,...  ...engineering, and AI/ML, defining the strategies...  ...that create reliable training data at scale. Our tools...  ...across teams and systems that unblock the next...  ...operational triage, etc) to accelerate understanding,... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    5 days ago
  •  ...Machine Learning Engineer Chicago, IL; New...  ...each month. Our systems process and move...  ...own our production ML systems and build...  ...systems, training/serving parity, retraining...  ...computing and GPU-accelerated workloads (e.g.,...  ...distributed training/inference), including... 
    Training
    Work at office
    Immediate start
    Remote work

    Attain

    Redwood City, CA
    1 day ago
  • Cerebras Systems builds the world's largest AI chip,...  ...deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud...  ...closely with hardware engineers to isolate and resolve...  ...failures quickly and accelerate debugging.Reproduce,... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  •  ...JPMorganChase is seeking a Software Engineer III to design and operate an end-to-end ML training platform on AWS and other clouds. You will run GPU workloads, optimize performance, and enable Gen AI workflows within a governed, secure environment. You will collaborate... 
    Training

    Jobleads-US

    Palo Alto, CA
    4 days ago
  •  ...for a Senior MLOps engineer to work closely...  ...build and deploy ML models on a modern...  ...model serving systems, hyper-parameter...  ...for distributed training on GPU-enabled clusters...  ...real-time and batch inference systems, ensuring...  ...quantize LLMs for accelerating inference on specific... 
    Training

    JP Morgan Chase

    Palo Alto, CA
    3 days ago
  • $180k

     ...mission is to create AI systems that can accurately...  ..., and focused on engineering excellence. This organization...  ...ROLE: As an ML Infrastructure...  ...building, and scaling GPU compute infrastructure, training frameworks, and...  ...data, training, and inference systems Collaborating... 
    Training
    Temporary work
    Work experience placement

    SpaceXAI

    Palo Alto, CA
    5 days ago
  • $182k - $242k

     ...technical expertise to accelerate breakthroughs...  ...of managing GPU infrastructure....  ...successfully training self-improving...  ...fraction of the total inference market, which...  ...for strong engineers with great...  ...experience with modern ML frameworks such...  ...model training systems Experience... 
    Training
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    3 days ago
  • $150k

     ...edge foundation model training, alongside world-class...  ...data scientists, and engineers, tackling the most fundamental...  ...The Role The GPU Kernel Engineer will...  ...at training and inference, and support the team...  ...new and cutting-edge systems. The ideal candidate will... 
    Training
    Full time
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    2 days ago
  • Cerebras Systems builds the world's largest...  ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale...  ...The RoleThe Core ML team develops novel...  ...Wafer-Scale Engine. Our work spans efficient...  ...-level assembly, accelerator programming, or a... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...a Senior MLOps engineer to work closely...  ...build and deploy ML models on a...  ...distributed model training on large...  ...distributed training on GPU-enabled...  ...throughput, real-time inference as well as...  ...pipelines to ensure system health,...  ...quantize LLMs for accelerating inference on specific... 
    Training

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  • $250k - $320k

    Staff Infrastructure Engineer We are partnered with a Stealth...  ...ll design and optimize the inference platform, GPU‑based training clusters, and data...  ...play a key role in scaling systems for both research and production...  ...experience in Software / ML Infrastructure Engineering... 
    Training
    Full time
    Immediate start

    Strativ Group

    Menlo Park, CA
    4 days ago
  •  ...and Conversational AI system that integrates...  ...ecosystem.About the AI & ML Platform TeamOur...  ...a Senior ML System Engineer on the AI & ML Platform’s Inference team, you will design...  ...level optimizations (GPU kernels, quantization...  ...reliablyBenchmark, fine-tune, and accelerate inference engines.... 
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    3 days ago
  •  ...life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)...  ...the central nervous system of SpaceX - we create...  ...throughout SpaceX to accelerate launch vehicle production...  ...support for training workloads.RESPONSIBILITIES...  ...including low-level GPU kernel work, quantization... 
    Training
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    3 days ago
  • $150k - $230k

     ..., recommendation systems, and adtech.Recognized...  ...Machine Learning Engineer to drive the post-training of our large...  ...on mid-to-large GPU clusters, applying...  ...engineering for ML. You can independently...  ...Hugging Face TRL/Accelerate, DeepSpeed or FSDP, and inference engines like vLLM... 
    Training
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    4 days ago
  • $173k - $253k

    Matterport - Senior ML Ops Engineer Job Description...  ...performance, optimize inference speed and resource utilization...  ...model architectures, training procedures, and...  ...optimization, hardware acceleration, and efficient AI....  ...with version control systems (e.g., Git) and agile... 
    Training
    Full time
    Work at office
    Work from home

    Matterport

    Sunnyvale, CA
    3 days ago
  • $119.8k - $234.7k

     ...and Infrastructure Engineering (SCHIE) is the...  ...Microsoft's Hardware Systems organization is...  ...combines custom accelerators, advanced networking...  ...-leading AI training and inference capabilities. The...  ...systems. Translate AI/ML workload...  ...:Experience with GPU, FPGA, TPU, or custom... 
    Training
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Mountain View, CA
    4 days ago
  • $180k - $280k

     ...you passionate about accelerating the future of autonomous...  ....As a Staff AI/ML Engineer within the Onboard Embodied...  ...onboard ML systems powering fully autonomous...  ...sophisticated neural networks trained from large-scale...  ...capable of real-time inference and robust autonomous... 
    Training
    Full time
    Local area
    Work from home
    Relocation
    Relocation package

    General Motors

    Sunnyvale, CA
    1 day ago
  • $197.5k - $272k

     ...trust Sonatus to accelerate this shift. Our...  ...Learning Engineer to join our seasoned...  ..., including system logs, traces, and...  ...the end-to-end ML pipeline—from data...  ...and model training to deployment on...  ...execution on CPU/GPU-bound targets...  ...(C++14/17 for inference).Deep proficiency... 
    Training
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    4 days ago
  • $153.2k - $234.1k

     ...hardware and battery systems to intuitive design,...  ...you passionate about accelerating the future of autonomous...  .... As a Senior ML Infra Engineer, you will work on the...  ...dataset generation, training, evaluation and iteration...  ...training across large GPU/CPU clusters or specialized... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...About the job ML Engineer Our Client Is a rapidly growing Tier 1 VC backed...  ...the evolving world of intelligent systems. Location : New York, NY Work...  ...acquisition, preprocessing, model training, deployment, inference, and monitoring in production environments... 
    Training
    Full time

    Catalyst Labs, LLC

    Menlo Park, CA
    1 day ago
  • $180.9k - $265.32k

     ...sustainability, design and engineering, ambition and...  ...driving systems. This role focuses...  ...on productizing ML models, ensuring...  ...Optimization Optimize inference pipelines using...  ...1/2, OpenCV, and GPU acceleration. Hands-on experience...  ..., education and training; certifications;... 
    Training
    Hourly pay
    Night shift

    Lucid Motors

    Newark, CA
    4 days ago
  • $150k - $230k

     ...About Clockwork Systems Clockwork.io – Software...  ...to increase GPU cluster utilization...  ...and veteran systems engineers who share a vision...  ...and performance acceleration that dynamically...  ...performance distributed GPU training. You'll work at...  ..., InfiniBand) ML framework or... 
    Training

    Clockwork.io

    Palo Alto, CA
    18 days ago
  • $295.25k - $345.04k

     ...Learning Infrastructure Engineer, you’ll build...  ...that powers ML systems across our organization...  ...boundaries of large-scale training and serving. Your...  ...to low-latency inference and production...  ...quantization, optimizing GPU utilization and...  ...ML platforms that accelerate model development... 
    Training
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer — Accelerate GPU Inference & Training. Be the first to apply!