Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Engineer - AWS Trainium & SageMaker [Remote]

Full-time

jobgether

United States
  • Remote job

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff ML Engineer – AWS Trainium & SageMaker based in United States.

As a Staff ML Engineer, you will design, train, optimize, and operate production machine learning workloads on AWS Trainium and Amazon SageMaker.
You will work deeply across the ML stack, from PyTorch training code and distributed execution to accelerator hardware and compiler behavior.
The role requires understanding how training workloads behave on custom silicon rather than treating infrastructure as a black box.
You will diagnose complex training issues, optimize throughput and cost, and build reliable pipelines for real production workloads.
You will work directly with engineering teams to translate business and technical requirements into scalable training solutions.
The environment is highly hands-on, technical, and client-facing, with a focus on solving specialized problems that require deep engineering expertise.
This is an opportunity to work at the intersection of machine learning, cloud infrastructure, distributed systems, and custom AI acceleration.

Accountabilities:

  • Train and operate machine learning models using Amazon SageMaker with AWS Trainium as the underlying compute infrastructure.
  • Develop and optimize PyTorch training workloads for execution on AWS Trainium.
  • Analyze how PyTorch code compiles and executes across Trainium and NeuronCore architecture.
  • Optimize memory utilization, throughput, and other performance characteristics of training workloads.
  • Diagnose training failures and performance issues caused by hardware, compiler behavior, device configuration, data, or model code.
  • Distinguish model- and data-level problems from accelerator- and compiler-level issues during troubleshooting.
  • Translate high-level requirements for Trainium workloads into complete, production-ready training pipelines.
  • Design training workflows that balance performance, scalability, reliability, and cloud infrastructure costs.
  • Tune distributed and multi-device training workloads for throughput and cost efficiency.
  • Operate production training workloads and help ensure their reliability throughout the ML lifecycle.
  • Work directly with client engineering teams to scope, design, and deliver specialized machine learning workloads.
  • Collaborate with internal engineering teams to develop production-grade solutions rather than isolated prototypes or notebook experiments.
  • Investigate complex technical issues across the ML software and hardware stack.
  • Apply a deep understanding of accelerator behavior to improve training architecture and implementation decisions.
  • Contribute to production engineering practices around deployment, monitoring, troubleshooting, and operational reliability.
  • Help translate emerging AI infrastructure capabilities into practical production solutions.
  • Communicate technical findings and tradeoffs clearly with both technical stakeholders and client teams.

Requirements:

  • Strong hands-on experience with PyTorch, including a deep understanding of training workflows.
  • Experience with distributed or multi-device model training is highly desirable.
  • Production experience using Amazon SageMaker for model training and/or inference.
  • Strong understanding of how machine learning workloads interact with accelerator hardware and device-specific compilation.
  • Ability to debug issues that originate at the hardware, accelerator, compiler, or runtime layer rather than solely within model or data code.
  • Experience optimizing training workloads for performance, throughput, memory utilization, or cost.
  • Strong Python programming fundamentals.
  • Understanding of distributed training concepts and production ML infrastructure.
  • Ability to design and operate end-to-end managed training pipelines.
  • AWS experience and familiarity with cloud-based machine learning infrastructure.
  • AWS Trainium or AWS Inferentia experience and familiarity with the AWS Neuron SDK is strongly preferred.
  • Candidates without direct Trainium or Inferentia experience may also be considered if they have deep PyTorch expertise and demonstrated ability to quickly learn new hardware targets.
  • Strong analytical and problem-solving skills, particularly when diagnosing complex system-level issues.
  • Ability to reason across multiple layers of the technology stack, from model code through frameworks, compilers, accelerators, and cloud infrastructure.
  • Experience working in a production engineering environment with high standards for reliability and delivery.
  • Strong communication and collaboration skills for working directly with client and internal engineering teams.
  • Comfortable operating in ambiguous environments where requirements and technical challenges may evolve.
  • Willingness and ability to travel approximately 20% within the United States.
  • Must be legally authorized to work in the United States.

Benefits:

  • Opportunity to work on production AI systems using AWS Trainium and Amazon SageMaker.
  • Exposure to custom AI accelerator hardware, compiler behavior, distributed training, and advanced ML infrastructure.
  • Hands-on work with complex machine learning workloads for enterprise clients.
  • Direct collaboration with client and internal engineering teams.
  • Opportunity to solve specialized technical problems that extend beyond conventional ML application development.
  • Production-focused engineering environment emphasizing systems that ship and operate at scale.
  • Approximately 20% U.S.-based travel associated with the role.
  • Final interview and onboarding may require onsite participation.
  • Professional growth through work across machine learning, cloud infrastructure, hardware acceleration, and production engineering.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 21 hours ago
Similar jobs that could be interesting for youBased on the Staff ML Engineer - AWS Trainium & SageMaker [Remote] in United States vacancy
  •  ...Position Overview We're seeking a Staff ML Engineer to advance our voice AI and clinical NLP...  ...infrastructure using Triton Inference Server on AWS GovCloud Create specialty-specific...  ...with cloud ML platforms (AWS SageMaker, GCP Vertex AI) ~ Master's or PhD in... 
    Amazon Web Service
    Work at office
    3 days per week

    SupportFinity

    San Francisco, CA
    1 day ago
  • $190k - $234k

     ...Staff ML Engineer Palo Alto, CA About Typeface: We help the world's biggest brands move from brief to fully personalized campaigns...  ...prototyping ~ Prior experience with AzureML, GCP Vertex, AWS Sagemaker or similar platforms ~ Bonus: Hands-on experience in deploying... 
    Amazon Web Service
    Work at office
    Local area
    Flexible hours
    3 days per week

    Typeface

    Palo Alto, CA
    23 hours ago
  •  ...You’ll Make an Impact As a Staff Machine Learning Engineer , you will play a key role...  ...scalability, and high quality of our ML systems from development to...  ..., distributed ML systems (Sagemaker) - Deep MLOps Proficiency:...  ...ML models at scale on AWS, with experience using... 
    Amazon Web Service
    Full time
    Remote work

    Cloudbeds

    Remote
    2 days ago
  •  ...Description: We are seeking a talented and experienced Deep Learning Engineer with a specialization in Large Language Models (LLMs) to join...  ...techniques and experience with cloud computing platforms (e.g., AWS, Azure, Google Cloud) • Excellent problem-solving skills and... 
    Amazon Web Service

    Xforia Inc

    San Jose, CA
    3 days ago
  • $300k

     ...term relationship. We're doubling down on ML as the future of Grindr, and in these...  ...leadership across teams, collaborating with engineering, data science and product teams to turn bold...  ...with any public cloud environment - AWS, GCP or Databricks. Expertise in... 
    Amazon Web Service
    Casual work
    Work at office
    Immediate start
    Worldwide
    Flexible hours

    Grindr

    Palo Alto, CA
    1 day ago
  • $220k - $280k

     ...industry together? As a Staff Machine Learning Engineer, you will lead the...  ...integrating robust, low-latency ML models across our sports...  ..., Kubeflow, Databricks, or SageMaker. ~ Strong Coding Skills:...  ...Functions, GKE, Vertex AI) or AWS equivalents. What... 
    Amazon Web Service
    Full time
    Remote work
    Work visa
    Flexible hours

    PrizePicks

    United States
    2 days ago
  •  ...terms of business impact. Own and prioritize the ML roadmap with product, engineering, and risk operations. Mentor ML engineers and...  ...routing experience is preferred. Familiarity with AWS ML tools such as SageMaker, feature stores, training and inference pipelines,... 
    Amazon Web Service
    Full time
    Flexible hours

    Payabli

    Remote
    13 days ago
  •  ...Staff Machine Learning Engineer In Brief We’re an early-stage startup on...  ...satisfied with training and tuning ML models that predict...  ...data as it moves through our AWS services, or improving the...  ...using MLOps tools such as SageMaker and MLFlow. Preferred... 
    Amazon Web Service
    Full time
    Local area
    Remote work

    Bayesian Health

    Remote
    15 days ago
  • $161k - $221.5k

     ...and technology solutions. About the Role We are seeking a Staff ML/AI Engineer to define and drive the architectural vision for Lyra’s machine...  ...: Strong experience architecting cloud-native solutions on AWS (or equivalent cloud providers). ~Strategic Communication: Exceptional... 
    Amazon Web Service
    Full time

    Lyra Health

    United States
    a month ago
  •  ...tech, and hyper-connected world. Role Overview As our Staff Software Engineer, ML infra Engineer for Search & Discovery organization, you...  ...integrating applications and platforms with cloud technologies (i.e, AWS and GCP) Strong verbal and written communication skills... 
    Amazon Web Service
    Temporary work
    Flexible hours

    Coupang

    Mountain View, CA
    2 days ago
  • $227.33k - $312.58k

     ...Job Description Job Description We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier Models organization. In this role...  ..., data quality and lineage tools Cloud & Infrastructure: AWS or GCP, containerized data workloads, CI/CD, infrastructure‑... 
    Amazon Web Service
    Work at office
    Local area
    Immediate start
    3 days per week

    Procore

    San Francisco, CA
    13 days ago
  • $298k - $351k

     ...are seeking an exceptional Staff Machine Learning Operations Engineer to join our Platform...  ...efficiency of Garner's production ML systems, including...  ...stack: model serving (e.g., Sagemaker, Triton, or equivalent), feature...  ...containerization, cloud (AWS preferred), Terraform/IaC,... 
    Amazon Web Service
    Work at office
    Work visa
    Flexible hours
    3 days per week

    Garner Health

    New York, NY
    8 days ago
  •  ...Job Description Are you a passionate Machine Learning Engineer with a strong background in SageMaker, prompt engineering, and LLM (Large Language Model)...  ...learning community. - Experience with cloud services (AWS, Azure, Google Cloud) and containerization technologies... 
    Amazon Web Service
    Remote work

    Maxiom Technology

    Ashburn, VA
    23 days ago
  •  ...full-time role partners with a fast-paced core engineering team to design, deploy, and optimize production-ready AI/ML solutions. Key Responsibilities Design, build...  ...centered on large language models using Python and AWS infrastructure. Architect scalable, high-... 
    Amazon Web Service
    Full time
    Remote work

    SaidGig

    United States
    a month ago
  •  ...Role We are seeking an exceptional Staff Machine Learning Engineer to lead the design and development of...  .... You will architect large-scale ML systems that detect and prevent fraud...  ...~ Familiarity with cloud platforms (AWS, GCP, or Azure) and containerized deployments... 
    Amazon Web Service

    AppGate

    New York, NY
    4 days ago
  •  ...Staff Machine Learning Engineer webAI is the first end-to-end private AI platform. Enterprises and Governments...  ...cross-functional teams to integrate ML models into our platform. Conduct...  ...with cloud computing services (AWS, Azure, GCP). Knowledge of Big Data... 
    Amazon Web Service
    Live out
    Work at office
    Local area
    Flexible hours

    webAI

    Austin, TX
    2 days ago
  •  ...Staff Machine Learning Engineer - Generative AI Lexington, KY Xometry powers the industries of today...  ...datasets Deploy in the cloud – Use AWS and other platforms to train, optimize...  ...and grow – Guide teammates on advanced ML methods, model architecture, and best... 
    Amazon Web Service
    Immediate start

    Xometry

    Lexington, KY
    3 days ago
  • $200k - $220k

     ...global manufacturing capacity. We are looking for a Staff Machine Learning Engineer to join our growing AI/ML team. This is a senior individual contributor role...  ...-time ML products at scale in cloud environments (AWS strongly preferred), including auto-scaling,... 
    Amazon Web Service

    Xometry

    Silver Spring, MD
    23 hours ago
  • $307k - $352k

     ...About the role: We are seeking a Staff Machine Learning Engineer to set the technical direction for machine learning at Kikoff. ML sits at the center of our business: our underwriting...  ...experience with cloud infrastructure (AWS or GCP), containerization (Docker,... 
    Amazon Web Service
    Local area
    Immediate start

    Kikoff

    San Francisco, CA
    2 days ago
  • $200k - $275k

     ...Our client, a growing FinTech company, are hiring a Staff Machine Learning Engineer to join their team in Colorado. The successful candidate will play...  ...systems and deploying cloud-native applications across AWS, Azure or Google Cloud Platform (GCP). Strong understanding... 
    Amazon Web Service
    Full time

    Alldus International Consulting Ltd

    Colorado
    more than 2 months ago
  •  ...our talented Team. Job Title: MLOps Platform Engineer (SageMaker) Location(s): Onsite Job Summary: This...  ...extensive experience in cloud infrastructure or ML platform operations, with a specific focus on AWS and Amazon SageMaker. The role involves designing... 
    Amazon Web Service

    Ampcus

    Plano, TX
    1 day ago
  •  ...Employment Type: Full-time Department: Engineering & Product As a Staff Machine Learning Engineer: You...  ...where refined data or internal ML/AI models can improve our product outcomes...  ...-driven architecture using Kafka, AWS (EKS), Python, Django/FastAPI, and Postgres... 
    Amazon Web Service
    Full time
    Work at office

    Steadily

    Austin, TX
    26 days ago
  • $292.5k - $409.5k

     ...teams. What You’ll Do: As a Senior Staff Software Engineer, you will help define and lead the...  ...Might Be: ~10+ years of experience in ML Engineering, AI Platform Engineering, or...  ...supporting an ML platform, including tools like AWS, Google Cloud Storage, infrastructure-... 
    Amazon Web Service
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    Remote
    20 days ago
  •  ...About the Role Payabli is looking for a Staff Machine Learning Engineer to set the technical direction for ML at Payabli. A few models are already live...  ...is a big plus Familiarity with AWS ML tools (e.g. SageMaker), and experience with feature stores, training... 
    Amazon Web Service
    Full time
    Flexible hours

    Payabli

    Miami, FL
    23 hours ago
  • $172.5k - $306.63k

     ...Design, develop, and maintain robust AI/ML infrastructure solutions to support the training...  ...AI models, using Kubernetes and Python on AWS cloud  Implement and improve distributed...  ...:  Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar  Experience with Pytorch... 
    Amazon Web Service
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    15 hours ago
  •  ...Job Title: ML Engineer Location: Malvern, PA Duration:6 Months Experience Required:...  ...Engineer with strong MLOps expertise on AWS, building, deploying, and monitoring scalable ML pipelines using services like SageMaker, S3, Lambda, Step Functions, and API Gateway... 
    Amazon Web Service

    I-Flow

    Malvern, PA
    3 days ago
  • $200k - $250k

     ...decision-making.Who We Are Looking ForWe're hiring a Staff Machine Learning Engineer to help move forward the ML platform that every AI initiative at AppFolio...  ...and operate AppFolio's ML infrastructure on AWS — ECS, SageMaker, GPU fleets, model serving, autoscaling, and... 
    Amazon Web Service
    Full time
    Flexible hours

    AppFolio

    Santa Barbara, CA
    3 days ago
  • $159k - $208.95k

     ...of our organization. As a Staff Machine Learning Engineer at FanDuel, you will help us...  ...will own critical production ML systems across real-time...  ...with vector store such as AWS OpenSearch, Elasticsearch,...  ...GenAI/LLM platforms (e.g., SageMaker, Bedrock, Databricks, etc.)... 
    Amazon Web Service
    Temporary work
    Local area
    Worldwide

    FanDuel

    Atlanta, GA
    2 days ago
  • $240k - $249.5k

     ...About the OpportunityGrubhub is looking for a Senior Staff Machine Learning Engineer to help lead the machine learning engine behind...  ...offline metric is misleading youExperience with cloud ML infrastructure (AWS/SageMaker or equivalent), model deployment, and production... 
    Amazon Web Service
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    2 days ago
  •  ...We are looking for an experienced ML Ops Engineer to build, automate, and maintain scalable machine...  ...and scripting. ~ Experience with AWS, Azure, or GCP. ~ Strong knowledge of...  .... ~ Experience with MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, or similar ML platforms... 
    Amazon Web Service

    Techvilla Solutions

    River Hills, WI
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Engineer - AWS Trainium & SageMaker [Remote]. Be the first to apply!