Staff ML Engineer - AWS Trainium & SageMaker [Remote]
jobgether
- Remote job
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff ML Engineer – AWS Trainium & SageMaker based in United States.
As a Staff ML Engineer, you will design, train, optimize, and operate production machine learning workloads on AWS Trainium and Amazon SageMaker.
You will work deeply across the ML stack, from PyTorch training code and distributed execution to accelerator hardware and compiler behavior.
The role requires understanding how training workloads behave on custom silicon rather than treating infrastructure as a black box.
You will diagnose complex training issues, optimize throughput and cost, and build reliable pipelines for real production workloads.
You will work directly with engineering teams to translate business and technical requirements into scalable training solutions.
The environment is highly hands-on, technical, and client-facing, with a focus on solving specialized problems that require deep engineering expertise.
This is an opportunity to work at the intersection of machine learning, cloud infrastructure, distributed systems, and custom AI acceleration.
Accountabilities:
- Train and operate machine learning models using Amazon SageMaker with AWS Trainium as the underlying compute infrastructure.
- Develop and optimize PyTorch training workloads for execution on AWS Trainium.
- Analyze how PyTorch code compiles and executes across Trainium and NeuronCore architecture.
- Optimize memory utilization, throughput, and other performance characteristics of training workloads.
- Diagnose training failures and performance issues caused by hardware, compiler behavior, device configuration, data, or model code.
- Distinguish model- and data-level problems from accelerator- and compiler-level issues during troubleshooting.
- Translate high-level requirements for Trainium workloads into complete, production-ready training pipelines.
- Design training workflows that balance performance, scalability, reliability, and cloud infrastructure costs.
- Tune distributed and multi-device training workloads for throughput and cost efficiency.
- Operate production training workloads and help ensure their reliability throughout the ML lifecycle.
- Work directly with client engineering teams to scope, design, and deliver specialized machine learning workloads.
- Collaborate with internal engineering teams to develop production-grade solutions rather than isolated prototypes or notebook experiments.
- Investigate complex technical issues across the ML software and hardware stack.
- Apply a deep understanding of accelerator behavior to improve training architecture and implementation decisions.
- Contribute to production engineering practices around deployment, monitoring, troubleshooting, and operational reliability.
- Help translate emerging AI infrastructure capabilities into practical production solutions.
- Communicate technical findings and tradeoffs clearly with both technical stakeholders and client teams.
Requirements:
- Strong hands-on experience with PyTorch, including a deep understanding of training workflows.
- Experience with distributed or multi-device model training is highly desirable.
- Production experience using Amazon SageMaker for model training and/or inference.
- Strong understanding of how machine learning workloads interact with accelerator hardware and device-specific compilation.
- Ability to debug issues that originate at the hardware, accelerator, compiler, or runtime layer rather than solely within model or data code.
- Experience optimizing training workloads for performance, throughput, memory utilization, or cost.
- Strong Python programming fundamentals.
- Understanding of distributed training concepts and production ML infrastructure.
- Ability to design and operate end-to-end managed training pipelines.
- AWS experience and familiarity with cloud-based machine learning infrastructure.
- AWS Trainium or AWS Inferentia experience and familiarity with the AWS Neuron SDK is strongly preferred.
- Candidates without direct Trainium or Inferentia experience may also be considered if they have deep PyTorch expertise and demonstrated ability to quickly learn new hardware targets.
- Strong analytical and problem-solving skills, particularly when diagnosing complex system-level issues.
- Ability to reason across multiple layers of the technology stack, from model code through frameworks, compilers, accelerators, and cloud infrastructure.
- Experience working in a production engineering environment with high standards for reliability and delivery.
- Strong communication and collaboration skills for working directly with client and internal engineering teams.
- Comfortable operating in ambiguous environments where requirements and technical challenges may evolve.
- Willingness and ability to travel approximately 20% within the United States.
- Must be legally authorized to work in the United States.
Benefits:
- Opportunity to work on production AI systems using AWS Trainium and Amazon SageMaker.
- Exposure to custom AI accelerator hardware, compiler behavior, distributed training, and advanced ML infrastructure.
- Hands-on work with complex machine learning workloads for enterprise clients.
- Direct collaboration with client and internal engineering teams.
- Opportunity to solve specialized technical problems that extend beyond conventional ML application development.
- Production-focused engineering environment emphasizing systems that ship and operate at scale.
- Approximately 20% U.S.-based travel associated with the role.
- Final interview and onboarding may require onsite participation.
- Professional growth through work across machine learning, cloud infrastructure, hardware acceleration, and production engineering.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
- ...Position Overview We're seeking a Staff ML Engineer to advance our voice AI and clinical NLP... ...infrastructure using Triton Inference Server on AWS GovCloud Create specialty-specific... ...with cloud ML platforms (AWS SageMaker, GCP Vertex AI) ~ Master's or PhD in...Amazon Web ServiceWork at office3 days per week
$190k - $234k
...Staff ML Engineer Palo Alto, CA About Typeface: We help the world's biggest brands move from brief to fully personalized campaigns... ...prototyping ~ Prior experience with AzureML, GCP Vertex, AWS Sagemaker or similar platforms ~ Bonus: Hands-on experience in deploying...Amazon Web ServiceWork at officeLocal areaFlexible hours3 days per week- ...You’ll Make an Impact As a Staff Machine Learning Engineer , you will play a key role... ...scalability, and high quality of our ML systems from development to... ..., distributed ML systems (Sagemaker) - Deep MLOps Proficiency:... ...ML models at scale on AWS, with experience using...Amazon Web ServiceFull timeRemote work
- ...Description: We are seeking a talented and experienced Deep Learning Engineer with a specialization in Large Language Models (LLMs) to join... ...techniques and experience with cloud computing platforms (e.g., AWS, Azure, Google Cloud) • Excellent problem-solving skills and...Amazon Web Service
$300k
...term relationship. We're doubling down on ML as the future of Grindr, and in these... ...leadership across teams, collaborating with engineering, data science and product teams to turn bold... ...with any public cloud environment - AWS, GCP or Databricks. Expertise in...Amazon Web ServiceCasual workWork at officeImmediate startWorldwideFlexible hours$220k - $280k
...industry together? As a Staff Machine Learning Engineer, you will lead the... ...integrating robust, low-latency ML models across our sports... ..., Kubeflow, Databricks, or SageMaker. ~ Strong Coding Skills:... ...Functions, GKE, Vertex AI) or AWS equivalents. What...Amazon Web ServiceFull timeRemote workWork visaFlexible hours- ...terms of business impact. Own and prioritize the ML roadmap with product, engineering, and risk operations. Mentor ML engineers and... ...routing experience is preferred. Familiarity with AWS ML tools such as SageMaker, feature stores, training and inference pipelines,...Amazon Web ServiceFull timeFlexible hours
- ...Staff Machine Learning Engineer In Brief We’re an early-stage startup on... ...satisfied with training and tuning ML models that predict... ...data as it moves through our AWS services, or improving the... ...using MLOps tools such as SageMaker and MLFlow. Preferred...Amazon Web ServiceFull timeLocal areaRemote work
$161k - $221.5k
...and technology solutions. About the Role We are seeking a Staff ML/AI Engineer to define and drive the architectural vision for Lyra’s machine... ...: Strong experience architecting cloud-native solutions on AWS (or equivalent cloud providers). ~Strategic Communication: Exceptional...Amazon Web ServiceFull time- ...tech, and hyper-connected world. Role Overview As our Staff Software Engineer, ML infra Engineer for Search & Discovery organization, you... ...integrating applications and platforms with cloud technologies (i.e, AWS and GCP) Strong verbal and written communication skills...Amazon Web ServiceTemporary workFlexible hours
$227.33k - $312.58k
...Job Description Job Description We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier Models organization. In this role... ..., data quality and lineage tools Cloud & Infrastructure: AWS or GCP, containerized data workloads, CI/CD, infrastructure‑...Amazon Web ServiceWork at officeLocal areaImmediate start3 days per week$298k - $351k
...are seeking an exceptional Staff Machine Learning Operations Engineer to join our Platform... ...efficiency of Garner's production ML systems, including... ...stack: model serving (e.g., Sagemaker, Triton, or equivalent), feature... ...containerization, cloud (AWS preferred), Terraform/IaC,...Amazon Web ServiceWork at officeWork visaFlexible hours3 days per week- ...Job Description Are you a passionate Machine Learning Engineer with a strong background in SageMaker, prompt engineering, and LLM (Large Language Model)... ...learning community. - Experience with cloud services (AWS, Azure, Google Cloud) and containerization technologies...Amazon Web ServiceRemote work
- ...full-time role partners with a fast-paced core engineering team to design, deploy, and optimize production-ready AI/ML solutions. Key Responsibilities Design, build... ...centered on large language models using Python and AWS infrastructure. Architect scalable, high-...Amazon Web ServiceFull timeRemote work
- ...Role We are seeking an exceptional Staff Machine Learning Engineer to lead the design and development of... .... You will architect large-scale ML systems that detect and prevent fraud... ...~ Familiarity with cloud platforms (AWS, GCP, or Azure) and containerized deployments...Amazon Web Service
- ...Staff Machine Learning Engineer webAI is the first end-to-end private AI platform. Enterprises and Governments... ...cross-functional teams to integrate ML models into our platform. Conduct... ...with cloud computing services (AWS, Azure, GCP). Knowledge of Big Data...Amazon Web ServiceLive outWork at officeLocal areaFlexible hours
- ...Staff Machine Learning Engineer - Generative AI Lexington, KY Xometry powers the industries of today... ...datasets Deploy in the cloud – Use AWS and other platforms to train, optimize... ...and grow – Guide teammates on advanced ML methods, model architecture, and best...Amazon Web ServiceImmediate start
$200k - $220k
...global manufacturing capacity. We are looking for a Staff Machine Learning Engineer to join our growing AI/ML team. This is a senior individual contributor role... ...-time ML products at scale in cloud environments (AWS strongly preferred), including auto-scaling,...Amazon Web Service$307k - $352k
...About the role: We are seeking a Staff Machine Learning Engineer to set the technical direction for machine learning at Kikoff. ML sits at the center of our business: our underwriting... ...experience with cloud infrastructure (AWS or GCP), containerization (Docker,...Amazon Web ServiceLocal areaImmediate start$200k - $275k
...Our client, a growing FinTech company, are hiring a Staff Machine Learning Engineer to join their team in Colorado. The successful candidate will play... ...systems and deploying cloud-native applications across AWS, Azure or Google Cloud Platform (GCP). Strong understanding...Amazon Web ServiceFull time- ...our talented Team. Job Title: MLOps Platform Engineer (SageMaker) Location(s): Onsite Job Summary: This... ...extensive experience in cloud infrastructure or ML platform operations, with a specific focus on AWS and Amazon SageMaker. The role involves designing...Amazon Web Service
- ...Employment Type: Full-time Department: Engineering & Product As a Staff Machine Learning Engineer: You... ...where refined data or internal ML/AI models can improve our product outcomes... ...-driven architecture using Kafka, AWS (EKS), Python, Django/FastAPI, and Postgres...Amazon Web ServiceFull timeWork at office
$292.5k - $409.5k
...teams. What You’ll Do: As a Senior Staff Software Engineer, you will help define and lead the... ...Might Be: ~10+ years of experience in ML Engineering, AI Platform Engineering, or... ...supporting an ML platform, including tools like AWS, Google Cloud Storage, infrastructure-...Amazon Web ServiceFull timeFor contractorsWork experience placementFlexible hours- ...About the Role Payabli is looking for a Staff Machine Learning Engineer to set the technical direction for ML at Payabli. A few models are already live... ...is a big plus Familiarity with AWS ML tools (e.g. SageMaker), and experience with feature stores, training...Amazon Web ServiceFull timeFlexible hours
$172.5k - $306.63k
...Design, develop, and maintain robust AI/ML infrastructure solutions to support the training... ...AI models, using Kubernetes and Python on AWS cloud Implement and improve distributed... ...: Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar Experience with Pytorch...Amazon Web ServiceFull timeTemporary workLocal areaWorldwide- ...Job Title: ML Engineer Location: Malvern, PA Duration:6 Months Experience Required:... ...Engineer with strong MLOps expertise on AWS, building, deploying, and monitoring scalable ML pipelines using services like SageMaker, S3, Lambda, Step Functions, and API Gateway...Amazon Web Service
$200k - $250k
...decision-making.Who We Are Looking ForWe're hiring a Staff Machine Learning Engineer to help move forward the ML platform that every AI initiative at AppFolio... ...and operate AppFolio's ML infrastructure on AWS — ECS, SageMaker, GPU fleets, model serving, autoscaling, and...Amazon Web ServiceFull timeFlexible hours$159k - $208.95k
...of our organization. As a Staff Machine Learning Engineer at FanDuel, you will help us... ...will own critical production ML systems across real-time... ...with vector store such as AWS OpenSearch, Elasticsearch,... ...GenAI/LLM platforms (e.g., SageMaker, Bedrock, Databricks, etc.)...Amazon Web ServiceTemporary workLocal areaWorldwide$240k - $249.5k
...About the OpportunityGrubhub is looking for a Senior Staff Machine Learning Engineer to help lead the machine learning engine behind... ...offline metric is misleading youExperience with cloud ML infrastructure (AWS/SageMaker or equivalent), model deployment, and production...Amazon Web ServiceFull timeTemporary workWork at officeFlexible hours3 days per week- ...We are looking for an experienced ML Ops Engineer to build, automate, and maintain scalable machine... ...and scripting. ~ Experience with AWS, Azure, or GCP. ~ Strong knowledge of... .... ~ Experience with MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, or similar ML platforms...Amazon Web Service
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff ML Engineer - AWS Trainium & SageMaker [Remote]. Be the first to apply!
- engineering aide United States
- technology administrator United States
- research assistant engineering United States
- staff security engineer United States
- information technology administrative assistant United States
- assistant engineering manager United States
- assistant building engineer United States
- project engineer assistant project manager United States
- staff qa engineer United States
- staff devops engineer United States



