ML Performance Engineer
$100k - $150kBright Vision Technologies
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: ML Performance Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization. The ideal candidate has demonstrated impact on production AI workloads, with strong instrumentation and measurement discipline that enables rigorous, data-driven optimization decisions. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production. Key Responsibilities
- Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
- Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
- Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
- Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
- Tune attention implementations using FlashAttention, paged attention, and related techniques.
- Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
- Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
- Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
- Build and maintain rigorous benchmark suites and regression frameworks across workloads.
- Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
- Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
- Evaluate new hardware and software offerings, and advise on adoption.
- Document performance tuning playbooks and share findings broadly across engineering teams.
- Stay current with AI systems research and translate advances into production improvements.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
- Six or more years of experience in performance engineering, ML systems, or HPC.
- Strong proficiency in Python and C++.
- Hands-on experience optimizing deep learning workloads on modern GPUs.
- Deep understanding of distributed training and inference techniques.
- Experience with profiling tools across CPU, GPU, and distributed systems.
- Familiarity with model compression techniques and their accuracy implications.
- Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
- Excellent measurement, debugging, and analytical reasoning skills.
- Strong communication and collaboration skills.
- Experience optimizing LLM inference at production scale.
- Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
- Familiarity with custom kernel authoring in Triton or CUTLASS.
- Experience with FinOps for AI workloads.
- Publications or talks on AI systems performance.
Would you like to know more about this opportunity? For immediate consideration, please send your resume to View email address on aiapply.co or contact us at View phone number on aiapply.co. Learn more about Bright Vision Technologies at
Bright Vision Technologies is an Equal Opportunity Employer.
Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
- ...you :- We are seeking a hands-on MLOps Engineer responsible for the infrastructure, deployment... ...• Monitor platform health, capacity, and performance • Troubleshoot production issues and... ...• Experience supporting AI or ML platforms • Experience with Vertex AI...PerformanceRemote workWorldwide
$170.1k - $258.3k
...pioneer new approaches to model export, kernel development, and performance engineering so that every cycle on our accelerators translates into... ...and custom libraries that sit at the heart of our on‑vehicle ML inference for ADAS and autonomous driving. We own making core...PerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...ML Engineer Location: Remote, Nationwide Our client is building next-generation AI systems designed to move beyond experimentation... ...operation. Create scalable model execution services designed to perform consistently under demanding latency and throughput...PerformanceRemote work
$193.3k - $261.5k
...and Trainium.The Acceleration Kernel Library team is at the forefront of maximizing performance for AWS's custom ML accelerators. Working at the hardware-software boundary, our engineers craft high-performance kernels for ML functions, ensuring every FLOP counts in...PerformanceInternshipLocal areaWork from homeFlexible hours$100.4k - $180.7k
Posting TitleML and Optimization Engineer.LocationCO - Golden.Position TypeLimited Term (Fixed... ..., applied mathematics, and high-performance computing within the Computational Science... ...strengths in high‑performance computing, AI/ML, modeling and simulation, and visualization...PerformanceFull timeFixed term contractLive inLocal areaRemote workRelocationShift work$159.3k - $230.7k
...means across the loop. The team delivers ML models that move the product up the data... ...for your next iteration.As a Senior AI/ML Engineer in the Embodied AI Data Foundations organization... ...that directly improve autonomous driving performance. You will design and run the data...PerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$173k - $253k
Matterport - Senior ML Ops Engineer Job Description CoStar Group is a leading global provider of commercial and residential real estate... ...part of CoStar Group, you will be pivotal in enhancing the performance, efficiency, and scalability of our machine learning models....PerformanceFull timeWork at officeWork from home- ...operating firm of Command Holdings, is seeking a Machine Learning (ML) Engineer to serve on a potential contract supporting USSOCOM's mission... ...making that learn from data to identify patterns and improve performance over time.Building and managing end-to-end ML pipelines,...PerformanceFull timeContract workWork at officeLocal areaVisa sponsorshipWork visa
- ...Active Learning – ML Engineer Location: Remote, Nationwide, EST and CST Our client is an early-stage technology company building... ...where the quality of the underlying data is central to product performance. They are seeking an Active Learning- ML Engineer to build...PerformanceFull timeRemote work
- ...Job Title : ML Engineer Location : Minnetonka, MN FULLTIME ONLY Job Description Must Have Technical... ...data contracts. • Own production health: drift detection, performance regression, rollback strategies, and incident response. Required...PerformanceFull time
$207k - $300k
...Notifications.Develop AI-generated content through advanced context engineering and agentic feedback loops to identify and deliver engaging... ...dense recommendation signals and feedback to improve the performance of the AIGC stack.Resolve key system-level bottlenecks in...Performance- ## ML Engineer**Atlanta,GA30339**Posted: 08/28/2026Employment Type:Contract to HireCategory: IT - AI/MLJob Number: 250989Work Location... ...observability practices, and continuously improve the reliability and performance of machine learning systems across multiple business domains....PerformanceContract workRemote work
$180k - $280k
...autonomous vehicle behavior across real-world scenarios.As a Staff AI/ML Engineer within the Onboard Embodied AI organization, you will be a... ...learning solutions directly impacting autonomous driving performance. Your role is pivotal in designing, architecting, and...PerformanceFull timeLocal areaWork from homeRelocationRelocation package$124k - $250k
...support of others.A Day in the LifeAs a member of our software engineering infra team, you'll solve technical challenges, including... ...state-of-the-art software infrastructure. The team builds a high-performance, high availability, globally distributed ecosystem platform...Performance$189k - $300k
...The team directly works on and delivers ML models to the product that successively... ..., thereby directly impacting AV product performance through smart use of data. As part of this... ..., high-impact team of AI/ML engineers, data scientists and engineers who are passionate...PerformanceFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$200k - $240k
...Senior Machine Learning Engineer (Computer Vision / Vision-Language Models) Remote (U.S... ...frameworks Optimize inference pipelines for performance, accuracy, and cost Collaborate... ...PyTorch skills ~ Proven record of deploying ML systems into production ~ Experience...PerformanceRemote work$153.2k - $234.1k
...autonomous vehicle behavior across real-world scenarios. As a Senior ML Infra Engineer, you will work on the core systems that enable rapid dataset... ...to next. You will develop model training pipelines that are performant, easy to use, and exceptionally reliable. Your success will...PerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$155.42k - $205.9k
...machine learning models, with a focus on performance, availability, concurrency, and scalability... ...development by prioritizing high-impact, ML-centric use cases. About the Role: We are seeking a Senior ML Infrastructure engineer to help build and scale robust Compute platforms...PerformanceFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$128.7k - $261.3k
...infrastructure that powers every machine learning engineer working on our cutting-edge Autonomous... ...entirely on enhancing the safety and performance of the car, rather than managing... ...advanced driverless vehicles.As a Senior ML Infra Engineer, you will build critical infrastructure...PerformanceFull timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- ...and hyper-connected world.Role OverviewAs our Staff Software Engineer, ML infra Engineer for Search & Discovery organization, you will... ...QualificationsComputer science fundamentals: data structures, algorithms, performance complexity, and implications of computer architecture on...PerformanceTemporary work
- AI/ML Ops EngineerLocation: Remote / Hybrid (Client-Facing Consulting Engagement)Employment... ...RoleWe are seeking an experienced AI/ML Engineer to design, deploy, and operate production... ..., and feature management platforms meet performance and quality requirements.* Establish and...PerformanceFull timeContract workLocal areaRemote workFlexible hours
$144.7k - $261.3k
...Foundations, solves critical evaluation challenges for autonomous vehicle development. We engineer high-performance tools that identify top-performing models and partner with data-intensive ML teams to drive rapid innovation.Why Join Us?Develop introspection and evaluation...PerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$295k
...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...enjoy working across the full stack of ML systems, this role gives you the opportunity... ...working on projects such as:Building a high-performance data loading and caching pipeline....PerformanceFull timeWork at officeLocal areaRemote workHome office$140k - $210k
...Senior Machine Learning Engineer Remote, United States Overview Harnham is partnering... ...on approved transactions , placing the performance of its machine learning technology... ...production engineering. You will own complex ML problems from initial exploration and model...PerformancePermanent employmentFull timeTemporary workH1bRemote workVisa sponsorshipWork visa$216.7k - $303.4k
...Senior Machine Learning Systems Engineer Remote - United States Reddit is a community... ...teams. What You’ll Do: As a Senior ML Infrastructure Engineer, you will lead development... ...Collaborate with ML engineers on performance tuning, including improving model...PerformanceFor contractorsWork experience placementRemote work$170k - $200k
...advantage in the robotics revolution. The Instawork Robotics ML Engineer will help build and scale the technology powering physical AI... ...develop methodologies and analytics to measure the quality and performance of our dataset and ML models.Research Reviews - stay on top...PerformanceHourly payInternshipLocal areaShift work$40 - $60 per hour
...Responsibilities Feature Engineering Data Integration Develop and maintain feature engineering... ...pipelines using Data bricks to support ML models effectively Data Pipeline... ...leveraging MLflow for experiment tracking and performance monitoring Query Optimization Low...PerformanceHourly payContract workRemote work$141k - $249k
...will... Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using an AI-first approach.... ...Comprehensively profile model runtime and memory to pinpoint performance bottlenecks. Qualifications: MS/PhD or Bachelors degree...PerformanceWork at officeWork from homeFlexible hours- ...New York City is seeking a Senior Machine Learning Engineer to design, build, train, and deploy production ML and LLM systems that extract actionable insights... ...models with robust metrics, and optimize performance in production at scale. #J-18808-Ljbffr Jobleads...PerformanceRemote job
$100k - $150k
...ML Infrastructure Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering... ..., distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers...PerformanceFull timeH1bLocal areaImmediate startRemote workVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Performance Engineer. Be the first to apply!
- computer vision machine learning engineer Remote
- junior machine learning research engineer Remote
- machine learning software engineer Remote
- ai ml engineer Remote
- senior ml engineer Remote
- machine learning ai engineer Remote
- junior machine learning engineer Remote
- data scientist machine learning engineer Remote
- machine learning engineer Remote
- performance food service Remote



