Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Performance Engineer

$100k - $150k
Full-time

Bright Vision Technologies

ML Performance Engineer - Remote


Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.


Job Title: ML Performance Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years


Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.


Job Summary
We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization. The ideal candidate has demonstrated impact on production AI workloads, with strong instrumentation and measurement discipline that enables rigorous, data-driven optimization decisions. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production.

Key Responsibilities
  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
  • Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
  • Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
  • Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Tune attention implementations using FlashAttention, paged attention, and related techniques.
  • Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
  • Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
  • Build and maintain rigorous benchmark suites and regression frameworks across workloads.
  • Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
  • Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
  • Evaluate new hardware and software offerings, and advise on adoption.
  • Document performance tuning playbooks and share findings broadly across engineering teams.
  • Stay current with AI systems research and translate advances into production improvements.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference techniques.
  • Experience with profiling tools across CPU, GPU, and distributed systems.
  • Familiarity with model compression techniques and their accuracy implications.
  • Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
  • Excellent measurement, debugging, and analytical reasoning skills.
  • Strong communication and collaboration skills.
Preferred Qualifications
  • Experience optimizing LLM inference at production scale.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Familiarity with custom kernel authoring in Triton or CUTLASS.
  • Experience with FinOps for AI workloads.
  • Publications or talks on AI systems performance.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to View email address on aiapply.co or contact us at View phone number on aiapply.co. Learn more about Bright Vision Technologies at
Bright Vision Technologies is an Equal Opportunity Employer.

Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Vacancy posted 7 days ago
Similar jobs that could be interesting for youBased on the ML Performance Engineer in Remote vacancy
  •  ...you :- We are seeking a hands-on MLOps Engineer responsible for the infrastructure, deployment...  ...• Monitor platform health, capacity, and performance • Troubleshoot production issues and...  ...• Experience supporting AI or ML platforms • Experience with Vertex AI... 
    Performance
    Remote work
    Worldwide

    CitiusTech

    United States
    3 hours ago
  • $170.1k - $258.3k

     ...pioneer new approaches to model export, kernel development, and performance engineering so that every cycle on our accelerators translates into...  ...and custom libraries that sit at the heart of our on‑vehicle ML inference for ADAS and autonomous driving. We own making core... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Center Line, MI
    5 days ago
  •  ...ML Engineer Location: Remote, Nationwide Our client is building next-generation AI systems designed to move beyond experimentation...  ...operation. Create scalable model execution services designed to perform consistently under demanding latency and throughput... 
    Performance
    Remote work

    Blue Signal Search

    United States
    2 days ago
  • $193.3k - $261.5k

     ...and Trainium.The Acceleration Kernel Library team is at the forefront of maximizing performance for AWS's custom ML accelerators. Working at the hardware-software boundary, our engineers craft high-performance kernels for ML functions, ensuring every FLOP counts in... 
    Performance
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago
  • $100.4k - $180.7k

    Posting TitleML and Optimization Engineer.LocationCO - Golden.Position TypeLimited Term (Fixed...  ..., applied mathematics, and high-performance computing within the Computational Science...  ...strengths in high‑performance computing, AI/ML, modeling and simulation, and visualization... 
    Performance
    Full time
    Fixed term contract
    Live in
    Local area
    Remote work
    Relocation
    Shift work

    National Renewable Energy Laboratory

    Golden, CO
    1 day ago
  • $159.3k - $230.7k

     ...means across the loop. The team delivers ML models that move the product up the data...  ...for your next iteration.As a Senior AI/ML Engineer in the Embodied AI Data Foundations organization...  ...that directly improve autonomous driving performance. You will design and run the data... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    1 day ago
  • $173k - $253k

    Matterport - Senior ML Ops Engineer Job Description CoStar Group is a leading global provider of commercial and residential real estate...  ...part of CoStar Group, you will be pivotal in enhancing the performance, efficiency, and scalability of our machine learning models.... 
    Performance
    Full time
    Work at office
    Work from home

    Matterport

    Sunnyvale, CA
    3 days ago
  •  ...operating firm of Command Holdings, is seeking a Machine Learning (ML) Engineer to serve on a potential contract supporting USSOCOM's mission...  ...making that learn from data to identify patterns and improve performance over time.Building and managing end-to-end ML pipelines,... 
    Performance
    Full time
    Contract work
    Work at office
    Local area
    Visa sponsorship
    Work visa

    Command Holdings

    Tampa, FL
    1 day ago
  •  ...Active Learning – ML Engineer Location: Remote, Nationwide, EST and CST Our client is an early-stage technology company building...  ...where the quality of the underlying data is central to product performance. They are seeking an Active Learning- ML Engineer to build... 
    Performance
    Full time
    Remote work

    Blue Signal Search

    United States
    4 days ago
  •  ...Job Title : ML Engineer Location : Minnetonka, MN FULLTIME ONLY Job Description Must Have Technical...  ...data contracts. • Own production health: drift detection, performance regression, rollback strategies, and incident response. Required... 
    Performance
    Full time

    AceStack LLC

    Minnetonka, MN
    4 days ago
  • $207k - $300k

     ...Notifications.Develop AI-generated content through advanced context engineering and agentic feedback loops to identify and deliver engaging...  ...dense recommendation signals and feedback to improve the performance of the AIGC stack.Resolve key system-level bottlenecks in... 
    Performance

    Google

    Mountain View, CA
    3 days ago
  • ## ML Engineer**Atlanta,GA30339**Posted: 08/28/2026Employment Type:Contract to HireCategory: IT - AI/MLJob Number: 250989Work Location...  ...observability practices, and continuously improve the reliability and performance of machine learning systems across multiple business domains.... 
    Performance
    Contract work
    Remote work

    The Intersect Group

    Atlanta, GA
    2 days ago
  • $180k - $280k

     ...autonomous vehicle behavior across real-world scenarios.As a Staff AI/ML Engineer within the Onboard Embodied AI organization, you will be a...  ...learning solutions directly impacting autonomous driving performance. Your role is pivotal in designing, architecting, and... 
    Performance
    Full time
    Local area
    Work from home
    Relocation
    Relocation package

    General Motors

    Sunnyvale, TX
    1 day ago
  • $124k - $250k

     ...support of others.A Day in the LifeAs a member of our software engineering infra team, you'll solve technical challenges, including...  ...state-of-the-art software infrastructure. The team builds a high-performance, high availability, globally distributed ecosystem platform... 
    Performance

    AppLovin

    Palo Alto, CA
    4 days ago
  • $189k - $300k

     ...The team directly works on and delivers ML models to the product that successively...  ..., thereby directly impacting AV product performance through smart use of data. As part of this...  ..., high-impact team of AI/ML engineers, data scientists and engineers who are passionate... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $200k - $240k

     ...Senior Machine Learning Engineer (Computer Vision / Vision-Language Models) Remote (U.S...  ...frameworks Optimize inference pipelines for performance, accuracy, and cost Collaborate...  ...PyTorch skills ~ Proven record of deploying ML systems into production ~ Experience... 
    Performance
    Remote work

    Harnham

    New York, NY
    4 days ago
  • $153.2k - $234.1k

     ...autonomous vehicle behavior across real-world scenarios. As a Senior ML Infra Engineer, you will work on the core systems that enable rapid dataset...  ...to next. You will develop model training pipelines that are performant, easy to use, and exceptionally reliable. Your success will... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    2 days ago
  • $155.42k - $205.9k

     ...machine learning models, with a focus on performance, availability, concurrency, and scalability...  ...development by prioritizing high-impact, ML-centric use cases. About the Role: We are seeking a Senior ML Infrastructure engineer to help build and scale robust Compute platforms... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $128.7k - $261.3k

     ...infrastructure that powers every machine learning engineer working on our cutting-edge Autonomous...  ...entirely on enhancing the safety and performance of the car, rather than managing...  ...advanced driverless vehicles.As a Senior ML Infra Engineer, you will build critical infrastructure... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    3 days ago
  •  ...and hyper-connected world.Role OverviewAs our Staff Software Engineer, ML infra Engineer for Search & Discovery organization, you will...  ...QualificationsComputer science fundamentals: data structures, algorithms, performance complexity, and implications of computer architecture on... 
    Performance
    Temporary work

    Coupang

    Mountain View, CA
    4 days ago
  • AI/ML Ops EngineerLocation: Remote / Hybrid (Client-Facing Consulting Engagement)Employment...  ...RoleWe are seeking an experienced AI/ML Engineer to design, deploy, and operate production...  ..., and feature management platforms meet performance and quality requirements.* Establish and... 
    Performance
    Full time
    Contract work
    Local area
    Remote work
    Flexible hours

    Slalom

    Houston, TX
    4 days ago
  • $144.7k - $261.3k

     ...Foundations, solves critical evaluation challenges for autonomous vehicle development. We engineer high-performance tools that identify top-performing models and partner with data-intensive ML teams to drive rapid innovation.Why Join Us?Develop introspection and evaluation... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Center Line, MI
    5 days ago
  • $295k

     ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...enjoy working across the full stack of ML systems, this role gives you the opportunity...  ...working on projects such as:Building a high-performance data loading and caching pipeline.... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    4 days ago
  • $140k - $210k

     ...Senior Machine Learning Engineer Remote, United States Overview Harnham is partnering...  ...on approved transactions , placing the performance of its machine learning technology...  ...production engineering. You will own complex ML problems from initial exploration and model... 
    Performance
    Permanent employment
    Full time
    Temporary work
    H1b
    Remote work
    Visa sponsorship
    Work visa

    Harnham

    United States
    4 days ago
  • $216.7k - $303.4k

     ...Senior Machine Learning Systems Engineer Remote - United States Reddit is a community...  ...teams. What You’ll Do: As a Senior ML Infrastructure Engineer, you will lead development...  ...Collaborate with ML engineers on performance tuning, including improving model... 
    Performance
    For contractors
    Work experience placement
    Remote work

    Reddit

    United States
    3 days ago
  • $170k - $200k

     ...advantage in the robotics revolution. The Instawork Robotics ML Engineer will help build and scale the technology powering physical AI...  ...develop methodologies and analytics to measure the quality and performance of our dataset and ML models.Research Reviews - stay on top... 
    Performance
    Hourly pay
    Internship
    Local area
    Shift work

    Instawork

    San Francisco, CA
    1 day ago
  • $40 - $60 per hour

     ...Responsibilities Feature Engineering Data Integration Develop and maintain feature engineering...  ...pipelines using Data bricks to support ML models effectively Data Pipeline...  ...leveraging MLflow for experiment tracking and performance monitoring Query Optimization Low... 
    Performance
    Hourly pay
    Contract work
    Remote work

    Cedent

    United States
    5 days ago
  • $141k - $249k

     ...will... Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using an AI-first approach....  ...Comprehensively profile model runtime and memory to pinpoint performance bottlenecks. Qualifications: MS/PhD or Bachelors degree... 
    Performance
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    2 days ago
  •  ...New York City is seeking a Senior Machine Learning Engineer to design, build, train, and deploy production ML and LLM systems that extract actionable insights...  ...models with robust metrics, and optimize performance in production at scale. #J-18808-Ljbffr Jobleads... 
    Performance
    Remote job

    Jobleads-US

    New York, NY
    2 days ago
  • $100k - $150k

     ...ML Infrastructure Engineer - Remote    Bright Vision Technologies is a technology consulting and software development company delivering...  ..., distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers... 
    Performance
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Plymouth, MN
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Performance Engineer. Be the first to apply!