Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Optimization Engineer

$100k

Jobgether

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Optimization Engineer based in United States.

This is a fully remote opportunity focused on improving the performance, scalability, and economics of large-scale AI systems.
You will optimize training and inference workloads across the stack, from low-level GPU kernels to distributed infrastructure.
The role combines systems engineering, performance analysis, machine learning infrastructure, and compiler-level optimization.
You'll work with modern GPUs and large neural networks, using rigorous measurement and profiling to identify and resolve performance bottlenecks.
The position offers the opportunity to influence production AI workloads where improvements in throughput, latency, and cost have meaningful business impact.
You'll collaborate closely with engineering, product, operations, and business teams while contributing to technical direction and engineering standards.
As a senior technical contributor, you'll also mentor engineers and help drive a culture of measurable, production-ready optimization.

Accountabilities:

  • Optimize training and inference workloads to maximize throughput, minimize latency, and improve cost efficiency across large-scale neural network systems.
  • Analyze and improve performance across the full technology stack, including GPU kernels, memory management, communication, distributed systems, and model execution.
  • Profile CPU, GPU, and distributed workloads to identify bottlenecks and use quantitative analysis to guide optimization decisions.
  • Design and implement performance improvements using Python, C++, and relevant AI systems technologies.
  • Optimize distributed training and inference architectures, including model parallelism, communication strategies, and resource utilization.
  • Evaluate and implement model compression techniques while carefully considering their impact on model accuracy and production performance.
  • Investigate complex performance and reliability issues through systematic debugging, instrumentation, benchmarking, and root-cause analysis.
  • Contribute to production-scale optimization of large language model inference and other demanding AI workloads.
  • Develop and improve low-level optimization techniques, including custom GPU kernels where appropriate.
  • Collaborate with product, design, engineering, operations, and business stakeholders to translate ambiguous requirements into scalable, well-engineered technical solutions.
  • Participate in architecture and code reviews, establish engineering best practices, and contribute to long-term technical strategy.
  • Mentor junior and mid-level engineers, helping raise technical quality and strengthen performance engineering capabilities.
  • Identify opportunities to improve the cost structure of AI workloads through infrastructure optimization and FinOps-oriented analysis.
Requirements:
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 6+ years of professional experience in performance engineering, machine learning systems, high-performance computing, or a closely related field.
  • Strong programming proficiency in Python and C++ , with the ability to develop production-quality, maintainable code.
  • Hands-on experience optimizing deep learning workloads on modern GPU architectures.
  • Deep understanding of distributed training and inference techniques, including parallelism strategies and communication primitives.
  • Strong knowledge of memory hierarchies, GPU/CPU performance characteristics, and systems-level optimization.
  • Experience using profiling and instrumentation tools across CPU, GPU, and distributed environments.
  • Familiarity with model compression methods and their implications for accuracy, performance, and production deployment.
  • Excellent measurement, debugging, analytical reasoning, and problem-solving abilities.
  • Strong communication and collaboration skills, with the ability to explain complex technical concepts to cross-functional stakeholders.
  • Demonstrated ability to work independently, make data-driven technical decisions, and deliver meaningful improvements in production environments.
  • Experience with production-scale LLM inference is strongly preferred.
  • Contributions to projects such as vLLM, TensorRT-LLM, DeepSpeed, or comparable AI systems projects are a plus.
  • Experience with custom kernel development using technologies such as Triton or CUTLASS is preferred.
  • Familiarity with FinOps and cost optimization for AI workloads is advantageous.
  • Publications, conference presentations, or technical talks focused on AI systems or performance engineering are a plus.
  • Must be currently based in the United States and authorized to work in the U.S.; U.S. citizens, permanent residents, EAD holders, and candidates eligible for H-1B transfer are encouraged to apply. New H-1B sponsorship is not available.
Benefits:
  • $100,000 annual salary for this full-time direct W2 position.
  • 100% remote work within the United States.
  • Opportunity to work on challenging AI optimization and high-performance computing problems.
  • Exposure to large-scale neural networks, modern GPU architectures, distributed systems, and production AI infrastructure.
  • Significant opportunities for technical ownership, mentorship, and career growth.
  • Collaborative environment spanning engineering, product, operations, design, and business teams.
  • Opportunity to contribute to impactful production AI systems and advance performance, scalability, and cost efficiency.
  • Equal employment opportunity and an inclusive workplace committed to fair treatment of employees and applicants.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?


Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Vacancy posted 22 hours ago
Similar jobs that could be interesting for youBased on the AI Optimization Engineer in United States vacancy
  • $100k

    Role Description We are seeking an AI Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization... 
    Suggested
    Full time
    Local area
    Immediate start

    Bright Vision Technologies

    Remote
    22 hours ago
  • $180k - $250k

     ...the first and best place to understand and optimize digital customer experiences. Our...  ...technology, business, and operations teams AI-powered insights into any issues impacting...  ...not a research role, and it is not prompt engineering; the team builds production systems that... 
    Suggested
    Full time

    Conviva

    Remote
    22 hours ago
  •  ...AI Engineer – Decision & Optimization Systems About Gallatin At Gallatin, we are rebuilding defense logistics for the warfighters of the United States and allied forces. We take an AI-first approach to modernizing how materiel, fuel, and equipment move from factory... 
    Suggested
    Full time

    Gallatin

    El Segundo, CA
    22 hours ago
  •  ...AI Engineer – Routing & Network Optimizatio About Gallatin At Gallatin, we are rebuilding defense logistics for the warfighters...  ...Role We’re looking for an AI Engineer – Routing & Network Optimization to own how supplies, vehicles, and materiel move across... 
    Suggested
    Full time

    Gallatin

    El Segundo, CA
    22 hours ago
  •  ...Texas Sports Academy is on the lookout for a Senior AI Engineer specializing in LLM (Large Language Model) Systems and RAG (Retrieval-Augmented Generation) Optimization. As we continue to push the boundaries of sports technology, your role will be pivotal in developing... 
    Suggested
    Remote job
    Full time

    Texas Sports Academy

    United States
    22 hours ago
  • Tavus in San Francisco is seeking an experienced Research Scientist/Engineer focused on model optimization to join our core AI team. We thrive in a startup environment, prioritizing independence, calculated risks, and rapid progress to push human-AI interaction forward... 

    Neura Market

    San Francisco, CA
    5 days ago
  • Framework Ventures is seeking an experienced AI model serving engineer to advance inference pipelines across edge and cloud environments in the...  ...with cross‑functional teams to benchmark performance, optimize memory usage, and push the boundaries of real‑world AI capabilities... 

    Framework Ventures

    New York, NY
    2 days ago
  • Intel is seeking a Software Research Engineer/Scientist in the United States (Arizona, Phoenix...  ...software applications, algorithms, and AI solutions. You will prototype concepts,...  ...of findings with a focus on scalable optimization models and impactful technology #J-1880... 

    Intel

    Phoenix, AZ
    2 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-... 
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    4 days ago
  •  ...is the moat — every competitor has humans optimizing campaigns manually. You are building the...  ...intelligence layer that compounds. The Platform engineer creates the tools, you create the...  ...systems, or automated recommendation loops. AI-first development workflow and ability to... 
    Full time

    Hellyeah AI

    San Francisco, CA
    6 days ago
  • $45 - $60 per hour

     ...purpose neural processing unit (GPNPU) architecture. Quadric’s co-optimized software and hardware is targeted to run neural network (NN)...  ...receive hands-on experience working alongside industry experts in AI and semiconductor technology, with access to mentorship and meaningful... 
    Hourly pay
    Temporary work
    Internship
    Work at office
    Relocation

    quadric, Inc

    Burlingame, CA
    4 days ago
  •  ...generation of autonomous system intelligence. As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient, production-...  ...to minimize latency and maximize memory bandwidth on AI accelerators. Write production-level, low latency, and... 
    Temporary work
    Relocation package

    Zoox

    Washington DC
    more than 2 months ago
  •  ...Job Description Job Description NextGen is seeking a highly motivated and technically skilled Edge AI/Model Optimization Engineer to support the deployment, optimization, and sustainment of AI and agentic AI capabilities within edge and tactical computing environments... 
    Local area

    NextGen Federal Systems

    Aberdeen, MD
    more than 2 months ago
  •  ...Laboratory of the Rockies (NLR) in Golden, CO seeks an ML and Optimization Engineer to design, develop, and test software supporting energy research and high-performance computing. You will work across AI/ML and optimization to build production-ready systems that operate... 

    National Renewable Energy Laboratory

    Golden, CO
    1 day ago
  • Forrest T. Jones & Company seeks an AI Process Optimization Engineer to identify and improve business processes using AI, automation, and data analytics. You will work with stakeholders, IT, and data professionals to evaluate AI opportunities, develop intelligent process... 

    Forrest T. Jones & Company

    Kansas City, MO
    4 days ago
  • A rapidly scaling technology company is seeking a Forward Deployed Engineer in New York. The role involves building, optimizing AI agents and working closely with customer data. You should have 2+ years in relevant engineering roles and be proficient in SQL. Excellent communication... 

    Searchability®

    New York, NY
    5 days ago
  • $100k

     ...position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Optimization Engineer based in United States. This is a fully remote opportunity focused on improving the performance, scalability, and... 
    Remote job
    Permanent employment
    Full time
    H1b

    jobgether

    United States
    6 days ago
  • $200k - $350k

     ...Figure is an AI Robotics company developing a general purpose humanoid. Our humanoid...  ...Our Helix team is looking for Perception Engineers to empower Figure humanoid robots to perform...  ...cross-functionally to evolve and optimize our autonomy stack. Requirements:... 
    Full time
    Work at office

    Figure

    Remote
    22 hours ago
  •  ...We are looking for a passionate and hands-on AI Builder to join our AI Engineering team. This role is ideal for someone with 1–2 years of experience...  .... Participate in testing, model evaluation, prompt optimization, and performance tuning. Document solutions and... 
    Full time

    Brillio

    Minneapolis, MN
    22 hours ago
  •  ...strong background in machine learning, artificial intelligence, and data science, the AI engineer will design, develop, and deploy AI solutions to address business challenges and optimize operations. Work closely with cross-functional teams to integrate AI models into... 
    Full time
    Work experience placement
    Work at office

    Luck Companies

    Remote
    22 hours ago
  •  ...one of the hardest problems in enterprise AI: AI models are generic but company...  ...London. We're a deeply technical team of engineers, AI researchers, and strategists with a...  ...time: human-in-the-loop feedback, prompt optimization and context engineering You have experience... 
    Full time

    Edra

    New York, NY
    22 hours ago
  •  ...About the role We’re seeking an experienced engineer to deploy enterprise-grade AI solutions, focusing on Retrieval-Augmented Generation (RAG) pipelines...  ...solutions for emerging needs. Responsibilities: Optimize and support solutions within strategic accounts on the... 
    Full time

    StackAI

    San Francisco, CA
    22 hours ago
  •  ...Founding AI Engineer (AI + Production) New York City (5 days on-site) · Top of market + equity + benefits TL;DR: Build AI that accelerates...  ...model for thermal analysis; pair with fullstack engineer to optimize inference latency; attend NVIDIA collaboration session on... 
    Full time
    Immediate start
    Weekend work

    Everstar

    New York, NY
    22 hours ago
  •  ...how Socure builds, deploys, and scales AI-driven identity solutions while enabling...  ...Reporting to the Head of New Product Engineering, you'll join a new Internal AI Engineering...  ...assessment, failure-mode analysis, and optimization of speed and accuracy of specific agentic... 
    Full time

    Socure

    Remote
    22 hours ago
  • $195k - $230k

     ...About Elicit Elicit is an AI research assistant that uses language models to help researchers...  ...on our mission. What is an "AI Engineer"? AI engineering is a new category of...  ...+ equity, depending on your level. We're optimizing for a hire who can contribute at a L4/... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    Elicit

    Oakland, CA
    22 hours ago
  •  ...We’re hiring an AI Engineer to build the intelligence layer for the leading AI companion for language learning. You’ll own the core AI...  ...conversations. Improve speech & conversation intelligence — optimize ASR feedback, turn-taking, and dialogue flow to make interactions... 
    Full time

    Pingo پینگو

    San Francisco, CA
    22 hours ago
  • $100.8k - $168k

     ...for designing, developing, and deploying AI-powered solutions across the company. Focuses...  ...closely with data scientists, engineers, and business leaders to deliver AI-driven...  ...Expectations Develops, trains, and optimizes machine learning and deep learning models... 
    Full time
    Flexible hours

    Finance Of America

    United States
    22 hours ago
  •  ...and increasing revenue through cutting-edge tools like AI chatbots, customized funnel optimization, targeted paid advertising, and powerful email/SMS...  ...seeking a talented and experienced Machine Learning/AI Engineer with a specialized focus on OpenAI technologies. As... 
    Full time

    Sales Hub Careers

    Chino Hills, CA
    22 hours ago
  •  ...front line needs and what is delivered. Our AI-native platform, Air Enterprise Readiness...  ...We are seeking an experienced AI Engineer to join our AI Enablement team, focused on...  ...identify and implement AI-driven solutions to optimize their workflows and reduce operational... 
    Full time
    Work at office

    Air Company

    Pittsburgh, PA
    22 hours ago
  • $187k - $253k

     ...Eigen Labs is building the coordination engine for a world run by humans and agents alike. AI is becoming abundant. Trust is not. We are building the...  ...Make agents cheap (cost-aware execution, performance optimization) Make agents useful in production (not demos -... 
    Remote job
    Full time
    Temporary work
    Flexible hours
    Shift work

    Eigen Labs

    Remote
    22 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Optimization Engineer. Be the first to apply!