AI Optimization Engineer
$100kJobgether
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Optimization Engineer based in United States. This is a fully remote opportunity focused on improving the performance, scalability, and economics of large-scale AI systems.
You will optimize training and inference workloads across the stack, from low-level GPU kernels to distributed infrastructure.
The role combines systems engineering, performance analysis, machine learning infrastructure, and compiler-level optimization.
You'll work with modern GPUs and large neural networks, using rigorous measurement and profiling to identify and resolve performance bottlenecks.
The position offers the opportunity to influence production AI workloads where improvements in throughput, latency, and cost have meaningful business impact.
You'll collaborate closely with engineering, product, operations, and business teams while contributing to technical direction and engineering standards.
As a senior technical contributor, you'll also mentor engineers and help drive a culture of measurable, production-ready optimization. Accountabilities:
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1 We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
You will optimize training and inference workloads across the stack, from low-level GPU kernels to distributed infrastructure.
The role combines systems engineering, performance analysis, machine learning infrastructure, and compiler-level optimization.
You'll work with modern GPUs and large neural networks, using rigorous measurement and profiling to identify and resolve performance bottlenecks.
The position offers the opportunity to influence production AI workloads where improvements in throughput, latency, and cost have meaningful business impact.
You'll collaborate closely with engineering, product, operations, and business teams while contributing to technical direction and engineering standards.
As a senior technical contributor, you'll also mentor engineers and help drive a culture of measurable, production-ready optimization. Accountabilities:
- Optimize training and inference workloads to maximize throughput, minimize latency, and improve cost efficiency across large-scale neural network systems.
- Analyze and improve performance across the full technology stack, including GPU kernels, memory management, communication, distributed systems, and model execution.
- Profile CPU, GPU, and distributed workloads to identify bottlenecks and use quantitative analysis to guide optimization decisions.
- Design and implement performance improvements using Python, C++, and relevant AI systems technologies.
- Optimize distributed training and inference architectures, including model parallelism, communication strategies, and resource utilization.
- Evaluate and implement model compression techniques while carefully considering their impact on model accuracy and production performance.
- Investigate complex performance and reliability issues through systematic debugging, instrumentation, benchmarking, and root-cause analysis.
- Contribute to production-scale optimization of large language model inference and other demanding AI workloads.
- Develop and improve low-level optimization techniques, including custom GPU kernels where appropriate.
- Collaborate with product, design, engineering, operations, and business stakeholders to translate ambiguous requirements into scalable, well-engineered technical solutions.
- Participate in architecture and code reviews, establish engineering best practices, and contribute to long-term technical strategy.
- Mentor junior and mid-level engineers, helping raise technical quality and strengthen performance engineering capabilities.
- Identify opportunities to improve the cost structure of AI workloads through infrastructure optimization and FinOps-oriented analysis.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
- 6+ years of professional experience in performance engineering, machine learning systems, high-performance computing, or a closely related field.
- Strong programming proficiency in Python and C++ , with the ability to develop production-quality, maintainable code.
- Hands-on experience optimizing deep learning workloads on modern GPU architectures.
- Deep understanding of distributed training and inference techniques, including parallelism strategies and communication primitives.
- Strong knowledge of memory hierarchies, GPU/CPU performance characteristics, and systems-level optimization.
- Experience using profiling and instrumentation tools across CPU, GPU, and distributed environments.
- Familiarity with model compression methods and their implications for accuracy, performance, and production deployment.
- Excellent measurement, debugging, analytical reasoning, and problem-solving abilities.
- Strong communication and collaboration skills, with the ability to explain complex technical concepts to cross-functional stakeholders.
- Demonstrated ability to work independently, make data-driven technical decisions, and deliver meaningful improvements in production environments.
- Experience with production-scale LLM inference is strongly preferred.
- Contributions to projects such as vLLM, TensorRT-LLM, DeepSpeed, or comparable AI systems projects are a plus.
- Experience with custom kernel development using technologies such as Triton or CUTLASS is preferred.
- Familiarity with FinOps and cost optimization for AI workloads is advantageous.
- Publications, conference presentations, or technical talks focused on AI systems or performance engineering are a plus.
- Must be currently based in the United States and authorized to work in the U.S.; U.S. citizens, permanent residents, EAD holders, and candidates eligible for H-1B transfer are encouraged to apply. New H-1B sponsorship is not available.
- $100,000 annual salary for this full-time direct W2 position.
- 100% remote work within the United States.
- Opportunity to work on challenging AI optimization and high-performance computing problems.
- Exposure to large-scale neural networks, modern GPU architectures, distributed systems, and production AI infrastructure.
- Significant opportunities for technical ownership, mentorship, and career growth.
- Collaborative environment spanning engineering, product, operations, design, and business teams.
- Opportunity to contribute to impactful production AI systems and advance performance, scalability, and cost efficiency.
- Equal employment opportunity and an inclusive workplace committed to fair treatment of employees and applicants.
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1 We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Vacancy posted 22 hours ago
Similar jobs that could be interesting for youBased on the AI Optimization Engineer in United States vacancy
$100k
Role Description We are seeking an AI Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization...SuggestedFull timeLocal areaImmediate start$180k - $250k
...the first and best place to understand and optimize digital customer experiences. Our... ...technology, business, and operations teams AI-powered insights into any issues impacting... ...not a research role, and it is not prompt engineering; the team builds production systems that...SuggestedFull time- ...AI Engineer – Decision & Optimization Systems About Gallatin At Gallatin, we are rebuilding defense logistics for the warfighters of the United States and allied forces. We take an AI-first approach to modernizing how materiel, fuel, and equipment move from factory...SuggestedFull time
- ...AI Engineer – Routing & Network Optimizatio About Gallatin At Gallatin, we are rebuilding defense logistics for the warfighters... ...Role We’re looking for an AI Engineer – Routing & Network Optimization to own how supplies, vehicles, and materiel move across...SuggestedFull time
- ...Texas Sports Academy is on the lookout for a Senior AI Engineer specializing in LLM (Large Language Model) Systems and RAG (Retrieval-Augmented Generation) Optimization. As we continue to push the boundaries of sports technology, your role will be pivotal in developing...SuggestedRemote jobFull time
- Tavus in San Francisco is seeking an experienced Research Scientist/Engineer focused on model optimization to join our core AI team. We thrive in a startup environment, prioritizing independence, calculated risks, and rapid progress to push human-AI interaction forward...
- Framework Ventures is seeking an experienced AI model serving engineer to advance inference pipelines across edge and cloud environments in the... ...with cross‑functional teams to benchmark performance, optimize memory usage, and push the boundaries of real‑world AI capabilities...
- Intel is seeking a Software Research Engineer/Scientist in the United States (Arizona, Phoenix... ...software applications, algorithms, and AI solutions. You will prototype concepts,... ...of findings with a focus on scalable optimization models and impactful technology #J-1880...
$229.9k - $262.4k
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-...Full timePart timeLocal area- ...is the moat — every competitor has humans optimizing campaigns manually. You are building the... ...intelligence layer that compounds. The Platform engineer creates the tools, you create the... ...systems, or automated recommendation loops. AI-first development workflow and ability to...Full time
$45 - $60 per hour
...purpose neural processing unit (GPNPU) architecture. Quadric’s co-optimized software and hardware is targeted to run neural network (NN)... ...receive hands-on experience working alongside industry experts in AI and semiconductor technology, with access to mentorship and meaningful...Hourly payTemporary workInternshipWork at officeRelocation- ...generation of autonomous system intelligence. As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient, production-... ...to minimize latency and maximize memory bandwidth on AI accelerators. Write production-level, low latency, and...Temporary workRelocation package
- ...Job Description Job Description NextGen is seeking a highly motivated and technically skilled Edge AI/Model Optimization Engineer to support the deployment, optimization, and sustainment of AI and agentic AI capabilities within edge and tactical computing environments...Local area
- ...Laboratory of the Rockies (NLR) in Golden, CO seeks an ML and Optimization Engineer to design, develop, and test software supporting energy research and high-performance computing. You will work across AI/ML and optimization to build production-ready systems that operate...
- Forrest T. Jones & Company seeks an AI Process Optimization Engineer to identify and improve business processes using AI, automation, and data analytics. You will work with stakeholders, IT, and data professionals to evaluate AI opportunities, develop intelligent process...
- A rapidly scaling technology company is seeking a Forward Deployed Engineer in New York. The role involves building, optimizing AI agents and working closely with customer data. You should have 2+ years in relevant engineering roles and be proficient in SQL. Excellent communication...
$100k
...position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Optimization Engineer based in United States. This is a fully remote opportunity focused on improving the performance, scalability, and...Remote jobPermanent employmentFull timeH1b$200k - $350k
...Figure is an AI Robotics company developing a general purpose humanoid. Our humanoid... ...Our Helix team is looking for Perception Engineers to empower Figure humanoid robots to perform... ...cross-functionally to evolve and optimize our autonomy stack. Requirements:...Full timeWork at office- ...We are looking for a passionate and hands-on AI Builder to join our AI Engineering team. This role is ideal for someone with 1–2 years of experience... .... Participate in testing, model evaluation, prompt optimization, and performance tuning. Document solutions and...Full time
- ...strong background in machine learning, artificial intelligence, and data science, the AI engineer will design, develop, and deploy AI solutions to address business challenges and optimize operations. Work closely with cross-functional teams to integrate AI models into...Full timeWork experience placementWork at office
- ...one of the hardest problems in enterprise AI: AI models are generic but company... ...London. We're a deeply technical team of engineers, AI researchers, and strategists with a... ...time: human-in-the-loop feedback, prompt optimization and context engineering You have experience...Full time
- ...About the role We’re seeking an experienced engineer to deploy enterprise-grade AI solutions, focusing on Retrieval-Augmented Generation (RAG) pipelines... ...solutions for emerging needs. Responsibilities: Optimize and support solutions within strategic accounts on the...Full time
- ...Founding AI Engineer (AI + Production) New York City (5 days on-site) · Top of market + equity + benefits TL;DR: Build AI that accelerates... ...model for thermal analysis; pair with fullstack engineer to optimize inference latency; attend NVIDIA collaboration session on...Full timeImmediate startWeekend work
- ...how Socure builds, deploys, and scales AI-driven identity solutions while enabling... ...Reporting to the Head of New Product Engineering, you'll join a new Internal AI Engineering... ...assessment, failure-mode analysis, and optimization of speed and accuracy of specific agentic...Full time
$195k - $230k
...About Elicit Elicit is an AI research assistant that uses language models to help researchers... ...on our mission. What is an "AI Engineer"? AI engineering is a new category of... ...+ equity, depending on your level. We're optimizing for a hire who can contribute at a L4/...Full timeTemporary workWork at officeRemote workFlexible hours- ...We’re hiring an AI Engineer to build the intelligence layer for the leading AI companion for language learning. You’ll own the core AI... ...conversations. Improve speech & conversation intelligence — optimize ASR feedback, turn-taking, and dialogue flow to make interactions...Full time
$100.8k - $168k
...for designing, developing, and deploying AI-powered solutions across the company. Focuses... ...closely with data scientists, engineers, and business leaders to deliver AI-driven... ...Expectations Develops, trains, and optimizes machine learning and deep learning models...Full timeFlexible hours- ...and increasing revenue through cutting-edge tools like AI chatbots, customized funnel optimization, targeted paid advertising, and powerful email/SMS... ...seeking a talented and experienced Machine Learning/AI Engineer with a specialized focus on OpenAI technologies. As...Full time
- ...front line needs and what is delivered. Our AI-native platform, Air Enterprise Readiness... ...We are seeking an experienced AI Engineer to join our AI Enablement team, focused on... ...identify and implement AI-driven solutions to optimize their workflows and reduce operational...Full timeWork at office
$187k - $253k
...Eigen Labs is building the coordination engine for a world run by humans and agents alike. AI is becoming abundant. Trust is not. We are building the... ...Make agents cheap (cost-aware execution, performance optimization) Make agents useful in production (not demos -...Remote jobFull timeTemporary workFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Optimization Engineer. Be the first to apply!
Related searches





