LLM Inference Optimization Engineer
GMI Cloud
GMI Cloud, a fast-growing AI infrastructure company, is hiring a Machine Learning Engineer, LLM Optimization to build a world-leading inference optimization team. You will drive research, validation, and productionization of advanced optimization techniques to boost latency, throughput, and cost efficiency across the inference platform. You will focus on B200-first optimization with H200 evolution, collaborating with platform and infra teams to turn new ideas into measurable customer-ready #J-18808-Ljbffr GMI Cloud
- ...Distributed LLM Inference Engineer At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software... ...LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large...SuggestedWork at office
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time$207k - $300k
Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure... ...degree in Computer Science, Computer Engineering, Electrical Engineering, Applied... ...:Experience with real world LLM inference serving environments or direct...Suggested$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational... ...they arelocked in• Implement and optimize the inference path for large-scale multimodal... ...models that fall outside standard LLM serving patterns — sustained low-latencyoutput...SuggestedInternshipLocal areaFlexible hours- ...leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter... ...and novel deployment patterns to deep optimization of inference kernels, to building proof... ...fabric. We are an applied research and engineering team that moves fast, ships real...Suggested
- ...advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role... ...environments.- Improve latency, throughput, and cost efficiency for LLM and multimodal model serving in production.- Analyze and...
$224k - $356.5k
NVIDIA is seeking an Engineering Manager to lead the development of an... ...for observing, debugging, and optimizing GenAI models deployed at... ...visibility into model behavior, inference performance, reliability, and... ...performance signals across large scale LLM and VLM deployments. It will...Full time$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale... ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry... ...building and optimizing LLM inference engines (e.g., vLLM, SGLang...Full time$184k - $287.5k
...revolution! The Algorithmic Model Optimization Team specifically focuses on... ...such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging... ...and externally by research and engineering teams alike developing best-in-class...Full time- ...and apply cutting-edge technologies to optimize Large Language Models (LLMs) and multimodal... ...architectures, for highly efficient LLM inference as well as deployment across... ...s degree in Computer Science, Computer Engineering, Applied Mathematics, Communications, Electronics...Full timeTemporary workFlexible hours
- ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...framework performance: Profile and optimize inference engines including vLLM, SGLang... ...(AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed...
$184k - $287.5k
...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior... ...of performance analysis and optimization to help us squeeze every last... ...and multimodal model inference as part of NVIDIA Inference Microservices... ...production code to TRT-LLM, NVIDIA’s open-source inference...Full time- ...of enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application... ...effort in Palo Alto, you will design and optimize large-scale model serving systems end-to... ...engines (e.g., SGLang, vLLM, TensorRT-LLM) Develop custom tools for tracing, replaying...Permanent employmentTemporary workRemote workWorldwideWeekend work
- ...researchers, data scientists, and engineers, tackling the most... ...As a member of the Diffusion LLM Team at MBZUAI, you will play... ...generation. Second, we improve inference-time scaling relative to standard... ...architectures and large-scale optimization techniques. Demonstrated...
$90 - $121.86 per hour
...Job Description Job Description LLM Research Engineer Key Responsibilities: Design, train... ...aligned with ethical AI standards. Optimize model architecture to improve accuracy... ...reduce latency, memory footprint, and inference time for real-time applications....Hourly pay$90k - $180k
...RAPIDS (cuDF/cuML/cuGraph) and optimize end‑to‑end performance. Use... ...reliable microservices for training/inference, vector indexing, and real-... ...to design reviews and engineering best practices. Mentor peers... ...feature stores. Familiarity with LLM and embedding services,...Full timeTemporary workPart time- ...automation with Moveworks’ Reasoning Engine and natural language... ...infrastructure for building and serving LLM’s at Moveworks. This role will be critical in building, optimizing and scaling end-to-end machine... ...distributed training and inference pipeline for large language...Work at officeRemote workFlexible hours
$160k - $275k
...compute infrastructure for large-scale LLM training and inference. We’re looking for an experienced, hands-on NPI Manufacturing Engineer to lead our AI chip, compute tray, and... ...issuesLead SMT process development and optimization (stencil design, solder paste, reflow profile...Daily paidFull timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours- ...industry-leading training and inference speeds; over 10 times faster than... ...on the Cerebras Wafer-Scale Engine.We are hiring a Software Engineer to productionize and optimize our GPU serving stack, working... ...as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an...
- ...a Senior Principal Software Engineer at JPMorganChase within the Commercial... ..., and governance.Drives optimization of model inferencing for high... ...and optimization using model inference servers such as Triton... ...success architecting and deploying LLM & GNN solutions on AWS (e.g.,...
$198k - $326k
...work is centered on trust and optimized for culture, connection,... ...power AI across LinkedIn. The LLM Serving team builds the critical... ...for a Senior Staff Software Engineer with deep expertise at the intersection... ..., and large-scale inference. This is a highly technical,...For contractorsWork at officeFlexible hours- ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-... ....About The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform. This team... ...systems you built directly. Experience optimizing latency, throughput, and efficiency in...
$142.8k - $274.8k
...Hardware, and Infrastructure Engineering (SCHIE) is the team behind... ...industry-leading AI training and inference. The Platform Systems... ...analysis. Hardware-Aware Workload Optimization Develop and optimize kernels... ...: HPL/HPC benchmarks LLM training workloads Transformer...Ongoing contractWork at officeLocal areaWorldwide3 days per week$165k - $242k
Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential... ...release-over-release. Implement advanced optimizations (e.g., micro-batch schedulers,... ...inference frameworks (vLLM, Triton, TensorRT‑LLM, Ray Serve, TorchServe). Experience...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$182k - $242k
...visual effects, rendering, and real-time inference. Our stack is engineered for speed, scale, and cost-... ...team, focused on kernel authoring and optimization. You will write, profile, and tune the... ...fused epilogues—on the critical path of LLM inference. Optimize for the...Permanent employmentTemporary workCasual workWork at officeFlexible hours- ...Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale... ...a skilled and motivated Manufacturing Automation Engineer to design, implement, maintain, and optimize automated manufacturing assembly and test systems....Contract workShift workWeekend work
- ...We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage... ...Triton Inference Server, and Ray Serve. Optimize model serving for latency, throughput,...Temporary work
$160k - $275k
What MatX Is BuildingMatX is building next-generation AI compute infrastructure for large-scale LLM training and inference. We are looking for a hands-on Mechanical Engineer to design, develop, and validate mechanical systems for our rack-scale AI platform from compute...Daily paidFull timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours- A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative platform. This role involves collaborating with cross-functional teams to push the performance limits of AI systems....
$272k - $431.25k
...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across... ...frameworks, and performance engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Optimization Engineer. Be the first to apply!


