Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Optimization Engineer

GMI Cloud

GMI Cloud, a fast-growing AI infrastructure company, is hiring a Machine Learning Engineer, LLM Optimization to build a world-leading inference optimization team. You will drive research, validation, and productionization of advanced optimization techniques to boost latency, throughput, and cost efficiency across the inference platform. You will focus on B200-first optimization with H200 evolution, collaborating with platform and infra teams to turn new ideas into measurable customer-ready #J-18808-Ljbffr GMI Cloud

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the LLM Inference Optimization Engineer in Mountain View, CA vacancy
  •  ...Distributed LLM Inference Engineer At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software...  ...LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large... 
    Suggested
    Work at office

    Anyscale

    Palo Alto, CA
    5 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $207k - $300k

    Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure...  ...degree in Computer Science, Computer Engineering, Electrical Engineering, Applied...  ...:Experience with real world LLM inference serving environments or direct... 
    Suggested

    Google

    Mountain View, CA
    3 days ago
  • $193.3k - $261.5k

    We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational...  ...they arelocked in• Implement and optimize the inference path for large-scale multimodal...  ...models that fall outside standard LLM serving patterns — sustained low-latencyoutput... 
    Suggested
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    1 day ago
  •  ...leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter...  ...and novel deployment patterns to deep optimization of inference kernels, to building proof...  ...fabric. We are an applied research and engineering team that moves fast, ships real... 
    Suggested

    d-Matrix

    Santa Clara, CA
    4 days ago
  •  ...advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role...  ...environments.- Improve latency, throughput, and cost efficiency for LLM and multimodal model serving in production.- Analyze and... 

    AMD

    San Jose, CA
    5 days ago
  • $224k - $356.5k

    NVIDIA is seeking an Engineering Manager to lead the development of an...  ...for observing, debugging, and optimizing GenAI models deployed at...  ...visibility into model behavior, inference performance, reliability, and...  ...performance signals across large scale LLM and VLM deployments. It will... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale...  ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry...  ...building and optimizing LLM inference engines (e.g., vLLM, SGLang... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...revolution! The Algorithmic Model Optimization Team specifically focuses on...  ...such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging...  ...and externally by research and engineering teams alike developing best-in-class... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...and apply cutting-edge technologies to optimize Large Language Models (LLMs) and multimodal...  ...architectures, for highly efficient LLM inference as well as deployment across...  ...s degree in Computer Science, Computer Engineering, Applied Mathematics, Communications, Electronics... 
    Full time
    Temporary work
    Flexible hours

    NIO USA, INC

    San Jose, CA
    17 days ago
  •  ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis...  ...framework performance: Profile and optimize inference engines including vLLM, SGLang...  ...(AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed... 

    AMD

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior...  ...of performance analysis and optimization to help us squeeze every last...  ...and multimodal model inference as part of NVIDIA Inference Microservices...  ...production code to TRT-LLM, NVIDIA’s open-source inference... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...of enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application...  ...effort in Palo Alto, you will design and optimize large-scale model serving systems end-to...  ...engines (e.g., SGLang, vLLM, TensorRT-LLM) Develop custom tools for tracing, replaying... 
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  •  ...researchers, data scientists, and engineers, tackling the most...  ...As a member of the Diffusion LLM Team at MBZUAI, you will play...  ...generation. Second, we improve inference-time scaling relative to standard...  ...architectures and large-scale optimization techniques. Demonstrated... 

    Institute of Foundation Models

    Sunnyvale, CA
    9 days ago
  • $90 - $121.86 per hour

     ...Job Description Job Description LLM Research Engineer Key Responsibilities: Design, train...  ...aligned with ethical AI standards. Optimize model architecture to improve accuracy...  ...reduce latency, memory footprint, and inference time for real-time applications.... 
    Hourly pay

    Cypress HCM

    Mountain View, CA
    8 days ago
  • $90k - $180k

     ...RAPIDS (cuDF/cuML/cuGraph) and optimize end‑to‑end performance. Use...  ...reliable microservices for training/inference, vector indexing, and real-...  ...to design reviews and engineering best practices. Mentor peers...  ...feature stores. Familiarity with LLM and embedding services,... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    1 day ago
  •  ...automation with Moveworks’ Reasoning Engine and natural language...  ...infrastructure for building and serving LLM’s at Moveworks. This role will be critical in building, optimizing and scaling end-to-end machine...  ...distributed training and inference pipeline for large language... 
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    1 day ago
  • $160k - $275k

     ...compute infrastructure for large-scale LLM training and inference. We’re looking for an experienced, hands-on NPI Manufacturing Engineer to lead our AI chip, compute tray, and...  ...issuesLead SMT process development and optimization (stencil design, solder paste, reflow profile... 
    Daily paid
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    2 days ago
  •  ...industry-leading training and inference speeds; over 10 times faster than...  ...on the Cerebras Wafer-Scale Engine.We are hiring a Software Engineer to productionize and optimize our GPU serving stack, working...  ...as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an... 

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...a Senior Principal Software Engineer at JPMorganChase within the Commercial...  ..., and governance.Drives optimization of model inferencing for high...  ...and optimization using model inference servers such as Triton...  ...success architecting and deploying LLM & GNN solutions on AWS (e.g.,... 

    JP Morgan Chase

    Palo Alto, CA
    1 day ago
  • $198k - $326k

     ...work is centered on trust and optimized for culture, connection,...  ...power AI across LinkedIn. The LLM Serving team builds the critical...  ...for a Senior Staff Software Engineer with deep expertise at the intersection...  ..., and large-scale inference. This is a highly technical,... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    1 day ago
  •  ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-...  ....About The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform. This team...  ...systems you built directly. Experience optimizing latency, throughput, and efficiency in... 

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $142.8k - $274.8k

     ...Hardware, and Infrastructure Engineering (SCHIE) is the team behind...  ...industry-leading AI training and inference. The Platform Systems...  ...analysis. Hardware-Aware Workload Optimization Develop and optimize kernels...  ...: HPL/HPC benchmarks LLM training workloads Transformer... 
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Mountain View, CA
    3 days ago
  • $165k - $242k

    Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential...  ...release-over-release. Implement advanced optimizations (e.g., micro-batch schedulers,...  ...inference frameworks (vLLM, Triton, TensorRT‑LLM, Ray Serve, TorchServe). Experience... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    1 day ago
  • $182k - $242k

     ...visual effects, rendering, and real-time inference. Our stack is engineered for speed, scale, and cost-...  ...team, focused on kernel authoring and optimization. You will write, profile, and tune the...  ...fused epilogues—on the critical path of LLM inference. Optimize for the... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    Socket

    Sunnyvale, CA
    2 days ago
  •  ...Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale...  ...a skilled and motivated Manufacturing Automation Engineer to design, implement, maintain, and optimize automated manufacturing assembly and test systems.... 
    Contract work
    Shift work
    Weekend work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage...  ...Triton Inference Server, and Ray Serve. Optimize model serving for latency, throughput,... 
    Temporary work

    2T Consulting

    Santa Clara, CA
    9 days ago
  • $160k - $275k

    What MatX Is BuildingMatX is building next-generation AI compute infrastructure for large-scale LLM training and inference. We are looking for a hands-on Mechanical Engineer to design, develop, and validate mechanical systems for our rack-scale AI platform from compute... 
    Daily paid
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    2 days ago
  • A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative platform. This role involves collaborating with cross-functional teams to push the performance limits of AI systems.... 

    SambaNova

    Palo Alto, CA
    5 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across...  ...frameworks, and performance engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Optimization Engineer. Be the first to apply!