Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM/VLM Inference Optimization Research Engineer

ByteDance

ByteDance Seed Infra in San Jose is seeking a Research Engineer to design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs. You will work on inference engines, serving frameworks, and end-to-end deployment pipelines, aiming to reduce latency and increase throughput. Ideal candidates will have strong C/C++ and Python skills, hands-on ML framework experience (PyTorch/TensorFlow), and background in GPU-accelerated optimization. #J-18808-Ljbffr ByteDance

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the LLM/VLM Inference Optimization Research Engineer in San Jose, CA vacancy
  • $244.8k

    Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Location San Jose Team Technology Employment Type Regular Job Code A02632A Responsibilities About the Team The Seed Infrastructures team oversees the distributed training, reinforcement learning framework... 
    Suggested
    Temporary work
    Local area

    Bytedance

    San Jose, CA
    4 days ago
  • $156k - $387.6k

    Research Engineer / Scientist - Storage for LLM Location: San Jose Team: Infrastructure Employment Type: Regular...  ...based LLMs across GPUs or nodes. Optimize low‑latency access and eviction...  ...reused embeddings. Collaborate with inference and serving teams to integrate... 
    Suggested
    Temporary work

    ByteDance

    San Jose, CA
    1 day ago
  • $115 - $131 per hour

     ...FocusKPI is seeking an LLM Research Engineer to join one of our clients, a high-tech SaaS company...  ...aligned with ethical AI standards. Optimize model architecture to improve accuracy...  ...reduce latency, memory footprint, and inference time for real-time applications. Collaborate... 
    Suggested
    Contract work
    Work at office

    FocusKPI Inc.

    Mountain View, CA
    3 days ago
  •  ...We are a dedicated research lab for building, understanding...  ...data scientists, and engineers, tackling the most...  ...of the Diffusion LLM Team at MBZUAI, you will...  ...generation. Second, we improve inference-time scaling relative...  ...and large-scale optimization techniques. Demonstrated... 
    Suggested

    Institute of Foundation Models

    Sunnyvale, CA
    4 days ago
  • $90 - $121.86 per hour

     ...Job Description Job Description LLM Research Engineer Key Responsibilities: Design, train...  ...aligned with ethical AI standards. Optimize model architecture to improve accuracy...  ...reduce latency, memory footprint, and inference time for real-time applications.... 
    Suggested
    Hourly pay

    Cypress HCM

    Mountain View, CA
    3 days ago
  • $136.8k - $259.2k

    Overview Research Engineer (LLM/ML/RL) - TikTok Ads Core ML, Ranking. Base pay range: $136,800.00/yr - $259,200.00/yr. Responsibilities Optimize efficiency across the entire advertising funnel, including Recall & Rough-sort, Fine-sort (CTR/CVR), format/creative personalization... 
    Full time
    Worldwide

    TikTok

    San Jose, CA
    1 day ago
  • $192k - $304.75k

    We are now looking for a Senior Research Engineer passionate about Generative AI inference. Are you excited to change the way...  ...team is dedicated to developing optimized inferencing technologies to...  ...and evaluate routing policies for LLM traffic to best use mixture of model... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $136.8k - $259.2k

     ...strong machine learning engineers who are excited to...  ...monetization. We are seeking Research Engineers who are...  ...detection/synthesis, optimizing performance across pre...  ...ecosystems. 3. Advance LLM-based agents with reinforcement...  ...model training and inference, exploring techniques... 
    Full time
    Temporary work
    Local area
    Flexible hours

    TikTok

    San Jose, CA
    1 day ago
  • ByteDance is seeking a Research Engineer / Scientist for their San Jose location. The role involves designing and optimizing distributed KV caching systems for transformer-based LLMs. Candidates must have a PhD in a relevant field and expertise in distributed systems and... 

    ByteDance

    San Jose, CA
    1 day ago
  • $166k - $225k

     ...missions. The Databricks AI Research organization enables companies...  ...to all. As a Sr. Research Engineer on the Scaling team, you will...  ...improvements through advanced optimization techniques including kernel fusion...  ...workflows and knowledge of LLM training dynamics including... 
    Local area

    Databricks

    Mountain View, CA
    4 days ago
  • $152k - $241.5k

     ...simulation, and virtual training.As a Research Engineer, you will contribute across model development...  ...you.What you'll be doing:Build and optimize the continuous autoregressive serving...  ...and KV-cache state management, GPU inference, frame streaming, and model integrations... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $192k - $304.75k

     ...people! We are seeking a world-class engineer to drive applied research at the intersection of AI and ASIC...  ...Formal verification, PPA prediction and optimization.Hands-on experience with LLMs, RL,...  ....Hands-on experience building LLM-based agents or AI tooling that real... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $164.6k - $313.3k

     ...looking for a driven Data/ML engineer to push the boundaries of audio...  ...collaborative and efficient research team looking for highly...  ...versioning.Run in-house models for inference over millions of audio/video...  ...ablations. Scale training and optimize throughput.Drive data... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    4 days ago
  • $224k - $356.5k

     ...Tools organization is seeking a Senior Research Engineer to join our Research team, where we build...  ...that help NVIDIA developers write, optimize, and maintain CUDA code — and that work...  ...data workFluency with the systems side of LLM-powered agents, including practical... 
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Tessera is a transformation engine: a governed, multi-agent platform...  ...and they're the reason the research is interesting. Governance:...  ..., evaluation, and inference machinery that turns a hypothesis...  ...training stack: SFT, preference optimization, and reinforcement learning... 

    Tessera Labs

    San Jose, CA
    3 days ago
  • $224k - $356.5k

     ...searching for a senior or principal engineer who specializes in building...  ...Generalist Embodied Agent Research (GEAR) group. Our team is...  ...foundation models for robotics.Optimize GPU and cluster utilization for...  ...at building large-scale LLM and multimodal LLM training infrastructure... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $244.14k - $413.16k

     ...smart connectivity.We are looking for exceptional Research Engineers / Scientists to design learning systems that...  ...Responsibilities:Reinforcement learning methods for LLM-driven agents and decision systems.Policy optimization for long-horizon reasoning and planning.Learning... 
    Full time

    XPENG Motors

    Santa Clara, CA
    4 days ago
  •  ...infrastructure for SFT, preference optimization, and reinforcement learning...  ...distributed training and inference throughput through kernels,...  .... Strong software engineering fundamentals and the ability...  ...communication and interest in taking research results into production.... 

    Tessera Labs

    San Jose, CA
    8 days ago
  • TikTok is seeking a Research Engineer/Scientist to design and implement efficient large-scale...  ..., compression, and hardware-efficient inference for scalable deployment. The role emphasizes...  ...efficient models, enabling training, optimization, and deployment at scale in a fast-... 

    TikTok

    San Jose, CA
    8 hours ago
  • $2,000 per month

     ...Machine Learning Research EngineerCupertino, CAEtched is building...  ...and scheduling algorithms to optimize for utilization, latency,...  ...metrics.Implement model-specific inference-time acceleration techniques...  ..., and greatly value engineering skills. We do not have boundaries... 

    ETCHED LLC

    Cupertino, CA
    1 day ago
  • TikTok is seeking a Machine Learning/Research Engineer Graduate for Monetization Technology-Ads Core Global, starting in...  ...will be part of the Ads Core ML Team to develop and optimize ad delivery using ML, DL, RL, and LLM techniques. Applicants should hold a PhD and be... 

    TikTok

    San Jose, CA
    3 days ago
  • $162k - $316.8k

     ...026 TikTok Technology Machine Learning/Research Engineer Graduate (Monetization Technology-Ads Core...  ...mission is to generate revenue and optimize the user experience on the world's leading...  ...technologies, including ML/DL, RL, LLM, and scaling laws in ad recommendation.... 
    Temporary work
    Local area

    Tik Tok

    San Jose, CA
    3 days ago
  • $184k - $287.5k

    We are recruiting top research engineers in the Autonomous Vehicles Research team at NVIDIA with strong expertise in software engineering...  ...graphics, natural language processing, autonomous driving, HW optimization, robotics, healthcare, and many more. NVIDIA has an open... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $174k - $252k

    Drive post-training research and engineering using reinforcement learning (RL) and supervised fine-tuning...  ...LLMs).Experience designing or running LLM evaluation benchmarks and measurement...  ...with reward modeling, preference optimization, or synthetic data generation for code... 

    Google

    Mountain View, CA
    3 days ago
  • $150k - $200k

    Santa Clara, CAUS Research and Development - Controls /Full-time /HybridPlusAI is a Physical...  ...its fast-growing teams.As a Research Engineer, you will deliver mission-critical improvements...  ...scientists, and domain experts to build optimal and data driven controls to realize... 
    Full time

    Plus.ai

    Santa Clara, CA
    1 day ago
  •  ...career. THE ROLE:We are hiring a AI Research Scientist - Infrastructure Engineer, Reinforcement Learning, to own...  ....THE PERSON:You profile before you optimize; you treat researcher time as expensive...  ...ownership of RL training infra, LLM post-training pipelines, or large-scale... 

    AMD

    Santa Clara, CA
    4 days ago
  •  ...seeking a Senior Performance Co-Design Engineer for TPU work in Sunnyvale to optimize serving performance of large language...  ...software partners. The role centers on LLM inference latency and throughput, collaborating with researchers and engineers to push the frontier of... 

    Google

    Sunnyvale, CA
    2 days ago
  • Innovation Engineer - Market Development (Optoelectronics)About the RoleWe are seeking a highly motivated Innovation Engineer - Market...  ...Power ElectronicsFamiliarity with optical path design, power optimization, and system validation setupsExperience working in global, cross... 
    Full time
    Work from home

    Vishay

    San Jose, CA
    17 hours ago
  •  ...Engineer Our Systems Technology group, a part of Roche Sequencing Solutions, is focused on creating and advancing technologies to...  ...microfluidic devices and systems Testing, characterizing and optimizing device and system performance Collecting data, documenting... 

    Tranzeal

    Santa Clara, CA
    4 days ago
  • $166k - $244k

     ...more of the following: Machine Learning Optimization (e.g., quantization, distillation),...  ...or PhD in Computer Science, Computer Engineering, or a related technical field. 5 years...  ...continue to push technology forward. Google researchers and engineers are turning science... 
    Full time
    Temporary work

    Google

    Mountain View, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM/VLM Inference Optimization Research Engineer. Be the first to apply!