Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Systems Engineer

ByteDance

ByteDance's Volcano Ark MaaS platform invites talented engineers to advance large-model inference systems across China and internal products. You will optimize inference performance, reduce costs, and contribute to stable, scalable compute for Douyin, Toutiao, and Xigua ecosystems.

Join a team that blends cutting-edge ML with practical systems, working on disaggregated inference, KV Cache, and multi-node deployments. Bachelor's/Master's in CS and strong C++/Python skills required.

#J-18808-Ljbffr
Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the LLM Inference Systems Engineer in San Jose, CA vacancy
  • NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models... 
    Suggested

    NVIDIA

    Santa Clara, CA
    2 days ago
  •  ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You...  ..., code generators, and GPU kernel innovations for LLM workloads. Join a team that designs extensible abstractions... 
    Suggested

    NVIDIA Corporation

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency...  ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache... 
    Suggested

    NVIDIA

    Santa Clara, CA
    4 days ago
  •  ...AI Services team is seeking Senior Machine Learning Engineers to help build scalable GenAI systems powering features across Adobe products like Firefly,...  ...Illustrator, Express, Stock, and Premiere. You will design inference pipelines, optimize models for latency, and develop... 
    Suggested

    Adobe Inc.

    San Jose, CA
    14 hours ago
  •  ...NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources. You will...  ...ecosystems, while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The... 
    Local area

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • $128k - $256k

     ...model service platform launched by Volcano Engine. It is a leading platform in China's...  ...provides end-to-end services including model inference, evaluation, fine-tuning, AI application...  .... It provides training and inference systems for recommendation, advertising,... 
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    14 hours ago
  •  ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware... 

    Jobleads-US

    Santa Clara, CA
    4 days ago
  • NVIDIA is seeking a highly capable Engineering Manager to lead the next generation of LLM/VLM inference software. You will architect and guide a team of engineers, interfacing with researchers and GPU architects to deliver production-grade software that sets the standard... 

    NVIDIA AI

    Santa Clara, CA
    4 days ago
  • $160k - $198k

     ...team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and manage...  ...for large-scale AI model training and inference. You will ensure our machine learning...  ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration... 
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  •  ...talented individuals to join the Data AML machine learning platform group. The role focuses on developing the Volcano Ark training system for post-training tasks such as SFT and RL, with emphasis on elastic, multi-tenant solutions across data centers and hardware. Strong... 

    ByteDance

    San Jose, CA
    14 hours ago
  • $195k - $285k

     ...the architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our hardware organization,...  ...Neural Processing Unit) to solve the industry's most massive LLM inference challenges.Required Qualifications• Education: BS/MS in Electrical... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    1 day ago
  • $190k - $237k

     ...team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage...  ...for large-scale AI model training and inference. You will ensure our machine learning...  ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration... 
    Local area

    Archer Aviation

    San Jose, CA
    2 days ago
  •  ...centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation...  ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis...  ...benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed... 

    AMD

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...”.NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class...  ...problems for AI workloads (both inference and training) and successfully transition...  ...and/or custom AI accelerator architectures.LLM Knowledge: Deep understanding of Large Language... 
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference...  ...tooling, and agentic optimization systems that improve GPU kernels at the assembly layer... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $195.2k - $361.2k

     ...fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge...  ...backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience... 
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior...  ...language and multimodal model inference as part of NVIDIA Inference Microservices...  ...production code to TRT-LLM, NVIDIA’s open-source...  ...experience with processor and system-level performance optimization... 
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $150k - $350k

     ...edge generative AI to assist engineers in RTL design, simulation, and...  ...We are seeking an ML Systems Engineer to optimize the performance...  ...efficiency of large language model inference powering our agentic AI...  ...inference that push the limits of LLM throughput and latency. Your... 

    ChipAgents

    San Jose, CA
    3 days ago
  • $136.3k - $231.7k

     ...into your hands without us. KLA invents systems and solutions for the manufacturing of wafers...  ...R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers...  ...infrastructure that powers model training, inference, networking, storage, data protection,... 
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    2 days ago
  •  ...days per week.The role: Senior toStaff Runtime Systems EngineerWhat You Will Do:d-Matrix is...  ...platform focusing on in-memory compute for AI inference in datacenters.This position is for runtime software engineering, working on the architecture, development, and... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    3 days ago
  •  ...TikTok USDS Joint Venture is hiring a PhD-level researcher to architect and build advanced LLM-based systems for processing high-dimensional data across diverse sources. The role emphasizes semantic understanding, including fine-grained intent classification and entity... 

    TikTok USDS Joint Venture

    San Jose, CA
    14 hours ago
  • $152k - $241.5k

     ...seeking talented and motivated engineers to join our TensorRT team in...  ...-leading deep learning inference software for NVIDIA AI accelerators...  ...NVIDIA TensorRT and TensorRT-LLM to supercharge inference...  ...Learning Frameworks, Compilers, or System Software.Excellent problem-... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $136k - $212.75k

     ...are now looking for a Senior Validation Engineer in the DGX Server Product Engineering Team...  ...computing products.What you will be doing:System architecture, design, performance...  ...systems.Developing/running real world ML/LLM workload.Dynamo, TensorRT, Slurm, BCM skills... 
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  •  ...NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute... 

    Nvidia Corporation in

    Santa Clara, CA
    14 hours ago
  • $152k - $241.5k

     ...of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation...  ...Science, Electrical/Computer Engineering, or a closely related field.4+ years...  ...on autoregressive LLM serving systems, including speculative decoding or... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...scalable, efficient end-to-end inference solutions aligned with...  ...single-node and distributed systems. Requirements Bachelor’...  ...computer science, electrical engineering, or a related field such as applied...  ...Preferred experience includes LLM or multimodal model training... 
    Full time
    Temporary work
    Flexible hours

    SambaNova Systems

    San Jose, CA
    3 days ago
  •  ...talented and driven ML performance engineer to optimize and scale state-...  ...between deep learning and systems performance, collaborating...  ...performance for large-scale AI inference. Responsibilities Bring...  ...Hands-on experience with LLM or multimodal model training... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Doist

    San Jose, CA
    1 day ago
  •  ...forward-thinking technology firm in San Jose is looking for an LLM Engineer to design and deploy large language models that drive...  ...innovative AI solutions. The role includes designing models, building inference pipelines, and collaborating with cross-functional teams.... 

    Take2 Consulting, LLC

    San Jose, CA
    14 hours ago
  • NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines... 

    NVIDIA

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Systems Engineer. Be the first to apply!