Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Systems Engineer

ByteDance

ByteDance's Volcano Ark MaaS platform invites talented engineers to advance large-model inference systems across China and internal products. You will optimize inference performance, reduce costs, and contribute to stable, scalable compute for Douyin, Toutiao, and Xigua ecosystems. Join a team that blends cutting-edge ML with practical systems, working on disaggregated inference, KV Cache, and multi-node deployments. Bachelor's/Master's in CS and strong C++/Python skills required. #J-18808-Ljbffr ByteDance

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the LLM Inference Systems Engineer in San Jose, CA vacancy
  • $150k - $350k

    ChipAgents is looking for a skilled ML Systems Engineer in San Jose, California. In this technical role, you will optimize large language model inference for our agentic AI platform, impacting chip design efficiency. Your responsibilities include implementing performance... 
    Suggested

    ChipAgents

    San Jose, CA
    3 days ago
  • NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models... 
    Suggested

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency...  ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $197.3k - $225.1k

    Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    5 days ago
  •  ..., seeks a pivotal role in AI performance evaluation and testing leadership. You will oversee the system test team focused on delivering high-quality generative AI inference systems. With over 7 years of experience, you'll have the chance to enhance test processes and work... 
    Suggested

    Recogni

    San Jose, CA
    2 days ago
  •  ...Applied Machine Learning Ark team in San Jose seeks engineers and researchers to advance LLM MaaS platforms. You will work across model...  ...multimodal AI. You will contribute to end-to-end systems spanning training, inference, and evaluation in a fast-growing tech #J-18808... 

    ByteDance

    San Jose, CA
    5 days ago
  •  ...AI Services team is seeking Senior Machine Learning Engineers to help build scalable GenAI systems powering features across Adobe products like Firefly,...  ...Illustrator, Express, Stock, and Premiere. You will design inference pipelines, optimize models for latency, and develop... 

    Adobe Inc.

    San Jose, CA
    4 days ago
  • $128k - $256k

     ...model service platform launched by Volcano Engine. It is a leading platform in China's...  ...provides end-to-end services including model inference, evaluation, fine-tuning, AI application...  .... It provides training and inference systems for recommendation, advertising,... 
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    2 days ago
  •  ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware... 

    Jobleads-US

    Santa Clara, CA
    4 days ago
  • $160k - $198k

     ...team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and manage...  ...for large-scale AI model training and inference. You will ensure our machine learning...  ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration... 
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  •  ...talented individuals to join the Data AML machine learning platform group. The role focuses on developing the Volcano Ark training system for post-training tasks such as SFT and RL, with emphasis on elastic, multi-tenant solutions across data centers and hardware. Strong... 

    ByteDance

    San Jose, CA
    2 days ago
  • $195k - $285k

     ...the architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our hardware organization,...  ...Neural Processing Unit) to solve the industry's most massive LLM inference challenges.Required Qualifications• Education: BS/MS in Electrical... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    1 day ago
  • $190k - $237k

     ...team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage...  ...for large-scale AI model training and inference. You will ensure our machine learning...  ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration... 
    Local area

    Archer Aviation

    San Jose, CA
    2 days ago
  •  ...centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation...  ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis...  ...benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed... 

    AMD

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior...  ...language and multimodal model inference as part of NVIDIA Inference Microservices...  ...production code to TRT-LLM, NVIDIA’s open-source...  ...experience with processor and system-level performance optimization... 
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference...  ...tooling, and agentic optimization systems that improve GPU kernels at the assembly layer... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $195.2k - $361.2k

     ...fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge...  ...backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience... 
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 days ago
  • $136.3k - $231.7k

     ...into your hands without us. KLA invents systems and solutions for the manufacturing of wafers...  ...R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers...  ...infrastructure that powers model training, inference, networking, storage, data protection,... 
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    2 days ago
  •  ...days per week.The role: Senior toStaff Runtime Systems EngineerWhat You Will Do:d-Matrix is...  ...platform focusing on in-memory compute for AI inference in datacenters.This position is for runtime software engineering, working on the architecture, development, and... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...application is built. We are seeking a Senior Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward...  ...performance limits on NVIDIA GPU-accelerated systems. Your work will span models, serving software, distributed... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • TikTok USDS Joint Venture is hiring a PhD-level researcher to architect and build advanced LLM-based systems for processing high-dimensional data across diverse sources. The role emphasizes semantic understanding, including fine-grained intent classification and entity... 

    TikTok USDS Joint Venture

    San Jose, CA
    1 day ago
  • $150k - $350k

     ...edge generative AI to assist engineers in RTL design, simulation, and...  ...We are seeking an ML Systems Engineer to optimize the performance...  ...efficiency of large language model inference powering our agentic AI...  ...inference that push the limits of LLM throughput and latency. Your... 

    ChipAgents

    San Jose, CA
    4 days ago
  • ByteDance is seeking a Research Engineer / Scientist for their San Jose location. The role involves designing and optimizing distributed KV caching systems for transformer-based LLMs. Candidates must have a PhD in a relevant field and expertise in distributed systems and... 

    ByteDance

    San Jose, CA
    2 days ago
  • $136k - $212.75k

     ...are now looking for a Senior Validation Engineer in the DGX Server Product Engineering Team...  ...computing products.What you will be doing:System architecture, design, performance...  ...systems.Developing/running real world ML/LLM workload.Dynamo, TensorRT, Slurm, BCM skills... 
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute... 

    Nvidia Corporation in

    Santa Clara, CA
    3 days ago
  •  ...scalable, efficient end-to-end inference solutions aligned with...  ...single-node and distributed systems. Requirements Bachelor’...  ...computer science, electrical engineering, or a related field such as applied...  ...Preferred experience includes LLM or multimodal model training... 
    Full time
    Temporary work
    Flexible hours

    SambaNova Systems

    San Jose, CA
    13 days ago
  • $272k - $431.25k

     ...a high-throughput, low-latency inference framework for serving generative...  ...accelerators feel like a single system at datacenter scale. As large language...  ...deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...generative AI models such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging from neural architecture...  ...across NVIDIA and externally by research and engineering teams alike developing best-in-class AI models.We... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $193.3k - $261.5k

     ...JAX enabling unparalleled ML inference and training performance.The...  ...hardware-software boundary, our engineers build systematic...  ...tuning of a wide variety of LLM model families, including massive...  ...models across the stack from system level optimizations through to... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  •  ...internal AI knowledge platform that lets engineers ask natural-language questions...  ...in retrieval, RAG, agentic systems, evaluation, and efficient inference, and evaluate promising techniques...  ...evaluation. Practical experience with LLM prompting, structured generation, tool... 
    Worldwide

    Socket.dev

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Systems Engineer. Be the first to apply!