LLM Inference Systems Engineer
ByteDance
ByteDance's Volcano Ark MaaS platform invites talented engineers to advance large-model inference systems across China and internal products. You will optimize inference performance, reduce costs, and contribute to stable, scalable compute for Douyin, Toutiao, and Xigua ecosystems.
Join a team that blends cutting-edge ML with practical systems, working on disaggregated inference, KV Cache, and multi-node deployments. Bachelor's/Master's in CS and strong C++/Python skills required.
#J-18808-Ljbffr- NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models...Suggested
- ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You... ..., code generators, and GPU kernel innovations for LLM workloads. Join a team that designs extensible abstractions...Suggested
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work...SuggestedFull time- NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache...Suggested
- ...AI Services team is seeking Senior Machine Learning Engineers to help build scalable GenAI systems powering features across Adobe products like Firefly,... ...Illustrator, Express, Stock, and Premiere. You will design inference pipelines, optimize models for latency, and develop...Suggested
- ...NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources. You will... ...ecosystems, while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The...Local area
$128k - $256k
...model service platform launched by Volcano Engine. It is a leading platform in China's... ...provides end-to-end services including model inference, evaluation, fine-tuning, AI application... .... It provides training and inference systems for recommendation, advertising,...Temporary workLocal area- ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware...
- NVIDIA is seeking a highly capable Engineering Manager to lead the next generation of LLM/VLM inference software. You will architect and guide a team of engineers, interfacing with researchers and GPU architects to deliver production-grade software that sets the standard...
$160k - $198k
...team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and manage... ...for large-scale AI model training and inference. You will ensure our machine learning... ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration...Local area- ...talented individuals to join the Data AML machine learning platform group. The role focuses on developing the Volcano Ark training system for post-training tasks such as SFT and RL, with emphasis on elastic, multi-tenant solutions across data centers and hardware. Strong...
$195k - $285k
...the architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our hardware organization,... ...Neural Processing Unit) to solve the industry's most massive LLM inference challenges.Required Qualifications• Education: BS/MS in Electrical...3 days per week$190k - $237k
...team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage... ...for large-scale AI model training and inference. You will ensure our machine learning... ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration...Local area- ...centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation... ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed...
$152k - $241.5k
...”.NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class... ...problems for AI workloads (both inference and training) and successfully transition... ...and/or custom AI accelerator architectures.LLM Knowledge: Deep understanding of Large Language...Full time$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference... ...tooling, and agentic optimization systems that improve GPU kernels at the assembly layer...Full time$195.2k - $361.2k
...fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge... ...backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience...Full timeInternshipLocal areaImmediate startShift work$184k - $287.5k
...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior... ...language and multimodal model inference as part of NVIDIA Inference Microservices... ...production code to TRT-LLM, NVIDIA’s open-source... ...experience with processor and system-level performance optimization...Full time$150k - $350k
...edge generative AI to assist engineers in RTL design, simulation, and... ...We are seeking an ML Systems Engineer to optimize the performance... ...efficiency of large language model inference powering our agentic AI... ...inference that push the limits of LLM throughput and latency. Your...$136.3k - $231.7k
...into your hands without us. KLA invents systems and solutions for the manufacturing of wafers... ...R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers... ...infrastructure that powers model training, inference, networking, storage, data protection,...Minimum wageFull timeWork experience placementFlexible hours- ...days per week.The role: Senior toStaff Runtime Systems EngineerWhat You Will Do:d-Matrix is... ...platform focusing on in-memory compute for AI inference in datacenters.This position is for runtime software engineering, working on the architecture, development, and...3 days per week
- ...TikTok USDS Joint Venture is hiring a PhD-level researcher to architect and build advanced LLM-based systems for processing high-dimensional data across diverse sources. The role emphasizes semantic understanding, including fine-grained intent classification and entity...
$152k - $241.5k
...seeking talented and motivated engineers to join our TensorRT team in... ...-leading deep learning inference software for NVIDIA AI accelerators... ...NVIDIA TensorRT and TensorRT-LLM to supercharge inference... ...Learning Frameworks, Compilers, or System Software.Excellent problem-...Full time$136k - $212.75k
...are now looking for a Senior Validation Engineer in the DGX Server Product Engineering Team... ...computing products.What you will be doing:System architecture, design, performance... ...systems.Developing/running real world ML/LLM workload.Dynamo, TensorRT, Slurm, BCM skills...Full time- ...NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute...
$152k - $241.5k
...of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation... ...Science, Electrical/Computer Engineering, or a closely related field.4+ years... ...on autoregressive LLM serving systems, including speculative decoding or...Full time- ...scalable, efficient end-to-end inference solutions aligned with... ...single-node and distributed systems. Requirements Bachelor’... ...computer science, electrical engineering, or a related field such as applied... ...Preferred experience includes LLM or multimodal model training...Full timeTemporary workFlexible hours
- ...talented and driven ML performance engineer to optimize and scale state-... ...between deep learning and systems performance, collaborating... ...performance for large-scale AI inference. Responsibilities Bring... ...Hands-on experience with LLM or multimodal model training...Full timeTemporary workLocal areaFlexible hours
- ...forward-thinking technology firm in San Jose is looking for an LLM Engineer to design and deploy large language models that drive... ...innovative AI solutions. The role includes designing models, building inference pipelines, and collaborating with cross-functional teams....
- NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Systems Engineer. Be the first to apply!
- senior windows systems engineer San Jose, CA
- software system engineer San Jose, CA
- system validation engineer San Jose, CA
- mission system engineer San Jose, CA
- healthcare systems engineer San Jose, CA
- system verification engineer San Jose, CA
- operating system engineer San Jose, CA
- system engineer remote San Jose, CA
- application system engineer San Jose, CA
- advanced systems engineer San Jose, CA



