LLM Inference Systems Engineer
ByteDance
ByteDance's Volcano Ark MaaS platform invites talented engineers to advance large-model inference systems across China and internal products. You will optimize inference performance, reduce costs, and contribute to stable, scalable compute for Douyin, Toutiao, and Xigua ecosystems. Join a team that blends cutting-edge ML with practical systems, working on disaggregated inference, KV Cache, and multi-node deployments. Bachelor's/Master's in CS and strong C++/Python skills required. #J-18808-Ljbffr ByteDance
$150k - $350k
ChipAgents is looking for a skilled ML Systems Engineer in San Jose, California. In this technical role, you will optimize large language model inference for our agentic AI platform, impacting chip design efficiency. Your responsibilities include implementing performance...Suggested- NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models...Suggested
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work...SuggestedFull time$197.3k - $225.1k
Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer...SuggestedFull timePart timeLocal area- ..., seeks a pivotal role in AI performance evaluation and testing leadership. You will oversee the system test team focused on delivering high-quality generative AI inference systems. With over 7 years of experience, you'll have the chance to enhance test processes and work...Suggested
- ...Applied Machine Learning Ark team in San Jose seeks engineers and researchers to advance LLM MaaS platforms. You will work across model... ...multimodal AI. You will contribute to end-to-end systems spanning training, inference, and evaluation in a fast-growing tech #J-18808...
- ...AI Services team is seeking Senior Machine Learning Engineers to help build scalable GenAI systems powering features across Adobe products like Firefly,... ...Illustrator, Express, Stock, and Premiere. You will design inference pipelines, optimize models for latency, and develop...
$128k - $256k
...model service platform launched by Volcano Engine. It is a leading platform in China's... ...provides end-to-end services including model inference, evaluation, fine-tuning, AI application... .... It provides training and inference systems for recommendation, advertising,...Temporary workLocal area- ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware...
$160k - $198k
...team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and manage... ...for large-scale AI model training and inference. You will ensure our machine learning... ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration...Local area- ...talented individuals to join the Data AML machine learning platform group. The role focuses on developing the Volcano Ark training system for post-training tasks such as SFT and RL, with emphasis on elastic, multi-tenant solutions across data centers and hardware. Strong...
$195k - $285k
...the architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our hardware organization,... ...Neural Processing Unit) to solve the industry's most massive LLM inference challenges.Required Qualifications• Education: BS/MS in Electrical...3 days per week$190k - $237k
...team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage... ...for large-scale AI model training and inference. You will ensure our machine learning... ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration...Local area- ...centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation... ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed...
$184k - $287.5k
...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior... ...language and multimodal model inference as part of NVIDIA Inference Microservices... ...production code to TRT-LLM, NVIDIA’s open-source... ...experience with processor and system-level performance optimization...Full time$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference... ...tooling, and agentic optimization systems that improve GPU kernels at the assembly layer...Full time$195.2k - $361.2k
...fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge... ...backgroundStrong in C++ and/or Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience...Full timeInternshipLocal areaImmediate startShift work$136.3k - $231.7k
...into your hands without us. KLA invents systems and solutions for the manufacturing of wafers... ...R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers... ...infrastructure that powers model training, inference, networking, storage, data protection,...Minimum wageFull timeWork experience placementFlexible hours- ...days per week.The role: Senior toStaff Runtime Systems EngineerWhat You Will Do:d-Matrix is... ...platform focusing on in-memory compute for AI inference in datacenters.This position is for runtime software engineering, working on the architecture, development, and...3 days per week
$184k - $287.5k
...application is built. We are seeking a Senior Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward... ...performance limits on NVIDIA GPU-accelerated systems. Your work will span models, serving software, distributed...Full time- TikTok USDS Joint Venture is hiring a PhD-level researcher to architect and build advanced LLM-based systems for processing high-dimensional data across diverse sources. The role emphasizes semantic understanding, including fine-grained intent classification and entity...
$150k - $350k
...edge generative AI to assist engineers in RTL design, simulation, and... ...We are seeking an ML Systems Engineer to optimize the performance... ...efficiency of large language model inference powering our agentic AI... ...inference that push the limits of LLM throughput and latency. Your...- ByteDance is seeking a Research Engineer / Scientist for their San Jose location. The role involves designing and optimizing distributed KV caching systems for transformer-based LLMs. Candidates must have a PhD in a relevant field and expertise in distributed systems and...
$136k - $212.75k
...are now looking for a Senior Validation Engineer in the DGX Server Product Engineering Team... ...computing products.What you will be doing:System architecture, design, performance... ...systems.Developing/running real world ML/LLM workload.Dynamo, TensorRT, Slurm, BCM skills...Full time- NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute...
- ...scalable, efficient end-to-end inference solutions aligned with... ...single-node and distributed systems. Requirements Bachelor’... ...computer science, electrical engineering, or a related field such as applied... ...Preferred experience includes LLM or multimodal model training...Full timeTemporary workFlexible hours
$272k - $431.25k
...a high-throughput, low-latency inference framework for serving generative... ...accelerators feel like a single system at datacenter scale. As large language... ...deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap...Full timeLocal areaRemote work$184k - $287.5k
...generative AI models such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging from neural architecture... ...across NVIDIA and externally by research and engineering teams alike developing best-in-class AI models.We...Full time$193.3k - $261.5k
...JAX enabling unparalleled ML inference and training performance.The... ...hardware-software boundary, our engineers build systematic... ...tuning of a wide variety of LLM model families, including massive... ...models across the stack from system level optimizations through to...Work experience placementInternshipLocal areaFlexible hours- ...internal AI knowledge platform that lets engineers ask natural-language questions... ...in retrieval, RAG, agentic systems, evaluation, and efficient inference, and evaluate promising techniques... ...evaluation. Practical experience with LLM prompting, structured generation, tool...Worldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Systems Engineer. Be the first to apply!
- system engineer remote San Jose, CA
- senior windows systems engineer San Jose, CA
- systems engineer intern San Jose, CA
- senior linux systems engineer San Jose, CA
- ground systems engineer San Jose, CA
- advanced systems engineer San Jose, CA
- system verification engineer San Jose, CA
- wireless systems engineer San Jose, CA
- senior staff systems engineer San Jose, CA
- mission system engineer San Jose, CA


