LLM Inference Systems Architect
Netpreme
Netpreme is seeking an LLM Systems Engineer to prototype and optimize advanced inference systems on cutting-edge hardware. The role blends engineering with research, guiding hardware teams on product definitions and pushing the frontiers of ML inference software. The candidate should have a strong track record in ML systems research and deep familiarity with industry-standard LLM inference systems, accelerator programming, and performance engineering. #J-18808-Ljbffr Netpreme
- ...Responsibilities Define and drive the technical strategy for inference-systems performance across workload capture, benchmarking, modeling,... ...product or roadmap impact. ~ Preferred: direct experience with LLM inference serving, continuous batching, prompt/KV caching,...SuggestedFull timeTemporary workFlexible hours
$184k - $287.5k
.... We are seeking an expert Solutions Architect to assist customers in building AI/ML... ...aspects related to tasks like large scale LLM training and inference.Conducting regular technical customer... ...performance issues for both AI and systems performance.What we need to see:BS/MS...SuggestedFull time$152k - $241.5k
...innovators to roll out and enhance AI inference solutions at scale,... ...and Kubernetes. As a Solutions Architect focused on inference, you’ll... ...inference pipelines using TensorRT-LLM, vLLM, SGLang, and other... ...of disaggregated inference systems and resolving complex issues....SuggestedFull time$174.72k - $295.68k
...mission is to build strong foundation for LLM deployment and quality sign-off for next-... ...model fine tuning, PTQ, QAT, on-vehicle inference and related fields.Key ResponsibilitiesDevelop... ...to work effectively across research, systems, infrastructure, and product teams....SuggestedFull time$119.25k - $150.85k
...scale. Within GM AV, the Model Deployment & Inference Solutions team deploys machine learning... ..., data structures, algorithms, operating systems, computer architecture) and solid coding... ...Kubeflow. Experience building agentic or LLM-powered tools or workflows. Open-source contributions...SuggestedFull timeInternshipLocal areaWork from homeRelocation packageFlexible hours- ...development, and deployment of advanced AI agents and agentic systems. Architect and implement complex multi-agent systems, including... ...of fine-tuning strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in agentic AI or related...Full timeWork experience placement
$232k - $310k
...defining the next era of agentic talent systems.What sets Eightfold apart is not... ...AI agents and agentic systems.Architect and implement complex multi-agent... ...tuning strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills & Experience:Research...Work experience placementWork at officeRemote workFlexible hours3 days per week$152k - $241.5k
...seeking an outstanding Solutions Architect, Foundation Models to join... ...models, and production inference! In this role, you will act as... ..., Nemotron, Dynamo, TensorRT-LLM, Triton, NIMs, and related tooling... ..., and large-scale inference systems, with hands-on expertise in fine...Full time$184k - $287.5k
NVIDIA is looking for an ambitious and forward-thinking solution architect to help in the enablement of Network Industry Software Vendors... ...). These ISVs are developing a network stack for distributed inference which will be used to orchestrate wide area networks/cloud...Full timeWork experience placement$195.2k - $262.2k
...engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems... ...sits at the intersection of distributed systems, GPU performance, model training... ...during production rollouts. Optimize LLM and VLM endpoints for latency, throughput...Full timeTemporary workImmediate startRemote work$150k - $230k
...information powered by advanced AI, recommendation systems, and adtech.Recognized by Fast Company as... ...-ready code.RequirementsHands-on LLM post-training experience. You have... ...Face TRL/Accelerate, DeepSpeed or FSDP, and inference engines like vLLM.Solid understanding of...Full timeLocal areaWork from home- ...SolutionsJob Title: ML Engineer with LLMLocation: Sunnyvale, CA(onsite)Job Description:6-8 years of experience in machine learning and LLM, with a proven track record in image processing and analysis.Development and optimization of Computer Vision algorithms and ML models...
- ...at scale, delivering measurable business outcomes through advanced AI, data, and engineering capabilities. A Technical Architect — AI Systems, Inference & Platform Internals is sought to design, scale, and optimize systems powering ChatGPT, OpenAI API, Codex, and agentic...
- ...is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines for...
$208k - $327.75k
NVIDIA Enterprise Platforms Group is seeking a Senior System Architect to define, design, and validate enterprise AI factory reference architectures... ...-cloud deployments optimized for training, fine-tuning, inference, agentic AI, physical AI, and HPC workloadsEvaluate...Full time$208k - $416k
...ever.The DRAM Architecture team is responsible for defining the system‑level and silicon‑level architectures that enable Micron’s memory... ...domains (e.g., silicon, packaging, systems, or platforms).• AI/LLM/ML knowledge and use for productivity.Preferred Qualifications:•...Full timeLocal areaImmediate start$184k - $356.5k
...We are looking for an architect that wants to change how the AI inference industry works. With agent adoption taking off... ...developers how to scale their businesses, systems and infrastructure. This role... ...pipelines using TensorRT-LLM, vLLM, SGLang, and other backends...Full time$184.7k - $324.8k
...Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara,... ...powered with the largest foundation models.Our systems serve billions of queries daily across... ...end to end. Hands-on experience with LLM inference stacks. Working knowledge of GPU...WorldwideRelocation$215k - $285k
...industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as... ...training runs spanning many nodes, and high-throughput batch inference sweeping petabytes of real-world autonomy logs for auto-labeling...Full timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift- ...deploy ML models for autonomous driving. You will work on the ML deployment platform, model-optimization workflows, and on-vehicle inference with mentorship and a structured onboarding plan. You’ll collaborate with cross-functional teams across kernels, compilers, and...
$246.5k
...Advertisers, Publishers, and Roku. The systems and solutions span multiple disciplines... ...core of this is our Machine Learning and Inference Platform that powers the entire landscape... .... About the role In this role, you will architect, design, and lead the development of a SOTA...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours$152k - $241.5k
...for an Infrastructure Solutions Architect to lead deployment and bring‑up of... ...identifying product health trends, system bottlenecks, and operational... ...Familiarity with modern deep learning, LLM architectures, and distributed training/inference challenges at scale.Your base...Full timeRemote workWorldwide$184k - $287.5k
...expert AV and GenAI Solutions Architect to help assist customers with... ....Strong understanding of AV systems (Sensors, dynamics, perception... ...etc.Experience in deploying LLM models at scale on mainstream... ...record to profile and optimize inference latency and throughput,...Full time$148k - $235.75k
...NVIDIA's latest products and systems; gathering install and bring-... ...an ambitious Senior Solutions Architect to drive validation of NVIDIA... ..., running and debugging AI/LLM workloads and benchmarks on Linux... ...with LLM training and/or inference workflows using frameworks such...Full timeRemote work$184k - $287.5k
...seeking outstanding AI Solutions Architects to assist and support... ...infrastructure for training, fine-tuning, inference, retrieval, and agentic AI... ...advisor for accelerated systems architecture, GPU and... ...platformsExperience deploying LLM training, fine-tuning, RAG, and...Full timeRemote work$152k - $241.5k
...multi-step model training, and inference, all on a large scale!We are seeking a hands-on Solutions Architect with deep expertise in backend... ...such as NIM, TensorRT-LLM, vLLM, and SGLang.Collaborate... ...Engineering, advancing AI/ML systems from proof of concept to production...Full time- ...technology that moves the world forward.THE ROLE:As a Wired System Architect, you will be responsible for participating in standards bodies... ...service provider broadband access, and emerging AI ML model inference applications in wired networks. THE PERSON:The ideal person...
$262k - $364k
...Experience integrating Generative AI tools or Large Language Model (LLM) interfaces into workflows.Experience with architecture and... ...including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial...Worldwide$184k - $287.5k
NVIDIA is seeking outstanding AI Solutions Architects to assist and support customers that are... ...and proof-of-concepts focused on inference for Generative AI and Large Language Models... ...knowledge of the theory and practice of LLM and DL inferenceExcellent presentation,...Full timeRemote work$152k - $241.5k
...is seeking an experienced Solutions Architect to be a trusted technical advisor,... ...Dynamo, NeMo Retriever, NVIDIA Triton Inference Server, TensorRT, TensorRT-LLM, NVIDIA CUDA-XHands-on expertise... ...as GPUs, networking, storage) and systems technology such as NCCL, DCGM, UFM,...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Systems Architect. Be the first to apply!


