Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Systems Architect

Netpreme

Netpreme is seeking an LLM Systems Engineer to prototype and optimize advanced inference systems on cutting-edge hardware. The role blends engineering with research, guiding hardware teams on product definitions and pushing the frontiers of ML inference software. The candidate should have a strong track record in ML systems research and deep familiarity with industry-standard LLM inference systems, accelerator programming, and performance engineering. #J-18808-Ljbffr Netpreme

Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the LLM Inference Systems Architect in Santa Clara, CA vacancy
  •  ...Responsibilities Define and drive the technical strategy for inference-systems performance across workload capture, benchmarking, modeling,...  ...product or roadmap impact. ~ Preferred: direct experience with LLM inference serving, continuous batching, prompt/KV caching,... 
    Suggested
    Full time
    Temporary work
    Flexible hours

    SambaNova Systems

    San Jose, CA
    4 days ago
  • $184k - $287.5k

     .... We are seeking an expert Solutions Architect to assist customers in building AI/ML...  ...aspects related to tasks like large scale LLM training and inference.Conducting regular technical customer...  ...performance issues for both AI and systems performance.What we need to see:BS/MS... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    12 hours ago
  • $152k - $241.5k

     ...innovators to roll out and enhance AI inference solutions at scale,...  ...and Kubernetes. As a Solutions Architect focused on inference, you’ll...  ...inference pipelines using TensorRT-LLM, vLLM, SGLang, and other...  ...of disaggregated inference systems and resolving complex issues.... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $174.72k - $295.68k

     ...mission is to build strong foundation for LLM deployment and quality sign-off for next-...  ...model fine tuning, PTQ, QAT, on-vehicle inference and related fields.Key ResponsibilitiesDevelop...  ...to work effectively across research, systems, infrastructure, and product teams.... 
    Suggested
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $119.25k - $150.85k

     ...scale. Within GM AV, the Model Deployment & Inference Solutions team deploys machine learning...  ..., data structures, algorithms, operating systems, computer architecture) and solid coding...  ...Kubeflow. Experience building agentic or LLM-powered tools or workflows. Open-source contributions... 
    Suggested
    Full time
    Internship
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...development, and deployment of advanced AI agents and agentic systems. Architect and implement complex multi-agent systems, including...  ...of fine-tuning strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in agentic AI or related... 
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    12 hours ago
  • $232k - $310k

     ...defining the next era of agentic talent systems.What sets Eightfold apart is not...  ...AI agents and agentic systems.Architect and implement complex multi-agent...  ...tuning strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills & Experience:Research... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...seeking an outstanding Solutions Architect, Foundation Models to join...  ...models, and production inference! In this role, you will act as...  ..., Nemotron, Dynamo, TensorRT-LLM, Triton, NIMs, and related tooling...  ..., and large-scale inference systems, with hands-on expertise in fine... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is looking for an ambitious and forward-thinking solution architect to help in the enablement of Network Industry Software Vendors...  ...). These ISVs are developing a network stack for distributed inference which will be used to orchestrate wide area networks/cloud... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    3 days ago
  • $195.2k - $262.2k

     ...engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems...  ...sits at the intersection of distributed systems, GPU performance, model training...  ...during production rollouts. Optimize LLM and VLM endpoints for latency, throughput... 
    Full time
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    2 days ago
  • $150k - $230k

     ...information powered by advanced AI, recommendation systems, and adtech.Recognized by Fast Company as...  ...-ready code.RequirementsHands-on LLM post-training experience. You have...  ...Face TRL/Accelerate, DeepSpeed or FSDP, and inference engines like vLLM.Solid understanding of... 
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    3 days ago
  •  ...SolutionsJob Title: ML Engineer with LLMLocation: Sunnyvale, CA(onsite)Job Description:6-8 years of experience in machine learning and LLM, with a proven track record in image processing and analysis.Development and optimization of Computer Vision algorithms and ML models... 

    SRI Tech

    Sunnyvale, CA
    3 days ago
  •  ...at scale, delivering measurable business outcomes through advanced AI, data, and engineering capabilities. A Technical Architect — AI Systems, Inference & Platform Internals is sought to design, scale, and optimize systems powering ChatGPT, OpenAI API, Codex, and agentic... 

    Worky

    Mountain View, CA
    3 days ago
  •  ...is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines for... 

    NVIDIA

    Santa Clara, CA
    12 hours ago
  • $208k - $327.75k

    NVIDIA Enterprise Platforms Group is seeking a Senior System Architect to define, design, and validate enterprise AI factory reference architectures...  ...-cloud deployments optimized for training, fine-tuning, inference, agentic AI, physical AI, and HPC workloadsEvaluate... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $208k - $416k

     ...ever.The DRAM Architecture team is responsible for defining the system‑level and silicon‑level architectures that enable Micron’s memory...  ...domains (e.g., silicon, packaging, systems, or platforms).• AI/LLM/ML knowledge and use for productivity.Preferred Qualifications:•... 
    Full time
    Local area
    Immediate start

    Micron

    San Jose, CA
    12 hours ago
  • $184k - $356.5k

     ...We are looking for an architect that wants to change how the AI inference industry works. With agent adoption taking off...  ...developers how to scale their businesses, systems and infrastructure. This role...  ...pipelines using TensorRT-LLM, vLLM, SGLang, and other backends... 
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $184.7k - $324.8k

     ...Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara,...  ...powered with the largest foundation models.Our systems serve billions of queries daily across...  ...end to end. Hands-on experience with LLM inference stacks. Working knowledge of GPU... 
    Worldwide
    Relocation

    Apple Inc.

    Santa Clara, CA
    4 days ago
  • $215k - $285k

     ...industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as...  ...training runs spanning many nodes, and high-throughput batch inference sweeping petabytes of real-world autonomy logs for auto-labeling... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    NLP PEOPLE

    Sunnyvale, CA
    4 days ago
  •  ...deploy ML models for autonomous driving. You will work on the ML deployment platform, model-optimization workflows, and on-vehicle inference with mentorship and a structured onboarding plan. You’ll collaborate with cross-functional teams across kernels, compilers, and... 

    General Motors

    Sunnyvale, CA
    1 day ago
  • $246.5k

     ...Advertisers, Publishers, and Roku. The systems and solutions span multiple disciplines...  ...core of this is our Machine Learning and Inference Platform that powers the entire landscape...  .... About the role In this role, you will architect, design, and lead the development of a SOTA... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...for an Infrastructure Solutions Architect to lead deployment and bring‑up of...  ...identifying product health trends, system bottlenecks, and operational...  ...Familiarity with modern deep learning, LLM architectures, and distributed training/inference challenges at scale.Your base... 
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...expert AV and GenAI Solutions Architect to help assist customers with...  ....Strong understanding of AV systems (Sensors, dynamics, perception...  ...etc.Experience in deploying LLM models at scale on mainstream...  ...record to profile and optimize inference latency and throughput,... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $148k - $235.75k

     ...NVIDIA's latest products and systems; gathering install and bring-...  ...an ambitious Senior Solutions Architect to drive validation of NVIDIA...  ..., running and debugging AI/LLM workloads and benchmarks on Linux...  ...with LLM training and/or inference workflows using frameworks such... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...seeking outstanding AI Solutions Architects to assist and support...  ...infrastructure for training, fine-tuning, inference, retrieval, and agentic AI...  ...advisor for accelerated systems architecture, GPU and...  ...platformsExperience deploying LLM training, fine-tuning, RAG, and... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...multi-step model training, and inference, all on a large scale!We are seeking a hands-on Solutions Architect with deep expertise in backend...  ...such as NIM, TensorRT-LLM, vLLM, and SGLang.Collaborate...  ...Engineering, advancing AI/ML systems from proof of concept to production... 
    Full time

    Nvidia

    Santa Clara, CA
    12 hours ago
  •  ...technology that moves the world forward.THE ROLE:As a Wired System Architect, you will be responsible for participating in standards bodies...  ...service provider broadband access, and emerging AI ML model inference applications in wired networks. THE PERSON:The ideal person... 

    AMD

    San Jose, CA
    12 hours ago
  • $262k - $364k

     ...Experience integrating Generative AI tools or Large Language Model (LLM) interfaces into workflows.Experience with architecture and...  ...including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial... 
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  • $184k - $287.5k

    NVIDIA is seeking outstanding AI Solutions Architects to assist and support customers that are...  ...and proof-of-concepts focused on inference for Generative AI and Large Language Models...  ...knowledge of the theory and practice of LLM and DL inferenceExcellent presentation,... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...is seeking an experienced Solutions Architect to be a trusted technical advisor,...  ...Dynamo, NeMo Retriever, NVIDIA Triton Inference Server, TensorRT, TensorRT-LLM, NVIDIA CUDA-XHands-on expertise...  ...as GPUs, networking, storage) and systems technology such as NCCL, DCGM, UFM,... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Systems Architect. Be the first to apply!