Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GenAI Engineer - LLM Infrastructure & Inference Services

Temporary

2T Consulting

We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage scalable enterprise GenAI platforms across GPU infrastructure and cloud environments.

Key Responsibilities

  • Deploy, host, and manage Large Language Models (LLMs) on GPU infrastructure for production environments.
  • Build scalable, high-performance inference services using vLLM, TensorRT-LLM, Triton Inference Server, and Ray Serve.
  • Optimize model serving for latency, throughput, GPU utilization, and cost efficiency.
  • Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes.
  • Implement RAG pipelines, vector databases, and agentic AI frameworks such as LangChain and LangGraph.
  • Manage GPU infrastructure, containerization, and cloud deployments across AWS, Azure, or GCP.
  • Establish MLOps/LLMOps practices including CI/CD, model deployment, monitoring, observability, and governance.
  • Perform performance tuning, benchmarking, capacity planning, and production support for enterprise GenAI platforms.
  • Collaborate with architects, data scientists, and product teams to deliver scalable, secure, and reliable AI solutions.

Core Technologies

  • LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve
  • AI/GenAI: RAG, LangChain, LangGraph, Vector Databases
  • Development: Python, FastAPI, Microservices
  • Infrastructure: Kubernetes, Docker, GPU Infrastructure
  • Cloud: AWS, Azure, GCP
  • MLOps/LLMOps: CI/CD, Monitoring, Observability, Model Deployment, Governance
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the GenAI Engineer - LLM Infrastructure & Inference Services in Santa Clara, CA vacancy
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible...  ...experiences. Our investments in technology infrastructure and world-class talent — along with...  ...have come to love the products and services we build. Team Description:... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    5 days ago
  • $198k - $326k

     ...model training, feature engineering and serving with...  ...queries.Model Training Infrastructure: As an engineer on the...  ...complex models across LLM and Personalization models...  ..., enable GPU based inference for a large variety of...  ...locationBeing accompanied by a service dogHaving a sign... 
    Suggested
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    2 days ago
  • $229.9k - $262.4k

     ...Overview AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and...  ...experiences. Our investments in technology infrastructure and world-class talent — along...  ...come to love the products and services we build. Team Description:... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    13 days ago
  • $184k - $287.5k

     ...application is built. We are seeking a Senior Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward...  ...based on workload, hardware, model quality, and service-level objectives.Build and optimize performance-... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $159.2k - $301.6k

     ...like a startup founding engineer while having the...  ...Engineer to build Creator Services from the ground up. Creator...  ...cost-efficient inference, caching, and graceful...  ...experience building with LLM APIs and agent...  ...Colligo or similar ML infrastructure.About AdobeAdobe empowers... 
    Suggested
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    7 hours ago
  • $137.1k - $299.3k

     ...Team DoorDash’s GenAI Platform team sits...  ...builds the shared infrastructure that helps...  ...trust the quality of LLM and agent systems...  ...serving and batch inference, guardrails, and cost...  ...pipelines, backend services, and observability...  ...role is ideal for an engineer who enjoys... 
    Hourly pay
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash USA

    Sunnyvale, CA
    3 days ago
  • $307k - $427k

     ...Principal M-LLM Post-Training and Execution Software Engineer, XR Share Principal M-LLM Post...  ...with high-performance inference engines and custom AI accelerators...  ...party devices and services that combine the best...  ...the post-training infrastructure for Multimodal LLMs, including... 

    Google Inc.

    San Jose, CA
    4 days ago
  • $92k - $135k

     ...CoreWeave combines superior infrastructure performance with deep...  ...You'll Do: Join the Inference team to ship production...  ...mentorship from experienced engineers. About the role:...  ...Go/C++ for model-serving services (e.g., Triton, vLLM, TensorRT-LLM, Ray Serve). Write tests... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    28 days ago
  • $165.2k - $223.6k

     ...team at Amazon Web Services (AWS) builds AWS Neuron...  ...deep learning and GenAI workloads on Amazon...  ...unparalleled ML inference and training performance...  ...boundary, our engineers build systematic infrastructure, innovate new methods...  ...a wide variety of LLM model families,... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    1 day ago
  • $274k - $300k

     ...please visit  AI Platform Engineer – Training & Inference Saviynt's AI-powered...  ...+ H100s, the multi-engine LLM inference mesh (vLLM,...  ...LLMs • Build RL training infrastructure: define Flyte workflows for...  ...high-growth, Platform as a Service company focused on... 

    Saviynt

    Milpitas, CA
    20 days ago
  • $109k - $160k

    Software Engineer - Data Infrastructure Services Sunnyvale, CA / Bellevue, WA CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    4 days ago
  •  ...Summary: Hands-on full-stack software engineer supporting NVIDIA Test Infrastructure Development. The contractor will...  ...senior engineers, including backend services, REST APIs, database integrations,...  ...AI coding tools or experience with LLM APIs, prompt engineering, RAG/vector... 
    For contractors

    Katalyst Healthcares & Life Sciences

    Santa Clara, CA
    7 hours ago
  •  ...a leader in AI cloud infrastructure serving tens of thousands...  ...and enterprise engineering teams to design, scale...  ...evaluations across training and inference workloads to show...  ...optimization (vLLM, TensorRT-LLM) and observability...  ...APIs, gRPC, and service-oriented cloud architectures... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Corporation

    San Jose, CA
    2 days ago
  • $272k - $425.5k

    Principal Software Engineer - Large-Scale LLM Memory and Storage Systems page is...  ...-throughput, low-latency inference framework for serving generative...  ...storage, or ML systems infrastructure in C/C++ and Python, with...  ...of delivering production services.* Deep understanding of... 
    Local area
    Remote work

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $175.6k - $263.4k

     ...Computer Science, Computer Engineering, or a closely related discipline...  ...Experience developing AI/LLM-powered applications and integrations...  ...We develop scalable backend services and APIs using Python, Node....  ...engineering productivity, infrastructure management, analytics, and... 
    Full time
    Work from home

    Qualcomm

    Santa Clara, CA
    11 days ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (Inference Optimization, FM Hosting, AI Platform)At...  ...investments in technology infrastructure and world-class talent — along...  ...come to love the products and services we build.The Intelligent...  ...introduce state-of-the-art LLM optimization techniques to... 
    Full time
    Part time

    Capital One

    San Jose, CA
    5 days ago
  • $116k - $189.75k

     ...We are looking for a Software Engineer focused on bring-up, triage,...  ...of distributed training and inference workloads across NVIDIA GPU platforms...  ..., and debug distributed LLM workloads on multi-GPU and...  ...debug large-scale AI clusters, infrastructure, and end-to-end workloads.... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $184k - $287.5k

     ...looking for a Senior Software Engineer to lead the bring-up, triage,...  ...of distributed training and inference workloads across NVIDIA GPU platforms...  ...to ensure state-of-the-art LLM workloads run efficiently and...  ...of large-scale AI clusters, infrastructure, and end-to-end workloads,... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...small, highly motivated, and focused on engineering excellence. This organization is for...  ...: As part of the Network Software and Services for AI (nssAI) team at SpaceXAI, you'll...  ...used for AI training and serving customer inference queries. Implement IaC best... 
    Temporary work

    SpaceXAI

    Palo Alto, CA
    28 days ago
  • $175.6k - $263.4k

     ...Technologies, Inc. Job Area: Engineering Group, Engineering Group...  ...Summary: Qualcomm's CAD Infrastructure team develops and operates...  ...automation platforms, data services, and AI-enabled solutions...  ...tools. Integrate AI/ML and LLM technologies into internal engineering... 
    Work experience placement
    Work from home

    Qualcomm

    Santa Clara, CA
    7 days ago
  • $2,500 per month

     ...are heavily focused on inference . Backed by hundreds of...  ...and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing...  ...stack with LLM integration and a strong...  ...simulation clusters, and service endpoints. This role... 
    Work at office
    Relocation package

    Etched

    San Jose, CA
    13 days ago
  • $182k - $242k

     ...CoreWeave combines superior infrastructure performance with deep...  ...looking for a Senior Engineer to be a driving force...  ...training and inference workloads. You will own...  ...Where needed, build self-service views (Grafana, Looker...  ...Apache Spark, Trino, llm-d, vLLM, or PyTorch.... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    14 days ago
  • $262k - $364k

     ...Influence and coach a distributed team of engineers.Facilitate alignment and clarity...  ...Experience integrating generative AI tools or LLM interfaces into workflows.Preferred...  ...traffic engineering is a critical infrastructure service for Google. In this role, you will be... 
    Worldwide

    Google

    Sunnyvale, CA
    7 hours ago
  • $182k - $242k

    Senior Software Engineer, Network Services Livingston, NJ / New York, NY / Sunnyvale, CA CoreWeave is The Essential Cloud for AI™. Built for...  ...startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    4 days ago
  • $119.25k - $150.85k

    Description About the Team The HIL Platform and Services team in GM’s Autonomous Vehicle (AV) Organization is responsible for developing...  ...they reach the road.  About the Role As a Software Engineer on the HIL Platform and Services team, you will be at the... 
    Full time
    Internship
    Local area
    Work from home
    Relocation package

    General Motors

    Sunnyvale, CA
    7 days ago
  • $195k - $285k

     ...Matrix designs and manufactures purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable...  ...AWS, Azure, GCP), and customer-facing platform services. You will hire and grow the team, set technical... 
    Remote work

    d-Matrix

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...projects to critically important infrastructure. Be part of a dynamic team...  ...full lifecycle of agentic services: from prototype through evaluation...  ...improvementBuild robust LLM-powered pipelines for code generation...  ...closely with GPU SW kernel engineers to pinpoint high-impact... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...best talent. We’re hiring a Deep Learning Engineer with strong experience in generative AI,...  ...pipelines fast and iterate fasterLeverage LLM/VLM and agents in the data generation...  ...document understanding) — including handling inference latency, cost-per-call tradeoffs, and... 
    Full time
    Internship

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $186k - $282k

     ...products have outgrown the infrastructure patterns the rest of...  ...nothing like a CRUD service. They call foundation...  ...catches. Today, DevOps engineers carry this work...  ...including an agentic LLM thread runtime, a natural...  ...throughput, cross-region inference, quotas and throttles,... 

    FloQast

    San Jose, CA
    a month ago
  •  ...expertise in Google Cloud services, cloud-native...  ...Design Generative AI, LLM, RAG (Retrieval-Augmented...  ...teams. Mentor architects, engineers, and cloud...  ...Services Terraform / Infrastructure as Code Google AI & Machine...  ...modernization programs. AI/GenAI implementation for business... 

    Diamondpick

    Milpitas, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GenAI Engineer - LLM Infrastructure & Inference Services. Be the first to apply!