Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GenAI Engineer - LLM Infrastructure & Inference Services

Temporary

2T Consulting

We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage scalable enterprise GenAI platforms across GPU infrastructure and cloud environments.

Key Responsibilities

  • Deploy, host, and manage Large Language Models (LLMs) on GPU infrastructure for production environments.
  • Build scalable, high-performance inference services using vLLM, TensorRT-LLM, Triton Inference Server, and Ray Serve.
  • Optimize model serving for latency, throughput, GPU utilization, and cost efficiency.
  • Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes.
  • Implement RAG pipelines, vector databases, and agentic AI frameworks such as LangChain and LangGraph.
  • Manage GPU infrastructure, containerization, and cloud deployments across AWS, Azure, or GCP.
  • Establish MLOps/LLMOps practices including CI/CD, model deployment, monitoring, observability, and governance.
  • Perform performance tuning, benchmarking, capacity planning, and production support for enterprise GenAI platforms.
  • Collaborate with architects, data scientists, and product teams to deliver scalable, secure, and reliable AI solutions.

Core Technologies

  • LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve
  • AI/GenAI: RAG, LangChain, LangGraph, Vector Databases
  • Development: Python, FastAPI, Microservices
  • Infrastructure: Kubernetes, Docker, GPU Infrastructure
  • Cloud: AWS, Azure, GCP
  • MLOps/LLMOps: CI/CD, Monitoring, Observability, Model Deployment, Governance
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the GenAI Engineer - LLM Infrastructure & Inference Services in Santa Clara, CA vacancy
  •  ...THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications...  ...throughput, and cost efficiency for LLM and multimodal model serving in production...  ...agencies, or fee-based recruitment services. AMD and its subsidiaries are equal... 
    Suggested

    AMD

    San Jose, CA
    4 days ago
  • $197.3k - $225.1k

    Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable...  .... Our investments in technology infrastructure and world-class talent — along with...  ...have come to love the products and services we build. Team Description:... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    3 days ago
  • $224k - $356.5k

     ...now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is hiring software...  ...a fast-paced team building Generative AI inference platform to make design and deployment of...  ...inference engines (vLLM, SGLang, TRT-LLM) and expand these capabilities to support... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $177.1k - $387.5k

     ...platform that powers Zoom AI Services, enabling AI capabilities...  ...systems, cloud infrastructure, and AI platform engineering to build reliable, high-performance...  ...Translation (MT), LLM applications, AI Agents,...  ...including model serving, AI inference platforms, GPU/CPU resource... 
    Suggested
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    2 days ago
  • $224k - $356.5k

     ...advisor for Cloud Service Provider...  ...collaborating with engineering, product management...  ...systems, or AI/ML infrastructure — with direct work...  ...GCP), enterprise GenAI application development...  .../ML training and inference, and Amazon-...  ...including CUDA, TensorRT LLM, Dynamo, Triton... 
    Suggested
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184.7k - $324.8k

    Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa...  ...privacy-preserving cloud infrastructure for AI workloads. PCC represents...  ...Hands-on experience with LLM inference stacks....  ...operating high-throughput services at large distributed scale... 
    Worldwide
    Relocation

    Apple Inc.

    Santa Clara, CA
    1 day ago
  • $198k - $326k

     ...model training, feature engineering and serving with...  ...queries.Model Training Infrastructure: As an engineer on the...  ...complex models across LLM and Personalization models...  ..., enable GPU based inference for a large variety of...  ...locationBeing accompanied by a service dogHaving a sign... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    4 days ago
  •  ...week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role...  ...software experts to build out the deployment infrastructure, working closely with other software (...  ...serving frameworks (such as TensorRT-LLM, vLLM, SGLang, etc.)Experience with... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    20 hours ago
  • $313k - $353k

    Distinguished Software Engineer Distributed Systems | Platform Architecture | AI Infrastructure Overview Seeking a Distinguished Software...  ...develop distributed platform services supporting cloud, edge, and...  ...infrastructure, autonomous operations, LLM agents, LangGraph, AutoGen,... 

    Yoh

    Santa Clara, CA
    3 days ago
  •  ...industry-leading training and inference speeds; over 10 times faster...  ...hyperscale cloud inference services. This order of magnitude...  ...RoleWe're hiring a Principal Engineer for our Inference Cloud Platform...  ...with ML, Product and Infrastructure teams. ResponsibilitiesProblem... 

    Cerebras Systems

    Sunnyvale, CA
    20 hours ago
  • $165.2k - $223.6k

     ...are used by the world's largest cloud infrastructures?Would you enjoy broad yet equally deep...  ...globally?AWS Manufacturing Infrastructure Services continues to pioneer and our team is...  ...succeed in this role?Deeply technical engineers, who stay close to the customer as well... 
    Internship
    Work at office
    Local area
    Flexible hours
    Night shift
    Weekend work

    Amazon

    Cupertino, CA
    20 hours ago
  • $154k - $220k

     .... Staff Software Development Engineer-AI Security to join our team....  ...designing and implementing core infrastructure components and distributed...  ...for highly efficient I/O and service-oriented architectureBuild resilient...  ...the entire stack, including LLM models, by employing... 
    Full time
    Work at office
    Local area

    Zscaler

    San Jose, CA
    2 days ago
  •  ...industry-leading training and inference speeds; over 10 times faster...  ...hyperscale cloud inference services. This order of magnitude...  ...the RoleWe're hiring a Staff Engineer to own major areas of the architecture...  .... Partner with ML, Product, Infrastructure, and Platform teams to... 

    Cerebras Systems

    Sunnyvale, CA
    20 hours ago
  • $182k - $242k

     ...CoreWeave combines superior infrastructure performance with deep...  ...looking for a Senior Engineer to be a driving force...  ...training and inference workloads. You will own...  ...Where needed, build self-service views (Grafana, Looker...  ...Apache Spark, Trino, llm-d, vLLM, or PyTorch.... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    3 days ago
  •  ...We are hiring a Founding ML Infrastructure Engineer to own the end-to-end...  ...operating production-grade LLM systems . You will apply deep...  ...Responsibilities Own the end-to-end LLM inference stack , including: Model...  ...production-grade GPU services : Multi-model serving... 
    Work at office
    Visa sponsorship

    Realmlabs

    Sunnyvale, CA
    20 hours ago
  • $193.3k - $261.5k

     ...Senior Software Development Engineer to lead that charge, architecting...  ...expanding across Agentic AI, LLM-powered workflows,...  ..., security, and emerging AWS services, pioneering collaborative AI...  ...transformer architecture, training/inference lifecycles, and optimization... 
    Internship
    Local area
    Flexible hours

    AmazonWebServices

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...application is built. We are seeking a Senior Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward...  ...based on workload, hardware, model quality, and service-level objectives.Build and optimize performance-... 
    Full time

    Nvidia

    Santa Clara, CA
    15 hours ago
  • $152k - $241.5k

    The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks...  ...hybrid and multi-cloud infrastructure. We are building the next-...  ...GPU-based training and inference jobs. Our work gives NVIDIA...  ...platform teams, and partner engineering teams to understand requirements... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $109k - $160k

     ...Software Engineer - Data Infrastructure ServicesSunnyvale, CA / Bellevue, WA CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers...  ..., and scalability of our data platforms, and related services and participate in the teams on-call rotation.Establish guidelines... 
    Permanent employment
    Full time
    Casual work
    Work at office

    CoreWeave

    Sunnyvale, CA
    20 hours ago
  • $193.3k - $261.5k

     ...accelerate deep learning and GenAI workloads on Amazon’...  ...unparalleled ML inference and training...  ...software boundary, our engineers build systematic infrastructure, innovate new methods...  ...a wide variety of LLM model families, including...  ...customer service; and follow all federal... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $272k - $431.25k

     ...throughput, low-latency inference framework for serving...  ...of cutting-edge LLM workloads.We are seeking...  ...seeking a Principal Systems Engineer to define the vision...  ..., or ML systems infrastructure in C/C++ and Python, with...  ...delivering production services.Deep understanding of... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    20 hours ago
  • $151.8k - $332.2k

     ...expect We are looking for an AI Inference Engineer with a solid background in...  ...engineering teams, and infrastructure teams, to deliver high-impact...  ...state-of-the-art speech services for Zoom products. Devising...  ...speech recognition, speech-llm or AI model inference.Display... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    20 hours ago
  • $184k - $299k

     ...is changing how the largest service providers and telecom operators build and monetize AI infrastructure. AI is shifting from centralized...  ...training to distributed inference. Telcos and service providers...  ...ability to engage both engineering teams and C-suite executives... 
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    20 hours ago
  • $230k - $250k

     ...of Technical Staff in Sunnyvale, CA. This role involves designing resilient software features for cloud-based AI inference, leveraging AWS tools and services. Candidates should have a Master’s degree in Computer Science and experience with containerization tools like... 

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $165k - $242k

    A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this role, you'll lead design reviews, implement optimizations, and improve service reliability. The ideal candidate has extensive experience with distributed... 

    CoreWeave

    Sunnyvale, CA
    4 days ago
  •  ...that enhance system resiliency and high availability across distributed environments. The role includes developing scalable AI inference services and deploying cloud-based workflows. Ideal candidates have a master's degree and significant experience with deployment... 

    Cerebras Systems, Inc.

    Sunnyvale, CA
    3 days ago
  •  ...Inc. in Sunnyvale, CA, seeks a Senior Software Developer to shape GenAI products across the full development lifecycle, from debugging to...  ...GenAI initiatives such as FortiGPT, applying cutting-edge GenAI/LLM technologies to deliver innovative features for our next‑generation... 

    Fortinet, Inc.

    Sunnyvale, CA
    2 days ago
  • $157.3k - $212.8k

     ...the cloud for AI training and inference? Want to do industry leading...  ...— from foundational services such as Amazon’s Simple Storage...  ...software, hardware, and network engineers, supply chain specialists,...  ...large scale applications- AI infrastructure hardware development and debugging... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $193.93k - $352.29k

     ...is not fungible is the infrastructure that decides whether...  ...autonomously inside Nuro's own engineering organization, under...  ..., current taste in LLM research. You...  ...happens under the hood at inference. Attention and KV-cache...  ...cloud infrastructure, service design, storage, queuing... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago
  •  ...Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  ...distributed AI training and inference, raw GPU and CPU horsepower...  ....The Lambda Infrastructure Engineering organization forges the...  ...scalable and resilient storage services that power our AI and machine... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GenAI Engineer - LLM Infrastructure & Inference Services. Be the first to apply!