GenAI Engineer - LLM Infrastructure & Inference Services
2T Consulting
We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage scalable enterprise GenAI platforms across GPU infrastructure and cloud environments.
Key Responsibilities
- Deploy, host, and manage Large Language Models (LLMs) on GPU infrastructure for production environments.
- Build scalable, high-performance inference services using vLLM, TensorRT-LLM, Triton Inference Server, and Ray Serve.
- Optimize model serving for latency, throughput, GPU utilization, and cost efficiency.
- Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes.
- Implement RAG pipelines, vector databases, and agentic AI frameworks such as LangChain and LangGraph.
- Manage GPU infrastructure, containerization, and cloud deployments across AWS, Azure, or GCP.
- Establish MLOps/LLMOps practices including CI/CD, model deployment, monitoring, observability, and governance.
- Perform performance tuning, benchmarking, capacity planning, and production support for enterprise GenAI platforms.
- Collaborate with architects, data scientists, and product teams to deliver scalable, secure, and reliable AI solutions.
Core Technologies
- LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve
- AI/GenAI: RAG, LangChain, LangGraph, Vector Databases
- Development: Python, FastAPI, Microservices
- Infrastructure: Kubernetes, Docker, GPU Infrastructure
- Cloud: AWS, Azure, GCP
- MLOps/LLMOps: CI/CD, Monitoring, Observability, Model Deployment, Governance
$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible... ...experiences. Our investments in technology infrastructure and world-class talent — along with... ...have come to love the products and services we build. Team Description:...SuggestedFull timePart timeLocal area$198k - $326k
...model training, feature engineering and serving with... ...queries.Model Training Infrastructure: As an engineer on the... ...complex models across LLM and Personalization models... ..., enable GPU based inference for a large variety of... ...locationBeing accompanied by a service dogHaving a sign...SuggestedFor contractorsWork at officeFlexible hours$229.9k - $262.4k
...Overview AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and... ...experiences. Our investments in technology infrastructure and world-class talent — along... ...come to love the products and services we build. Team Description:...SuggestedFull timePart timeLocal area$184k - $287.5k
...application is built. We are seeking a Senior Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward... ...based on workload, hardware, model quality, and service-level objectives.Build and optimize performance-...SuggestedFull time$159.2k - $301.6k
...like a startup founding engineer while having the... ...Engineer to build Creator Services from the ground up. Creator... ...cost-efficient inference, caching, and graceful... ...experience building with LLM APIs and agent... ...Colligo or similar ML infrastructure.About AdobeAdobe empowers...SuggestedFull timeTemporary workLocal areaWorldwide$137.1k - $299.3k
...Team DoorDash’s GenAI Platform team sits... ...builds the shared infrastructure that helps... ...trust the quality of LLM and agent systems... ...serving and batch inference, guardrails, and cost... ...pipelines, backend services, and observability... ...role is ideal for an engineer who enjoys...Hourly payFull timeWork at officeLocal areaRemote workFlexible hours$307k - $427k
...Principal M-LLM Post-Training and Execution Software Engineer, XR Share Principal M-LLM Post... ...with high-performance inference engines and custom AI accelerators... ...party devices and services that combine the best... ...the post-training infrastructure for Multimodal LLMs, including...$92k - $135k
...CoreWeave combines superior infrastructure performance with deep... ...You'll Do: Join the Inference team to ship production... ...mentorship from experienced engineers. About the role:... ...Go/C++ for model-serving services (e.g., Triton, vLLM, TensorRT-LLM, Ray Serve). Write tests...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeFlexible hours$165.2k - $223.6k
...team at Amazon Web Services (AWS) builds AWS Neuron... ...deep learning and GenAI workloads on Amazon... ...unparalleled ML inference and training performance... ...boundary, our engineers build systematic infrastructure, innovate new methods... ...a wide variety of LLM model families,...Work experience placementInternshipLocal areaFlexible hours$274k - $300k
...please visit AI Platform Engineer – Training & Inference Saviynt's AI-powered... ...+ H100s, the multi-engine LLM inference mesh (vLLM,... ...LLMs • Build RL training infrastructure: define Flyte workflows for... ...high-growth, Platform as a Service company focused on...$109k - $160k
Software Engineer - Data Infrastructure Services Sunnyvale, CA / Bellevue, WA CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...Summary: Hands-on full-stack software engineer supporting NVIDIA Test Infrastructure Development. The contractor will... ...senior engineers, including backend services, REST APIs, database integrations,... ...AI coding tools or experience with LLM APIs, prompt engineering, RAG/vector...For contractors
- ...a leader in AI cloud infrastructure serving tens of thousands... ...and enterprise engineering teams to design, scale... ...evaluations across training and inference workloads to show... ...optimization (vLLM, TensorRT-LLM) and observability... ...APIs, gRPC, and service-oriented cloud architectures...Work at officeLocal areaWork from homeFlexible hours
$272k - $425.5k
Principal Software Engineer - Large-Scale LLM Memory and Storage Systems page is... ...-throughput, low-latency inference framework for serving generative... ...storage, or ML systems infrastructure in C/C++ and Python, with... ...of delivering production services.* Deep understanding of...Local areaRemote work$175.6k - $263.4k
...Computer Science, Computer Engineering, or a closely related discipline... ...Experience developing AI/LLM-powered applications and integrations... ...We develop scalable backend services and APIs using Python, Node.... ...engineering productivity, infrastructure management, analytics, and...Full timeWork from home$229.9k - $262.4k
...Sr. Lead AI Engineer (Inference Optimization, FM Hosting, AI Platform)At... ...investments in technology infrastructure and world-class talent — along... ...come to love the products and services we build.The Intelligent... ...introduce state-of-the-art LLM optimization techniques to...Full timePart time$116k - $189.75k
...We are looking for a Software Engineer focused on bring-up, triage,... ...of distributed training and inference workloads across NVIDIA GPU platforms... ..., and debug distributed LLM workloads on multi-GPU and... ...debug large-scale AI clusters, infrastructure, and end-to-end workloads....Full timeRemote work$184k - $287.5k
...looking for a Senior Software Engineer to lead the bring-up, triage,... ...of distributed training and inference workloads across NVIDIA GPU platforms... ...to ensure state-of-the-art LLM workloads run efficiently and... ...of large-scale AI clusters, infrastructure, and end-to-end workloads,...Full timeRemote work- ...small, highly motivated, and focused on engineering excellence. This organization is for... ...: As part of the Network Software and Services for AI (nssAI) team at SpaceXAI, you'll... ...used for AI training and serving customer inference queries. Implement IaC best...Temporary work
$175.6k - $263.4k
...Technologies, Inc. Job Area: Engineering Group, Engineering Group... ...Summary: Qualcomm's CAD Infrastructure team develops and operates... ...automation platforms, data services, and AI-enabled solutions... ...tools. Integrate AI/ML and LLM technologies into internal engineering...Work experience placementWork from home$2,500 per month
...are heavily focused on inference . Backed by hundreds of... ...and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing... ...stack with LLM integration and a strong... ...simulation clusters, and service endpoints. This role...Work at officeRelocation package$182k - $242k
...CoreWeave combines superior infrastructure performance with deep... ...looking for a Senior Engineer to be a driving force... ...training and inference workloads. You will own... ...Where needed, build self-service views (Grafana, Looker... ...Apache Spark, Trino, llm-d, vLLM, or PyTorch....Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$262k - $364k
...Influence and coach a distributed team of engineers.Facilitate alignment and clarity... ...Experience integrating generative AI tools or LLM interfaces into workflows.Preferred... ...traffic engineering is a critical infrastructure service for Google. In this role, you will be...Worldwide$182k - $242k
Senior Software Engineer, Network Services Livingston, NJ / New York, NY / Sunnyvale, CA CoreWeave is The Essential Cloud for AI™. Built for... ...startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$119.25k - $150.85k
Description About the Team The HIL Platform and Services team in GM’s Autonomous Vehicle (AV) Organization is responsible for developing... ...they reach the road. About the Role As a Software Engineer on the HIL Platform and Services team, you will be at the...Full timeInternshipLocal areaWork from homeRelocation package$195k - $285k
...Matrix designs and manufactures purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable... ...AWS, Azure, GCP), and customer-facing platform services. You will hire and grow the team, set technical...Remote work$184k - $287.5k
...projects to critically important infrastructure. Be part of a dynamic team... ...full lifecycle of agentic services: from prototype through evaluation... ...improvementBuild robust LLM-powered pipelines for code generation... ...closely with GPU SW kernel engineers to pinpoint high-impact...Full time$152k - $241.5k
...best talent. We’re hiring a Deep Learning Engineer with strong experience in generative AI,... ...pipelines fast and iterate fasterLeverage LLM/VLM and agents in the data generation... ...document understanding) — including handling inference latency, cost-per-call tradeoffs, and...Full timeInternship$186k - $282k
...products have outgrown the infrastructure patterns the rest of... ...nothing like a CRUD service. They call foundation... ...catches. Today, DevOps engineers carry this work... ...including an agentic LLM thread runtime, a natural... ...throughput, cross-region inference, quotas and throttles,...- ...expertise in Google Cloud services, cloud-native... ...Design Generative AI, LLM, RAG (Retrieval-Augmented... ...teams. Mentor architects, engineers, and cloud... ...Services Terraform / Infrastructure as Code Google AI & Machine... ...modernization programs. AI/GenAI implementation for business...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GenAI Engineer - LLM Infrastructure & Inference Services. Be the first to apply!
- data infrastructure engineer Santa Clara, CA
- infrastructure engineer Santa Clara, CA
- remote infrastructure engineer Santa Clara, CA
- senior infrastructure engineer Santa Clara, CA
- infrastructure developer Santa Clara, CA
- associate infrastructure engineer
- senior IT infrastructure engineer
- junior infrastructure engineer
- lead infrastructure engineer
- security infrastructure engineer





