On-Prem LLM Inference Engineer: GPU & AI Infra
Compunnel, Inc.
Compunnel, Inc. in Charlotte, North Carolina is seeking an On-Premise LLM Inference & GPU Systems Engineer. This role involves building, optimizing, and supporting a large-scale enterprise Generative AI infrastructure utilizing NVIDIA H200 GPU clusters and OpenShift AI. Candidates should have significant experience in GPU runtime optimization, Kubernetes orchestration, and managing open-source LLMs. Key qualifications include 5+ years of experience in relevant roles and hands-on expertise with NVIDIA GPU environments. This position offers a contract opportunity with significant responsibilities across enterprise AI workloads. #J-18808-Ljbffr Compunnel, Inc.
- ...in Charlotte, NC is seeking an LLM Inference & GPU Systems Consultant to build and maintain on-prem LLM infrastructure on NVIDIA H200 clusters with an OpenShift AI deployment. The role focuses on... ...have 8+ years in LLM systems or AI infra, experience with OpenShift AI...Suggested3 days per week
- ....We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina... ...(US).Role Overview We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an...SuggestedWork at officeRemote workFlexible hours
- ...We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina... ...). Job Description We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an...Suggested
- ...026 Contract Active Job Description: Job Summary We are seeking an On-Premise LLM Inference & GPU Systems Engineer to build, optimize, and support a large-scale enterprise Generative AI infrastructure environment. This role is focused exclusively on Large Language Model...SuggestedContract work
- ...in leveraging advanced AI and analytics to create... ...., forecasting models, LLM-based solutions), while... ...patterns• Helm, Operators• GPU orchestration concepts... ...development, data engineering, and software engineering... ...Design and implement LLM inference serving stacks using: o...SuggestedFull timeTemporary workRelocation
- ...Lead Systems Operations Engineer within the Branch... ...managementExperience supporting AI/ML-enabled production... ...risk management for LLM- or agent-based workflows... ...g., model deployments, inference services, retrieval... ...implementations, including on prem, client server, on prem...Full timeWork experience placementWeekend workAfternoon shift3 days per week
- Synechron is seeking a Gen AI Tester to ensure the quality, reliability, security, and... ...enterprise Generative AI applications and LLM solutions. The role requires 6+ years in QA... ...governance, security testing, and collaboration across AI Engineers, Data #J-18808-Ljbffr Synechron
- Coinbase is seeking a Senior Software Engineer for the AI Platform to build and operate the LLM and agent infrastructure used across the company. You will own core platform systems including the LLM gateway, AI hub, and agent runtime, ensuring scalability, governance, and...
- Gina’s Tech Jobs - IT Recruiting Agency seeks an Artificial Intelligence (AI) Engineer for onsite work in Charlotte, NC. You will develop GPU-accelerated video inference pipelines and optimize YOLO-based models for real-time safety monitoring across large fleets and industrial...
- ...are currently seeking a AI Architect to join our... ...Agentic Stack & AI Platform Engineering:Spearhead the growth... ....Infrastructure, Inference & Edge Computing:Design... .../ML platforms.Optimize LLM inference, implementing... ...strategy, including rigorous GPU management, utilization...Work at officeLocal areaRemote workFlexible hours
$91.1k - $179.5k
...significant impact on our clients’ success. We are hiring an AI Engineer to build and operate the data, features, and GenAI... ...pipelines and services that support model training, real-time inference, and LLM applications using Claude-, GPT/Codex-, and Gemini-class models...Local area$134.5k - $265.1k
...6.Work you will do:As a Cyber Forward Deployed Engineer (FDE) Sr Consultant, you will drive delivery of cybersecurity and AI-enabled solutions at the intersection of client... ...hands-on experience building and deploying GenAI/LLM-powered solutions in client or production...Local areaVisa sponsorship- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure... ...that powers AI platforms, GPU-accelerated workloads, large-scale models... ...NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks...Full timeWork experience placementLive inWork at officeLocal area
- Position Summary Our Deloitte AI & Engineering team to transform technology platforms, drive innovation, and help make a significant... ...ventures, and fuel growth through innovation. Work you'll do As a Infra and DevOps Cloud Architect on the team, you will be...Work at officeLocal area
- ...enterprise, reviewing and guiding all AI/ML use cases. Push the... .... Lead prompt and context engineering strategies to maximize model... ...and AI/ML platforms. Optimize LLM inference, implementing advanced... ...strategy, including rigorous GPU management, utilization, and...Local area
- AI/ML Ops EngineerLocation: Remote / Hybrid (Client-Facing Consulting... ...seeking an experienced AI/ML Engineer to design, deploy, and operate... ...real-time APIs, batch inference pipelines, and feature stores.... ...Generative AI or Large Language Model (LLM) solutions in production...Full timeContract workLocal areaRemote workFlexible hours
- ...hands-on Financial Services Consultant with AI experience to design, build, and scale... ...leveraging AI techniques, with experience in LLM application development, agent... ...knowledge-driven applicationsImplement prompt engineering, evaluation frameworks, and guardrailsPerform...Full timeTemporary workWork experience placement
- ...engagement. A Cybersecurity Forward Deployed Engineer is a production engineer who works... ...their security and engineering teams—to make AI systems secure, governed, and resilient in... ...complex multi-stakeholder client environments—LLM systems, multi-agent pipelines, RAG...Full timeWork experience placementLive inWork at officeLocal area
- ...requirements About the role:As an AI Architect at TQL, you will... ...You will partner closely with Engineering, Product Management, Data and... ...toolsEstablish standards for building LLM applications, retrieval-... ..., embedding pipelines, inference services, feature stores and model...H1b
$122k - $240.5k
Position Summary Google AI Architect/AI and EngineeringJoin our AI & Engineering team in transforming technology platforms... ...fine-tune, evaluate, and govern LLM solutions with Gemini on Vertex... ...); implement deployment, inference optimization, and monitoring.Build...Local areaVisa sponsorshipFlexible hours- About this role:This Senior Software Engineer will help design, build, and evolve next-generation AI-powered operational platforms. In this role, you will drive the... ...Develop AI-powered workflows leveraging enterprise LLM platforms, MCP integrations, vector search, and...Full time
- ...We are seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS,... ...observability, infrastructure automation, and AI-assisted troubleshooting. The role will... ...containerized environments. Leverage AI/LLM tools, including Claude AI, for pipeline troubleshooting...
- ...technology and leadership in cloud, data and AI with unmatched industry experience,... ...industry knowledge and applied AI and data engineering. We help the world’s leading Resources and... ...and cloud platforms. You develop LLM-powered applications, APIs, and pipelines...Full timeWork experience placementLive inWork at officeLocal area
- ...Role : AI Automation Engineer Location : Charlotte NC (100% Onsite) Job Description This role requires deep expertise in scripting... ...multiple AI agents via API endpoints Familiarity with LLM-based services, AI orchestration patterns, or AI-assisted decision...
- Job title:AI Automation Engineer Location: Local to Charlotte,NC Only! Duration:6 Months Experience Required: 4-6 JD Must Have Technical/Functional... ...automation and AI agent workflows using Python, RPA, and LLM orchestration frameworks (LangChain / LangGraph), deployed...Local area
- AI Engineer GenAI / Agentic Systems Location: Charlotte, NC or Dallas, TX - Onsite, 5 days/week Duration: 6 18 Months Job Summary... ...AI solutions, with strong expertise in agentic AI, GraphRAG, LLM applications, and modern AI engineering practices. Key...Local area
$128k - $252.5k
...executives and data scientists to AI strategists, machine learning specialists, and data engineers. SFL Scientific, a Deloitte... ...sources using cloud computing or on-prem technologiesDesign and lead... ...Architect)2+ years of experience with GPU computing (CUDA, OpenCL) and HPC...Local areaVisa sponsorship- TechDigital Group is seeking a seasoned software/data engineer to build production‑grade AI agents, tools, prompts, evals, and integrations on an established reference architecture. You will implement multi‑agent systems, RAG pipelines, and scalable APIs using Python and...
$133.37k - $156.9k
...DescriptionU.S. Bank is seeking an AI Scientist to join the... ...combines strong scientific and engineering expertise with a practical mindset... ..., and governance of LLM-based systems.Assess emerging... ...systems, model evaluation, and inference optimization.Experience building...Full timeLocal area- ...-Track platform integrates GPS tracking, AI-powered video telematics, and real-time data... ...You’ll help build quality engineering as a first-class capability inside a high... ...development/testing tools (e.g., Cursor IDE, LLM-based test generation, intelligent automation...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to On-Prem LLM Inference Engineer: GPU & AI Infra. Be the first to apply!


