LLM Engineer
Talent Software Services
Job Details
- Job Title: LLM Engineer
- Location: Cincinnati, OH
- Work Location: Remote - USA
- Duration: 1 year
- Experience Required: 8+ years
- Role Category: AI and Automation
- Education: Any degree
- Start Date: 15-September-2026
- Design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities.
- Build secure, reliable, reusable, and enterprise-ready AI capabilities.
- Support:
- Agentic AI workflows
- AI for SDLC
- Knowledge retrieval
- Model evaluation
- Private AI hosting
- AgentOps
- Work closely with:
- Principal AI Architect
- AI Engineering Lead
- Platform Engineers
- Security teams
- Enterprise Architecture
- Product Owners
- Domain teams
- Large Language Models (LLMs)
- Small Language Models (SLMs)
- Prompt Engineering
- Context Engineering
- Retrieval-Augmented Generation (RAG)
- Embeddings
- Semantic Search
- Agentic AI Patterns
- Multi-Agent Workflows
- Tool Calling
- Function Calling
- Model Evaluation
- LLM Observability
- Fine-Tuning
- Supervised Fine-Tuning
- LoRA
- QLoRA
- Quantization
- Distillation
- Model Compression
- Synthetic Data Generation
- Model Benchmarking
- Model Selection
- Model Routing
- Private LLM Hosting
- On-Prem Model Deployment
- GPU-Based Inference
- Model Serving APIs
- High-Availability Inference
- Autoscaling
- Load Balancing
- Caching
- Batch and Real-Time Inference
- Kubernetes
- Docker
- Kubeflow
- KServe
- Ray Serve
- MLflow
- Hugging Face
- Transformers
- PyTorch
- PEFT
- DeepSpeed
- NVIDIA NIM
- Triton Inference Server
- TensorRT-LLM
- vLLM
- TGI
- SGLang
- Python
- TypeScript or JavaScript
- REST APIs
- Microservices
- CI/CD
- GitHub or Azure DevOps
- API Design
- Distributed Systems
- Cloud-Native Engineering
- Test Automation
- Vector Databases
- Knowledge Graphs
- Document Processing
- Metadata Management
- Data Pipelines
- Object Storage
- Enterprise Search
- Structured and Unstructured Data Integration
- Build enterprise-grade LLM-powered applications and intelligent agent capabilities.
- Design reusable LLM patterns, services, APIs, and accelerators.
- Develop model interaction patterns for:
- Reasoning
- Summarization
- Classification
- Extraction
- Planning
- Decision support
- Build reusable prompt, context, retrieval, memory, and evaluation components.
- Support AI-for-SDLC agents across:
- Requirements
- Design
- Coding
- Testing
- Security Review
- Deployment
- Operations
- Convert AI use cases into scalable production solutions.
- Build core intelligence services for enterprise agents.
- Develop reusable capabilities for:
- Planning
- Task decomposition
- Reasoning
- Tool usage
- Agent collaboration
- Enable agent-to-agent interaction and multi-agent orchestration.
- Integrate LLMs with:
- Agent runtimes
- Tool registries
- Workflow engines
- MCP-based gateways
- Support human-in-the-loop, approval, escalation, and feedback workflows.
- Improve agent quality, accuracy, safety, and task completion.
- Design reusable prompt engineering standards, templates, and libraries.
- Create:
- System prompts
- Task prompts
- Role prompts
- Guardrail prompts
- Evaluation prompts
- Develop context engineering strategies for better grounding, relevance, and personalization.
- Optimize:
- Token usage
- Context windows
- Memory injection
- Retrieval inputs
- Establish prompt versioning, testing, and governance practices.
- Design and implement enterprise RAG architectures.
- Build retrieval pipelines using:
- Enterprise documents
- Knowledge repositories
- Structured data
- Metadata
- Optimize:
- Chunking
- Embeddings
- Indexing
- Ranking
- Reranking
- Retrieval strategies
- Improve grounding, citation quality, precision, recall, and factual accuracy.
- Build reusable retrieval services for agents and business domains.
- Partner with data and knowledge management teams to onboard trusted data sources.
- Evaluate, build, fine-tune, deploy, and optimize LLMs and SLMs.
- Support domain-specific model development using approved datasets.
- Build supervised fine-tuning and model adaptation pipelines.
- Apply:
- LoRA
- QLoRA
- Distillation
- Quantization
- Model compression
- Evaluate commercial, open-source, and internally hosted models.
- Select models based on:
- Accuracy
- Latency
- Cost
- Data residency
- Security
- Operational requirements
- Build and support private AI capabilities for LLM/SLM hosting.
- Deploy models across:
- On-premises
- Hybrid
- Private cloud environments
- Support GPU-enabled model hosting.
- Optimize latency, throughput, concurrency, resiliency, and GPU utilization.
- Build secure inference endpoints for internal applications and agents.
- Support air-gapped and restricted AI environments.
- Partner with infrastructure and platform teams on private AI hosting.
- Implement scalable model serving using modern inference frameworks.
- Build high-availability inference architectures.
- Optimize:
- Token throughput
- Response latency
- Cost efficiency
- Inference performance
- Implement:
- Model routing
- Load balancing
- Caching
- Fallback strategies
- Support batch and real-time inference.
- Develop reusable deployment templates for different model families.
- Build operational practices for managing models and agents throughout their lifecycle.
- Implement observability for:
- Prompts
- Retrieval
- Model responses
- Latency
- Cost
- Failures
- Develop evaluation pipelines for regression testing and continuous quality improvement.
- Monitor:
- Model drift
- Response quality
- Hallucination indicators
- Safety risks
- Support CI/CD and release management for:
- Prompts
- Models
- Agents
- Retrieval pipelines
- Build dashboards and metrics for AI quality, reliability, adoption, and operational readiness.
- Define and implement LLM evaluation frameworks.
- Measure:
- Accuracy
- Groundedness
- Relevance
- Hallucination rate
- Toxicity risk
- Safety compliance
- Task completion
- User satisfaction
- Build automated test suites for prompts, agents, tools, and RAG pipelines.
- Benchmark models across enterprise use cases.
- Compare cloud, open-source, and on-prem models based on performance, cost, quality, and risk.
- Establish quality gates for production AI releases.
- Implement Responsible AI controls in LLM applications and agent workflows.
- Develop guardrails for:
- Safe output
- Tool usage
- Data access
- Enterprise policy compliance
- Support:
- Model risk management
- Auditability
- Transparency
- Traceability
- Ensure sensitive data is handled according to security and privacy requirements.
- Collaborate with Security, Enterprise Architecture, Risk, and Compliance teams.
- Support model and agent approval and production-readiness governance.
- Experience deploying open-source models such as:
- Llama
- Mistral
- Mixtral
- Phi
- Gemma
- Qwen
- DeepSeek
- Granite
- Falcon
- Domain-specific models
- Experience with GPU infrastructure such as:
- NVIDIA H100
- H200
- B200
- B300
- A100
- L40S
- GH200
- AMD MI300X
- Experience with:
- Private AI
- Hybrid AI
- Air-gapped AI environments
- Experience in regulated industries such as:
- Healthcare
- Financial Services
- Insurance
- Experience building:
- Enterprise copilots
- AI assistants
- Agent platforms
- Experience with:
- MCP
- Tool registries
- Agent runtimes
- Enterprise integration patterns
- Experience with Responsible AI, model governance, model risk management, and AI compliance.
- Experience optimizing AI workloads for:
- Cost
- Performance
- Latency
- Security
- Role: LLM Engineer
- Essential Skill: LLM
- Primary Skill: AI and Automation
- Experience: 8-10+ years
- Work Location: Remote USA
- Duration: 1 year
- Core Technologies: LLM, SLM, RAG, Generative AI, Agentic AI, Python, Kubernetes, MLflow, Hugging Face, PyTorch, vLLM, KServe, Docker, Vector Databases
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the LLM Engineer in United States vacancy
- ...Ventures. About The Role In this role, you'll own the core LLM infrastructure powering two products redefining B2B research:... ...translate user needs into technical solutions while maintaining engineering best practices Nice To Have RAG systems, embeddings,...SuggestedFull timeImmediate startFlexible hours
- ...part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC), United States (US).Role Overview We are seeking an AI Infrastructure...SuggestedWork at officeRemote workFlexible hours
- General Information Job Title ML Staff Engineer - LLM & Production Systems Job ID 107242 Work Areas Technology & Engineering Employment Type Permanent Full-Time Location(s) Boston, Chicago, Dallas, New York Description & Requirements...SuggestedPermanent employmentFull timeWork at officeLocal area1 day per week
- ...Texas Sports Academy is on the lookout for a Senior AI Engineer specializing in LLM (Large Language Model) Systems and RAG (Retrieval-Augmented Generation) Optimization. As we continue to push the boundaries of sports technology, your role will be pivotal in developing...SuggestedRemote jobFull time
$85k - $115k
Vein Clinics of America, Inc. is seeking an experienced LLM Engineer to serve as the AI technical lead. In this full-time, on-site role in Northbrook, IL, you will design, implement, and optimize LLM solutions that improve business processes. Responsibilities include developing...SuggestedFull time- B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a...
- ID.me is hiring a Staff Software Development Engineer to lead the Tools Team in McLean, VA. You will champion AI-native engineering, architect... ...has 10+ years in full-stack engineering, including hands-on LLM integration experience. This on-site role fosters collaboration...
- ...latency, and robustness. Build shared APIs and platform components used broadly across engineering teams. Key Responsibilities Design and implement orchestration patterns for LLM-powered agents. Evaluate and select models, tools, and providers based on...Full time
- Siamo alla ricerca di un/una AI Engineer da inserire nel nostro team dedicato alla trasformazione digitale dei processi del settore bancario... ...di automazione basati su Node.js Integrazione di modelli AI (LLM) all’interno dei processi aziendali Progettazione e implementazione...
- ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-performance KV cache layer for LLM inference across GPUs and nodes. This role focuses on reducing latency, increasing throughput, and lowering cost for...
$191k - $253k
...that platform so the documentation pipeline handles more work at lower cost and lower latency.Role SpecificsAs the Applied LLM Systems Engineer, your mission is to design, build, and operate production AI systems that improve how technical documentation is created, transformed...Full timeWork experience placementImmediate start- ...AI Software Engineer LLM Evaluation & Automation (Remote) We are looking for an AI Software Engineer for a B2B high-tech company. In this role, you will build and integrate evaluation harnesses and automation for software development use cases, including turning...Remote work
- ...Hiring: AI / LLM Developer (Florida – In-Person Collaboration)I'm seeking an experienced AI / LLM developer based in Florida for... ...extensibility and future customization in mindThis is a hands-on engineering role, not prompt engineering or theoretical research.Required ExperienceDemonstrated...Contract workLocal area
- ...We are looking for an LLM Ops Engineer with deep Databricks experience to build, automate, and scale our machine learning delivery pipelines on the Lakehouse. You’ll own the model lifecycle end‑to‑end—from data ingestion and feature engineering to CI/CD, deployment, monitoring...
- ...looking for someone who "has experience with AI" we're looking for someone who builds with it every day. An engineer whose default instinct is to reach for an LLM, an agent framework, or a durable workflow instead of writing another for-loop. This is a high-ownership...Full timeLive inImmediate startWeekend work1 day per week
- Anduril Industries seeks an Applied LLM Systems Engineer to design, build, and operate production AI systems that improve technical documentation workflows. You will develop AI-assisted authoring, automated publishing, and multi-agent tooling for the Technical Publications...
$87.95k - $203.95k
...want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a API LLM Integration / ReactJS Engineer - Hybrid to join our team in Santa Clara, California (US-CA), United States (US).This API LLM Integration Engineer...Temporary workWork at officeRemote workFlexible hours$184k - $287.5k
...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing...Full timeWork experience placement$272k - $431.25k
...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full... ...engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across...Full time$170k - $245k
...Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an...Work at office- ...Senior LLM Engineer Location(s) Austin, Texas | Fremont, California | Lisle, Illinois Eligible for remote Yes Company Molex Career Field Data & Analytics Job Number 192106 As a Koch company, Molex is a leading supplier of connectors and interconnect components,...Remote work
- ...As an LLM Systems / AI Agent Engineer , you will focus on building and evolving production AI agents on foundation models, covering orchestration, context engineering, evaluation pipelines, and production observability and monitoring. This is an LLM systems and AI agent...Full timeSummer workImmediate startShift work
$81.34k - $141.5k
LLM Engineer As LLM Engineer, you will make an impact by designing, building, optimizing, and deploying Large Language Model (LLM) and Small Language Model (SLM) solutions that power the enterprise Agent Factory. You will be a valued member of the AI Engineering team and...Remote jobTemporary work- ...LLM Inference Engineer San Francisco or Remote Locations: San Francisco or Remote About The Role The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient...Remote work
- ...To support the Army VANTAGE platform, the full-time LLM Engineer will focus on LLM integration and fine-tuning, prompt engineering, and deploying AI/ML solutions while working remotely. Key responsibilities Collaborate with the OptiSource development team to implement...Full timeRemote work
- ...LLM Engineer Remote- Local to Chicago, IL preferred 12 months Video VISA RESTRICTIONS Required Skills: +3 Years Python, Client, AWS Cloud Computing Experience Familiarity with agent-based LLM frameworks (e.g.,LangChain) TensorFlow, PyTorch, and Hugging Face Transformers...Local areaRemote work
- ...Role Mission As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses...Work at office
$85k - $115k
...LLM Engineer The LLM Engineer serves as the organization's AI technical lead responsible for designing, implementing, and optimizing Large Language Model (LLM) solutions that automate business processes, improve operational efficiency, and support strategic initiatives...Full timeMonday to Friday- ...Job Title: LLM Engineer (Large Language Model Engineer) Job Summary: We are seeking a highly skilled LLM Engineer to design, develop, fine-tune, and deploy Large Language Model (LLM) based applications and AI-powered solutions. The ideal candidate...Full timeRemote work
$184k - $287.5k
We are now looking for a Senior High-Performance LLM Training Engineer! NVIDIA is seeking experienced engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing...Work experience placement
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Engineer. Be the first to apply!



