Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Engineer

Talent Software Services

Job Details
  • Job Title: LLM Engineer
  • Location: Cincinnati, OH
  • Work Location: Remote - USA
  • Duration: 1 year
  • Experience Required: 8+ years
  • Role Category: AI and Automation
  • Education: Any degree
  • Start Date: 15-September-2026
Role Overview
  • Design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities.
  • Build secure, reliable, reusable, and enterprise-ready AI capabilities.
  • Support:
    • Agentic AI workflows
    • AI for SDLC
    • Knowledge retrieval
    • Model evaluation
    • Private AI hosting
    • AgentOps
  • Work closely with:
    • Principal AI Architect
    • AI Engineering Lead
    • Platform Engineers
    • Security teams
    • Enterprise Architecture
    • Product Owners
    • Domain teams
Must-Have Technical Skills

LLM & Generative AI
  • Large Language Models (LLMs)
  • Small Language Models (SLMs)
  • Prompt Engineering
  • Context Engineering
  • Retrieval-Augmented Generation (RAG)
  • Embeddings
  • Semantic Search
  • Agentic AI Patterns
  • Multi-Agent Workflows
  • Tool Calling
  • Function Calling
  • Model Evaluation
  • LLM Observability
Model Engineering
  • Fine-Tuning
  • Supervised Fine-Tuning
  • LoRA
  • QLoRA
  • Quantization
  • Distillation
  • Model Compression
  • Synthetic Data Generation
  • Model Benchmarking
  • Model Selection
  • Model Routing
Model Hosting & Serving
  • Private LLM Hosting
  • On-Prem Model Deployment
  • GPU-Based Inference
  • Model Serving APIs
  • High-Availability Inference
  • Autoscaling
  • Load Balancing
  • Caching
  • Batch and Real-Time Inference
AI Infrastructure & Frameworks
  • Kubernetes
  • Docker
  • Kubeflow
  • KServe
  • Ray Serve
  • MLflow
  • Hugging Face
  • Transformers
  • PyTorch
  • PEFT
  • DeepSpeed
  • NVIDIA NIM
  • Triton Inference Server
  • TensorRT-LLM
  • vLLM
  • TGI
  • SGLang
Programming & Engineering
  • Python
  • TypeScript or JavaScript
  • REST APIs
  • Microservices
  • CI/CD
  • GitHub or Azure DevOps
  • API Design
  • Distributed Systems
  • Cloud-Native Engineering
  • Test Automation
Data & Knowledge Systems
  • Vector Databases
  • Knowledge Graphs
  • Document Processing
  • Metadata Management
  • Data Pipelines
  • Object Storage
  • Enterprise Search
  • Structured and Unstructured Data Integration
Roles & Responsibilities

LLM Application Engineering
  • Build enterprise-grade LLM-powered applications and intelligent agent capabilities.
  • Design reusable LLM patterns, services, APIs, and accelerators.
  • Develop model interaction patterns for:
    • Reasoning
    • Summarization
    • Classification
    • Extraction
    • Planning
    • Decision support
  • Build reusable prompt, context, retrieval, memory, and evaluation components.
  • Support AI-for-SDLC agents across:
    • Requirements
    • Design
    • Coding
    • Testing
    • Security Review
    • Deployment
    • Operations
  • Convert AI use cases into scalable production solutions.
Agent Factory Intelligence Layer
  • Build core intelligence services for enterprise agents.
  • Develop reusable capabilities for:
    • Planning
    • Task decomposition
    • Reasoning
    • Tool usage
    • Agent collaboration
  • Enable agent-to-agent interaction and multi-agent orchestration.
  • Integrate LLMs with:
    • Agent runtimes
    • Tool registries
    • Workflow engines
    • MCP-based gateways
  • Support human-in-the-loop, approval, escalation, and feedback workflows.
  • Improve agent quality, accuracy, safety, and task completion.
Prompt & Context Engineering
  • Design reusable prompt engineering standards, templates, and libraries.
  • Create:
    • System prompts
    • Task prompts
    • Role prompts
    • Guardrail prompts
    • Evaluation prompts
  • Develop context engineering strategies for better grounding, relevance, and personalization.
  • Optimize:
    • Token usage
    • Context windows
    • Memory injection
    • Retrieval inputs
  • Establish prompt versioning, testing, and governance practices.
Retrieval-Augmented Generation (RAG)
  • Design and implement enterprise RAG architectures.
  • Build retrieval pipelines using:
    • Enterprise documents
    • Knowledge repositories
    • Structured data
    • Metadata
  • Optimize:
    • Chunking
    • Embeddings
    • Indexing
    • Ranking
    • Reranking
    • Retrieval strategies
  • Improve grounding, citation quality, precision, recall, and factual accuracy.
  • Build reusable retrieval services for agents and business domains.
  • Partner with data and knowledge management teams to onboard trusted data sources.
LLM / SLM Model Engineering
  • Evaluate, build, fine-tune, deploy, and optimize LLMs and SLMs.
  • Support domain-specific model development using approved datasets.
  • Build supervised fine-tuning and model adaptation pipelines.
  • Apply:
    • LoRA
    • QLoRA
    • Distillation
    • Quantization
    • Model compression
  • Evaluate commercial, open-source, and internally hosted models.
  • Select models based on:
    • Accuracy
    • Latency
    • Cost
    • Data residency
    • Security
    • Operational requirements
Private AI & On-Prem Hosting
  • Build and support private AI capabilities for LLM/SLM hosting.
  • Deploy models across:
    • On-premises
    • Hybrid
    • Private cloud environments
  • Support GPU-enabled model hosting.
  • Optimize latency, throughput, concurrency, resiliency, and GPU utilization.
  • Build secure inference endpoints for internal applications and agents.
  • Support air-gapped and restricted AI environments.
  • Partner with infrastructure and platform teams on private AI hosting.
Model Serving & Inference Optimization
  • Implement scalable model serving using modern inference frameworks.
  • Build high-availability inference architectures.
  • Optimize:
    • Token throughput
    • Response latency
    • Cost efficiency
    • Inference performance
  • Implement:
    • Model routing
    • Load balancing
    • Caching
    • Fallback strategies
  • Support batch and real-time inference.
  • Develop reusable deployment templates for different model families.
LLMOps / ModelOps / AgentOps
  • Build operational practices for managing models and agents throughout their lifecycle.
  • Implement observability for:
    • Prompts
    • Retrieval
    • Model responses
    • Latency
    • Cost
    • Failures
  • Develop evaluation pipelines for regression testing and continuous quality improvement.
  • Monitor:
    • Model drift
    • Response quality
    • Hallucination indicators
    • Safety risks
  • Support CI/CD and release management for:
    • Prompts
    • Models
    • Agents
    • Retrieval pipelines
  • Build dashboards and metrics for AI quality, reliability, adoption, and operational readiness.
AI Evaluation & Benchmarking
  • Define and implement LLM evaluation frameworks.
  • Measure:
    • Accuracy
    • Groundedness
    • Relevance
    • Hallucination rate
    • Toxicity risk
    • Safety compliance
    • Task completion
    • User satisfaction
  • Build automated test suites for prompts, agents, tools, and RAG pipelines.
  • Benchmark models across enterprise use cases.
  • Compare cloud, open-source, and on-prem models based on performance, cost, quality, and risk.
  • Establish quality gates for production AI releases.
Responsible AI, Security & Governance
  • Implement Responsible AI controls in LLM applications and agent workflows.
  • Develop guardrails for:
    • Safe output
    • Tool usage
    • Data access
    • Enterprise policy compliance
  • Support:
    • Model risk management
    • Auditability
    • Transparency
    • Traceability
  • Ensure sensitive data is handled according to security and privacy requirements.
  • Collaborate with Security, Enterprise Architecture, Risk, and Compliance teams.
  • Support model and agent approval and production-readiness governance.
Preferred / Additional Experience
  • Experience deploying open-source models such as:
    • Llama
    • Mistral
    • Mixtral
    • Phi
    • Gemma
    • Qwen
    • DeepSeek
    • Granite
    • Falcon
    • Domain-specific models
  • Experience with GPU infrastructure such as:
    • NVIDIA H100
    • H200
    • B200
    • B300
    • A100
    • L40S
    • GH200
    • AMD MI300X
  • Experience with:
    • Private AI
    • Hybrid AI
    • Air-gapped AI environments
  • Experience in regulated industries such as:
    • Healthcare
    • Financial Services
    • Insurance
  • Experience building:
    • Enterprise copilots
    • AI assistants
    • Agent platforms
  • Experience with:
    • MCP
    • Tool registries
    • Agent runtimes
    • Enterprise integration patterns
  • Experience with Responsible AI, model governance, model risk management, and AI compliance.
  • Experience optimizing AI workloads for:
    • Cost
    • Performance
    • Latency
    • Security
Role / Skills Summary
  • Role: LLM Engineer
  • Essential Skill: LLM
  • Primary Skill: AI and Automation
  • Experience: 8-10+ years
  • Work Location: Remote USA
  • Duration: 1 year
  • Core Technologies: LLM, SLM, RAG, Generative AI, Agentic AI, Python, Kubernetes, MLflow, Hugging Face, PyTorch, vLLM, KServe, Docker, Vector Databases
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the LLM Engineer in United States vacancy
  •  ...Ventures. About The Role In this role, you'll own the core LLM infrastructure powering two products redefining B2B research:...  ...translate user needs into technical solutions while maintaining engineering best practices Nice To Have RAG systems, embeddings,... 
    Suggested
    Full time
    Immediate start
    Flexible hours

    Newtonx

    United States
    1 day ago
  •  ...part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC), United States (US).Role Overview We are seeking an AI Infrastructure... 
    Suggested
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Charlotte, NC
    4 days ago
  • General Information Job Title ML Staff Engineer - LLM & Production Systems Job ID 107242 Work Areas Technology & Engineering Employment Type Permanent Full-Time Location(s) Boston, Chicago, Dallas, New York Description & Requirements... 
    Suggested
    Permanent employment
    Full time
    Work at office
    Local area
    1 day per week

    Bain & Company

    Boston, MA
    4 days ago
  •  ...Texas Sports Academy is on the lookout for a Senior AI Engineer specializing in LLM (Large Language Model) Systems and RAG (Retrieval-Augmented Generation) Optimization. As we continue to push the boundaries of sports technology, your role will be pivotal in developing... 
    Suggested
    Remote job
    Full time

    Texas Sports Academy

    United States
    1 day ago
  • $85k - $115k

    Vein Clinics of America, Inc. is seeking an experienced LLM Engineer to serve as the AI technical lead. In this full-time, on-site role in Northbrook, IL, you will design, implement, and optimize LLM solutions that improve business processes. Responsibilities include developing... 
    Suggested
    Full time

    Vein Clinics of America, Inc.

    Northbrook, IL
    2 days ago
  • B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a... 

    B Capital

    San Francisco, CA
    3 days ago
  • ID.me is hiring a Staff Software Development Engineer to lead the Tools Team in McLean, VA. You will champion AI-native engineering, architect...  ...has 10+ years in full-stack engineering, including hands-on LLM integration experience. This on-site role fosters collaboration... 

    ID.me

    Mc Lean, VA
    3 days ago
  •  ...latency, and robustness. Build shared APIs and platform components used broadly across engineering teams. Key Responsibilities Design and implement orchestration patterns for LLM-powered agents. Evaluate and select models, tools, and providers based on... 
    Full time

    Calliere

    Remote
    1 day ago
  • Siamo alla ricerca di un/una AI Engineer da inserire nel nostro team dedicato alla trasformazione digitale dei processi del settore bancario...  ...di automazione basati su Node.js Integrazione di modelli AI (LLM) all’interno dei processi aziendali Progettazione e implementazione... 

    Gruppo MOL

    New York State
    6 days ago
  • ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-performance KV cache layer for LLM inference across GPUs and nodes. This role focuses on reducing latency, increasing throughput, and lowering cost for... 

    ByteDance

    Seattle, WA
    6 days ago
  • $191k - $253k

     ...that platform so the documentation pipeline handles more work at lower cost and lower latency.Role SpecificsAs the Applied LLM Systems Engineer, your mission is to design, build, and operate production AI systems that improve how technical documentation is created, transformed... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  •  ...AI Software Engineer LLM Evaluation & Automation (Remote) We are looking for an AI Software Engineer for a B2B high-tech company. In this role, you will build and integrate evaluation harnesses and automation for software development use cases, including turning... 
    Remote work

    Stage 4 Solutions Inc

    Remote
    3 days ago
  •  ...Hiring: AI / LLM Developer (Florida – In-Person Collaboration)I'm seeking an experienced AI / LLM developer based in Florida for...  ...extensibility and future customization in mindThis is a hands-on engineering role, not prompt engineering or theoretical research.Required ExperienceDemonstrated... 
    Contract work
    Local area

    Caring Transitions

    Boca Raton, FL
    4 days ago
  •  ...We are looking for an LLM Ops Engineer with deep Databricks experience to build, automate, and scale our machine learning delivery pipelines on the Lakehouse. You’ll own the model lifecycle end‑to‑end—from data ingestion and feature engineering to CI/CD, deployment, monitoring... 

    Health Business Solutions LLC

    Cooper, FL
    more than 2 months ago
  •  ...looking for someone who "has experience with AI" we're looking for someone who builds with it every day. An engineer whose default instinct is to reach for an LLM, an agent framework, or a durable workflow instead of writing another for-loop. This is a high-ownership... 
    Full time
    Live in
    Immediate start
    Weekend work
    1 day per week

    Tradeify

    Remote
    21 days ago
  • Anduril Industries seeks an Applied LLM Systems Engineer to design, build, and operate production AI systems that improve technical documentation workflows. You will develop AI-assisted authoring, automated publishing, and multi-agent tooling for the Technical Publications... 

    Anduril-1

    Costa Mesa, CA
    3 days ago
  • $87.95k - $203.95k

     ...want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a API LLM Integration / ReactJS Engineer - Hybrid to join our team in Santa Clara, California (US-CA), United States (US).This API LLM Integration Engineer... 
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full...  ...engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $170k - $245k

     ...Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an... 
    Work at office

    Anyscale

    San Francisco, CA
    4 hours ago
  •  ...Senior LLM Engineer Location(s) Austin, Texas | Fremont, California | Lisle, Illinois Eligible for remote Yes Company Molex Career Field Data & Analytics Job Number 192106 As a Koch company, Molex is a leading supplier of connectors and interconnect components,... 
    Remote work

    Koch Industries

    Lisle, IL
    15 hours ago
  •  ...As an LLM Systems / AI Agent Engineer , you will focus on building and evolving production AI agents on foundation models, covering orchestration, context engineering, evaluation pipelines, and production observability and monitoring. This is an LLM systems and AI agent... 
    Full time
    Summer work
    Immediate start
    Shift work

    Smartworking

    Remote
    28 days ago
  • $81.34k - $141.5k

    LLM Engineer As LLM Engineer, you will make an impact by designing, building, optimizing, and deploying Large Language Model (LLM) and Small Language Model (SLM) solutions that power the enterprise Agent Factory. You will be a valued member of the AI Engineering team and... 
    Remote job
    Temporary work

    Latitude

    Brooklyn, NY
    4 days ago
  •  ...LLM Inference Engineer San Francisco or Remote Locations: San Francisco or Remote About The Role The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient... 
    Remote work

    NEAR.AI

    United States
    2 days ago
  •  ...To support the Army VANTAGE platform, the full-time LLM Engineer will focus on LLM integration and fine-tuning, prompt engineering, and deploying AI/ML solutions while working remotely. Key responsibilities Collaborate with the OptiSource development team to implement... 
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    1 day ago
  •  ...LLM Engineer Remote- Local to Chicago, IL preferred 12 months Video VISA RESTRICTIONS Required Skills: +3 Years Python, Client, AWS Cloud Computing Experience Familiarity with agent-based LLM frameworks (e.g.,LangChain) TensorFlow, PyTorch, and Hugging Face Transformers... 
    Local area
    Remote work

    RIT Solutions

    Tampa, FL
    1 day ago
  •  ...Role Mission As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses... 
    Work at office

    Hippocratic AI

    Menlo Park, CA
    2 days ago
  • $85k - $115k

     ...LLM Engineer The LLM Engineer serves as the organization's AI technical lead responsible for designing, implementing, and optimizing Large Language Model (LLM) solutions that automate business processes, improve operational efficiency, and support strategic initiatives... 
    Full time
    Monday to Friday

    USA Vein Clinics

    Northbrook, IL
    1 day ago
  •  ...Job Title: LLM Engineer (Large Language Model Engineer) Job Summary: We are seeking a highly skilled LLM Engineer to design, develop, fine-tune, and deploy Large Language Model (LLM) based applications and AI-powered solutions. The ideal candidate... 
    Full time
    Remote work

    Ova Technologies

    New York, NY
    2 days ago
  • $184k - $287.5k

    We are now looking for a Senior High-Performance LLM Training Engineer! NVIDIA is seeking experienced engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing... 
    Work experience placement

    NVIDIA

    Santa Clara, CA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Engineer. Be the first to apply!