On-Prem AI Engineer: LLaMA/Mistral, GPUs & RAG
Jobleads-US
Texas Integrated Services seeks experts to deploy and maintain on-prem AI infrastructure for healthcare-adjacent clients. You will work across GPU deployment, model quantization, vector databases like Qdrant, and RAG pipelines, while safeguarding data and staying within client networks.
The role requires hands-on server hardware setup and a thoughtful approach to data leaving the building. Travel within Texas for install days is required; familiarity with HIPAA and client privacy is important to
#J-18808-Ljbffr Jobleads-USVacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the On-Prem AI Engineer: LLaMA/Mistral, GPUs & RAG in Kentucky vacancy
- ...Texas Integrated Services builds on-prem AI servers for small Texas businesses so they can run LLaMA, Mistral, and similar models on hardware they own. You'll deploy, tune... ...model quantization, vector databases (Qdrant), RAG pipelines, and hands-on server hardware setup. The...Suggested
- ...Join our AI team to deploy and fine-tune open-source... ...into reliable, on-prem AI products for the financial... ...function calling, and RAG against bank data and... ...’s in CS, Software Engineering, AI/ML, Data Science, or... ...Bonus: open-source LLMs (Llama, Mistral, Qwen), orchestration frameworks...Suggested
$125k - $250k
...East West Bank is seeking an experienced Senior AI Engineering to design, build, and operationalize... ..., AWS Bedrock, and open-source models such as Llama or Mistral. Practical experience with prompt engineering, RAG, embeddings, vector databases, LLM orchestration...Suggested- ...SpaceX in Hawthorne, CA is seeking a Software Engineer for AI Infrastructure (Starshield). You’ll design, operate, and scale the on-prem GPU and AI infrastructure to support critical national security missions, collaborating with AI engineers and cross-functional teams...Suggested
$130k - $180k
...development company delivering cloud, AI, data, and enterprise... ...Title Conversational AI Engineer Location: Remote (U.S.)... ...Retrieval-Augmented Generation (RAG) solutions. Build scalable... ...Anthropic Claude, Google Gemini, Llama, or Mistral . ~ Strong experience with...SuggestedFull timeH1bLocal areaRemote workVisa sponsorship- InterImage is looking for engineers who thrive where innovation meets... .... You'll take Generative AI concepts from proof of concept... ...-Augmented Generation (RAG) solutions using vector databases... ...Models including GPT, Llama, Claude, Mistral, or similar models. Experience...
$150k - $200k
...Applied AI Engineer Soulside AI · US On-Site · Reports to the CTO About Soulside... ...fine-tuning open-source models (e.g., Llama, Qwen, Mistral) using SFT, LoRA/PEFT, or preference-based... ...handling sensitive clinical data. RAG systems, retrieval quality tuning, or...H1bImmediate startRemote workVisa sponsorshipFlexible hours- ...We are seeking a highly skilled AI Engineer with hands-on experience in agent development, Retrieval-Augmented Generation (RAG), agentic workflows , and platforms such as Cursor AI . The ideal candidate will have a strong developer mindset and expertise in integrating...Permanent employmentContract workLocal area
- ...About the Role Join our engineering team at the forefront of applied AI to design and build production-grade ML... ...retrieval-augmented generation (RAG) pipelines including document ingestion... ...distilling open-source models (LLaMA, Mistral, etc.) Experience with multi-...
- ...Reuters is seeking a Senior Software Engineer, AI, to build AI-driven software for professionals. You will implement AI orchestration, RAG pipelines, and AI agents, while shipping robust backend APIs and modern UIs. The role emphasizes collaboration with AI researchers...
- Kriya Therapeutics, Inc. is seeking an AI Engineer II/III to design and operate agentic AI systems, RAG pipelines, and data integrations that power production AI capabilities. You’ll collaborate with a broad cross-functional team to translate scientific and business needs...
- Dutech Systems in Austin, TX seeks a senior AI solutions architect to design, develop, and deploy Generative AI and automation solutions... ...AI software. You will lead end-to-end AI/ML pipelines, apply RAG architectures with vector databases, and collaborate across teams...
- BSPS is seeking an AI Software Engineer in Anchorage, AK to design and implement AI-powered capabilities across LogIT applications. You will build... ...integrate AI features via APIs and microservices, maintain RAG pipelines, monitor LLM performance, and ensure data...Remote job
- ...25, 2026 Senior | Agentic AI & Applied ML US-based — CA... ...research on topics including: engineering-diagram and technical-document... ...Serving: Open-weight LLMs (Llama, Mistral, Qwen, or similar), vLLM/TGI/... ...similar) Data & Retrieval: RAG pipelines, vector databases,...Part timeRemote work
- ...Thomson Reuters in the United States is seeking a Senior Software Engineer, AI, to build AI-driven software for professionals through CLEAR. You will design multi-component pipelines, RAG architectures, and autonomous agents as part of scalable backend and frontend features...
- ...TLDR: We're looking for an AI Red Team Engineer to break LLM-powered systems responsibly, automate... ...leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others... ...-powered systems: chatbots, copilots, RAG pipelines, AI agents, tool-calling...Local area
- ...REI Systems is seeking a mid-senior AI/ML & LLM Engineer to support technology modernization and digital transformation initiatives for the FDA... ...Python, ML, and LLM engineering skills with practical experience building production-quality RAG, #J-18808-Ljbffr Jobleads-US
$102k - $170k
## AI EngineerApply: US - Remote (Any location): Full time: Posted... ...with data scientists, platform engineers, and product teams to iterate... ...experience in software, data, RAG architecture/vector databases or... ...architecting scalable, hybrid (on-prem + cloud) AI/ML solutions end-to...Full timeTemporary workRemote workFlexible hours- S27a in Washington, DC is seeking an AI/ML Engineer to advance mission-critical government initiatives and help build secure, scalable AI infrastructure. You will leverage LLMs, RAG, and multi-agent orchestration to deliver production-ready ML systems that meet strict...
- Typeform is seeking a Senior AI Engineer to build and evolve the AI capabilities behind Research Flow and other AI products. You will ship... ...Analytics to turn ideas into scalable AI solutions using Python, LLMs, RAG, and modern data tooling. This is a remote role for US-based...Remote job
- HTC Global Services, Inc. is seeking an AI/ML Engineer to build intelligent data products and production RAG systems across data platforms and cloud engineering. You will design and implement architectures that handle large-scale structured and unstructured information...
- Applied Research Solutions (ARS) is seeking an AI Systems Engineer to support, operate, and enhance AI platforms across cloud and on‑prem environments. This role emphasizes operational support, platform administration, and secure governance within a CMMC Level 2 framework...
- ...design, build and scale intelligent AI, digital, cybersecurity, cloud... ...is hiring a Senior AI Security Engineer to help secure the organization... ...cloud-based, hybrid or on-prem solutions. Deep knowledge... ...against LLMs, agentic workflows, or RAG systems Strong programming...
$110.3k - $183.8k
...seeking a highly skilled and motivated AI Full Stack Software Engineer to join our innovative team at CMM. In... ..., image parsing, text to speech, RAG) into new and legacy codebases Design... ...performance Fine-tune & serve GPT-family, Llama Develop AI agents Semantic Kernel...- Dover is seeking a Software Engineer specialized in AI/ML applications to independently drive end-to-end lifecycle management of AI-powered systems, blending advanced AI development with robust software engineering and automation. You will deploy and scale open-source...
- ...SpaceX is seeking a Sr. Software Engineer focused on AI infrastructure for Starshield. This role designs, operates, and scales GPU-enabled on-premises infrastructure and Kubernetes-based AI clusters to support national security missions. You will automate deployments...
- **Job Title - Senior AI Engineer Document Intelligence & LLM Infrastructure**Nature: ContractTime Zone : US ShiftWorking Hours : 5... ...Hugging Face TGI§ GPU-based LLM serving§ Open-source LLMs (LLaMA, Qwen, Mistral, etc.)○ Experience building deterministic validation systems...Remote workShift work
- # Agentic AI EngineerBenchlingSan Francisco, United StatesPosted 8 Aug 2026Benchling is looking for a founding engineer to join their Intelligence Engineering & Enablement team to build agentic... ...agentic frameworks* Experience with RAG and vector databases* Strong systems...
- ...States Job Description StatusNeo is a global AI-native transformation firm helping enterprises design, engineer, and govern AI-led systems with trust at the core... ...Build scalable Retrieval-AugmentedGeneration (RAG) pipelines using enterprise knowledge sources....Work experience placement
- ...The AI Engineer III serves as the primary institutional owner and operational steward of the NebulaONE generative AI platform (TulaneAI)... ...agentic AI system design Familiarity with API integration, RAG (Retrieval-Augmented Generation) agents, and low-code/no-code AI...Full timeWork at officeLocal areaImmediate startShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to On-Prem AI Engineer: LLaMA/Mistral, GPUs & RAG. Be the first to apply!

