LLM Inference Engineer
NEAR.AI
LLM Inference Engineer
San Francisco or Remote Locations: San Francisco or Remote
About The Role
The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.
What You'll Be Doing
- Architect and maintain production high-traffic LLM serving systems.
- Optimize throughput, latency, and cost for leading open-source LLMs.
What We're Looking For
- Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
- Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
- Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
- Strong problem-solving skills and ability to communicate technical ideas clearly.
We'd Love If You Have
- Experience with Trusted Execution Environments (TEE).
- Active contributor to open-source LLM inference engines.
Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.
- ...part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC), United States (US).Role Overview We are seeking an AI Infrastructure...SuggestedWork at officeRemote workFlexible hours
$170k - $245k
...Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an...SuggestedWork at office- ...Role Mission As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses...SuggestedWork at office
- ...Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack:... ...unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly...Suggested
$224k - $356.5k
NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for... ...provides visibility into model behavior, inference performance, reliability, and cost across... ...performance signals across large scale LLM and VLM deployments. It will help engineers...SuggestedFull time- ...complex use cases (e.g., forecasting models, LLM-based solutions), while refining... ...Expertise in ML model development, data engineering, and software engineering principles.Knowledge... .../ML contexts.• Design and implement LLM inference serving stacks using: o vLLM, TensorRT-...Full timeTemporary workRelocation
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work...Full time- General Information Job Title ML Staff Engineer - LLM & Production Systems Job ID 107242 Work Areas Technology & Engineering... ...efficiencyDevelop deployment strategies and optimize inference workloads through capacity planningOwn ML Platforms End-to-EndLead...Permanent employmentFull timeWork at officeLocal area1 day per week
$236k - $330k
...Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization.Our mission is to build the next... ...configurations with minimal manual tuning. We embrace AI-native engineering, using AI not only as the workload we optimize, but also as a...Shift work$200k - $420k
...from scratch: personal hardware for local inference, bespoke training infrastructure, next-... ...research. Who we are We are scientists, engineers, and builders from the industry's top... ...Experience extending SGLang, vLLM, TensorRT-LLM, or similar frameworks. Work on...Full timeLocal areaVisa sponsorshipRelocation package- ...Title: Applied AI Engineer — Inference & Agent Systems Location: United States What We're Building Arcana is building AI agents that... ...enforcement: JSON schema validation, retry loops on malformed LLM output, graceful degradation - Tool call design: schema...Full time
- ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-performance KV cache layer for LLM inference across GPUs and nodes. This role focuses on reducing latency, increasing throughput, and lowering cost for...
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...Full time$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational AI. This is a full-stack inference role... ...-time streaminggenerative models that fall outside standard LLM serving patterns — sustained low-latencyoutput under concurrent...InternshipLocal areaFlexible hours- ...Hiring: AI / LLM Developer (Florida – In-Person Collaboration)I'm seeking an experienced... ...customization in mindThis is a hands-on engineering role, not prompt engineering or... ...or agent frameworksExperience with local inference, persistence, and memory systemsComfortable...Contract workLocal area
- ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...head-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed explanations of why...
$197.3k - $225.1k
...Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer...Full timePart timeLocal area$167k - $209k
...builders in the world. We are seeking a Senior Engineer II to implement and contribute to the design and optimization of our Serverless Inference infrastructure and APIs. In this role,... ...Knowledge: Understanding of modern LLM serving architectures and familiarity with...Full timeLocal areaWorldwideFlexible hours$184k - $287.5k
...are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who... ...:Implement language and multimodal model inference as part of NVIDIA Inference Microservices... ...bugs and deliver production code to TRT-LLM, NVIDIA’s open-source inference serving library...Full time- ...Ventures. About The Role In this role, you'll own the core LLM infrastructure powering two products redefining B2B research:... ...translate user needs into technical solutions while maintaining engineering best practices Nice To Have RAG systems, embeddings,...Full timeImmediate startFlexible hours
$200.8k - $251k
...a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251...Full time$193.4k - $240.8k
...Requirements: ~ Bachelors degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or a related discipline plus a minimum... ...of developing AI and ML algorithms or technologies (e.g., LLM Inference, Similarity Search, VectorDBs, Guardrails, Memory) using...Full time- ...Job Title: LLM Engineer (Large Language Model Engineer) Job Summary: We are seeking a highly skilled LLM Engineer... ...). Knowledge of AI evaluation, model optimization, and inference strategies. Strong problem-solving and system design...Full timeRemote work
- ...Job Details Job Title: LLM Engineer Location: Cincinnati, OH Work Location: Remote - USA Duration: 1 year... ...Private LLM Hosting On-Prem Model Deployment GPU-Based Inference Model Serving APIs High-Availability Inference...Remote work
- ...Texas Sports Academy is on the lookout for a Senior AI Engineer specializing in LLM (Large Language Model) Systems and RAG (Retrieval-Augmented Generation) Optimization. As we continue to push the boundaries of sports technology, your role will be pivotal in developing...Remote jobFull time
$85k - $115k
Vein Clinics of America, Inc. is seeking an experienced LLM Engineer to serve as the AI technical lead. In this full-time, on-site role in Northbrook, IL, you will design, implement, and optimize LLM solutions that improve business processes. Responsibilities include developing...Full time- B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a...
- ID.me is hiring a Staff Software Development Engineer to lead the Tools Team in McLean, VA. You will champion AI-native engineering, architect... ...has 10+ years in full-stack engineering, including hands-on LLM integration experience. This on-site role fosters collaboration...
- ...latency, and robustness. Build shared APIs and platform components used broadly across engineering teams. Key Responsibilities Design and implement orchestration patterns for LLM-powered agents. Evaluate and select models, tools, and providers based on...Full time
- Siamo alla ricerca di un/una AI Engineer da inserire nel nostro team dedicato alla trasformazione digitale dei processi del settore bancario... ...di automazione basati su Node.js Integrazione di modelli AI (LLM) all’interno dei processi aziendali Progettazione e implementazione...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Engineer. Be the first to apply!



