Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Engineer

NEAR.AI

LLM Inference Engineer

San Francisco or Remote Locations: San Francisco or Remote

About The Role

The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.

We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.

What You'll Be Doing

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.

What We're Looking For

  • Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
  • Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
  • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
  • Strong problem-solving skills and ability to communicate technical ideas clearly.

We'd Love If You Have

  • Experience with Trusted Execution Environments (TEE).
  • Active contributor to open-source LLM inference engines.

Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the LLM Inference Engineer in United States vacancy
  •  ...part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC), United States (US).Role Overview We are seeking an AI Infrastructure... 
    Suggested
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Charlotte, NC
    4 days ago
  • $170k - $245k

     ...Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    7 hours ago
  •  ...Role Mission As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses... 
    Suggested
    Work at office

    Hippocratic AI

    Menlo Park, CA
    2 days ago
  •  ...Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack:...  ...unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly... 
    Suggested

    d-Matrix

    Santa Clara, CA
    7 hours ago
  • $224k - $356.5k

    NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for...  ...provides visibility into model behavior, inference performance, reliability, and cost across...  ...performance signals across large scale LLM and VLM deployments. It will help engineers... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...complex use cases (e.g., forecasting models, LLM-based solutions), while refining...  ...Expertise in ML model development, data engineering, and software engineering principles.Knowledge...  .../ML contexts.• Design and implement LLM inference serving stacks using: o vLLM, TensorRT-... 
    Full time
    Temporary work
    Relocation

    Infosys Technologies

    Charlotte, NC
    1 day ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency...  ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • General Information Job Title ML Staff Engineer - LLM & Production Systems Job ID 107242 Work Areas Technology & Engineering...  ...efficiencyDevelop deployment strategies and optimize inference workloads through capacity planningOwn ML Platforms End-to-EndLead... 
    Permanent employment
    Full time
    Work at office
    Local area
    1 day per week

    Bain & Company

    Boston, MA
    4 days ago
  • $236k - $330k

     ...Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization.Our mission is to build the next...  ...configurations with minimal manual tuning. We embrace AI-native engineering, using AI not only as the workload we optimize, but also as a... 
    Shift work

    Snowflake

    Bellevue, WA
    3 days ago
  • $200k - $420k

     ...from scratch: personal hardware for local inference, bespoke training infrastructure, next-...  ...research. Who we are We are scientists, engineers, and builders from the industry's top...  ...Experience extending SGLang, vLLM, TensorRT-LLM, or similar frameworks. Work on... 
    Full time
    Local area
    Visa sponsorship
    Relocation package

    River AI Inc.

    Palo Alto, CA
    3 days ago
  •  ...Title:  Applied AI Engineer — Inference & Agent Systems Location: United States What We're Building Arcana is building AI agents that...  ...enforcement: JSON schema validation, retry loops on malformed LLM output, graceful degradation - Tool call design: schema... 
    Full time

    Arcana Analytics

    United States
    1 day ago
  • ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-performance KV cache layer for LLM inference across GPUs and nodes. This role focuses on reducing latency, increasing throughput, and lowering cost for... 

    ByteDance

    Seattle, WA
    6 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Full time

    Nvidia

    Seattle, WA
    4 days ago
  • $193.3k - $261.5k

    We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational AI. This is a full-stack inference role...  ...-time streaminggenerative models that fall outside standard LLM serving patterns — sustained low-latencyoutput under concurrent... 
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    3 days ago
  •  ...Hiring: AI / LLM Developer (Florida – In-Person Collaboration)I'm seeking an experienced...  ...customization in mindThis is a hands-on engineering role, not prompt engineering or...  ...or agent frameworksExperience with local inference, persistence, and memory systemsComfortable... 
    Contract work
    Local area

    Caring Transitions

    Boca Raton, FL
    4 days ago
  •  ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated...  ...head-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed explanations of why... 

    AMD

    Santa Clara, CA
    2 days ago
  • $197.3k - $225.1k

     ...Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    3 days ago
  • $167k - $209k

     ...builders in the world. We are seeking a Senior Engineer II to implement and contribute to the design and optimization of our Serverless Inference infrastructure and APIs. In this role,...  ...Knowledge: Understanding of modern LLM serving architectures and familiarity with... 
    Full time
    Local area
    Worldwide
    Flexible hours

    DigitalOcean

    Seattle, WA
    4 days ago
  • $184k - $287.5k

     ...are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who...  ...:Implement language and multimodal model inference as part of NVIDIA Inference Microservices...  ...bugs and deliver production code to TRT-LLM, NVIDIA’s open-source inference serving library... 
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  •  ...Ventures. About The Role In this role, you'll own the core LLM infrastructure powering two products redefining B2B research:...  ...translate user needs into technical solutions while maintaining engineering best practices Nice To Have RAG systems, embeddings,... 
    Full time
    Immediate start
    Flexible hours

    Newtonx

    United States
    1 day ago
  • $200.8k - $251k

     ...a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251... 
    Full time

    Scale AI

    Seattle, WA
    3 days ago
  • $193.4k - $240.8k

     ...Requirements: ~ Bachelors degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or a related discipline plus a minimum...  ...of developing AI and ML algorithms or technologies (e.g., LLM Inference, Similarity Search, VectorDBs, Guardrails, Memory) using... 
    Full time

    Capital One

    Charlottesville, VA
    15 days ago
  •  ...Job Title: LLM Engineer (Large Language Model Engineer) Job Summary: We are seeking a highly skilled LLM Engineer...  ...). Knowledge of AI evaluation, model optimization, and inference strategies. Strong problem-solving and system design... 
    Full time
    Remote work

    Ova Technologies

    New York, NY
    2 days ago
  •  ...Job Details Job Title: LLM Engineer Location: Cincinnati, OH Work Location: Remote - USA Duration: 1 year...  ...Private LLM Hosting On-Prem Model Deployment GPU-Based Inference Model Serving APIs High-Availability Inference... 
    Remote work

    Talent Software Services

    United States
    2 days ago
  •  ...Texas Sports Academy is on the lookout for a Senior AI Engineer specializing in LLM (Large Language Model) Systems and RAG (Retrieval-Augmented Generation) Optimization. As we continue to push the boundaries of sports technology, your role will be pivotal in developing... 
    Remote job
    Full time

    Texas Sports Academy

    United States
    1 day ago
  • $85k - $115k

    Vein Clinics of America, Inc. is seeking an experienced LLM Engineer to serve as the AI technical lead. In this full-time, on-site role in Northbrook, IL, you will design, implement, and optimize LLM solutions that improve business processes. Responsibilities include developing... 
    Full time

    Vein Clinics of America, Inc.

    Northbrook, IL
    2 days ago
  • B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a... 

    B Capital

    San Francisco, CA
    3 days ago
  • ID.me is hiring a Staff Software Development Engineer to lead the Tools Team in McLean, VA. You will champion AI-native engineering, architect...  ...has 10+ years in full-stack engineering, including hands-on LLM integration experience. This on-site role fosters collaboration... 

    ID.me

    Mc Lean, VA
    3 days ago
  •  ...latency, and robustness. Build shared APIs and platform components used broadly across engineering teams. Key Responsibilities Design and implement orchestration patterns for LLM-powered agents. Evaluate and select models, tools, and providers based on... 
    Full time

    Calliere

    Remote
    1 day ago
  • Siamo alla ricerca di un/una AI Engineer da inserire nel nostro team dedicato alla trasformazione digitale dei processi del settore bancario...  ...di automazione basati su Node.js Integrazione di modelli AI (LLM) all’interno dei processi aziendali Progettazione e implementazione... 

    Gruppo MOL

    New York State
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Engineer. Be the first to apply!