LLM Inference Engineer
NEAR.AI
NEAR AI was started by Illia Polosukhin, co-author of the landmark paper Attention Is All You Need. We are building fast, efficient, decentralized and confidential machine learning infrastructure to enable user‑owned AI. Our mission is to build highly scalable and efficient infrastructure for open‑source AI at a global scale. We are specifically seeking an expert in high‑performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served. What You’ll Be Doing: Architect and maintain production high‑traffic LLM serving systems. Optimize throughput, latency, and cost for leading open‑source LLMs. What We’re Looking For: Strong hands‑on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT. Deep knowledge of state‑of‑the‑art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc. Proven track record in designing and maintaining end‑to‑end high‑traffic LLM serving systems. Strong problem‑solving skills and ability to communicate technical ideas clearly. We’d Love If You Have: Experience with Trusted Execution Environments (TEE). Active contributor to open‑source LLM inference engines. Please let us know if you require any special requirements for your interview and we’ll do our best to accommodate. #J-18808-Ljbffr
$170k - $245k
...Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an...SuggestedWork at office$160k - $230k
...infrastructure to enable efficient and scalable inference for large language models (LLMs). Our... ...anInference Frameworks and Optimization Engineer to design, develop, and optimize... ...unique opportunity to shape the future of LLM inference infrastructure, ensuring scalable...SuggestedFull time- B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a...Suggested
- Reflection in San Francisco is looking for a data engineer dedicated to enhancing data quality for LLM pre-training. Your role will involve collaborating with world-class researchers to establish high standards for data collection and processing. You will design automated...Suggested
$249.5k - $273.5k
...Product, Applied Research, Design, and Engineering leadership, you will lead a small team of... ...pragmatically: Evaluate emerging agent frameworks, inference optimization techniques, retrieval... .... Experience in AI, ML platforms, LLM systems, or agentic architectures is strongly...SuggestedWork at office$142.2k - $204.6k
P-1284About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks... ...and production, ensuring our large language model (LLM) serving systems are fast, scalable, and efficient. Your work...Local areaWorldwide$264.8k - $331k
...enterprise clients. As an ML Sys Research Engineer, you'll work on building out the... ...Build, profile and optimize our training and inference framework. Post-train state of the art models... ...Ideally you'd have: At least 1-3 years of LLM training in a production environment...Full time- ...deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity... ...infrastructure management. Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks is a significant plus....Worldwide
$300k
...growing group of committed researchers, engineers, policy experts, and business leaders working... .... About the role Our Inference team is responsible for building and maintaining... ..., or traffic management systems LLM inference optimization, batching, and caching...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours- ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies... ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE... ...runtime that powers large-scale LLM inference across our platform. We operate...Full timeFlexible hours
- ...About the Team OpenAI’s Inference team powers the deployment of our most advanced models -... ...world. We're a small, fast-moving team of engineers focused on delivering a world-class developer... ...inference tooling like vLLM, TensorRT-LLM, or custom model parallel systems. Own...Full time
- ...We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises... ...in low-latency, high-throughput inference for OCR and multimodal models. Own profiling... ...model graphs Evaluate vLLM, TensorRT LLM, and Triton tradeoffs Implement autoscaling...Full timeWork at officeVisa sponsorshipRelocation package
- ...more. We focus on high-performance model inference and accelerating research through... ...inference systems. In this role, you’ll lead engineering efforts to ensure our largest models run... ...with TPU, AMD GPUs, ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron, MPI, or Horovod....Full time
- ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies... ...Conviction. Join us and help build the platform engineers turn to to ship AI products. At... ...AI hardware. We believe that as LLM and multi-modal workloads scale, the network...Full timeFlexible hours
$189.6k - $237k
...framework for large language model training and inference. The platform has been powering MLEs,... ...automatic training and evaluation of LLM's, as well as evaluation of data quality... ...distributed ML systemsStrong software engineering skills, proficient in frameworks and tools...Full time$300k
...growing group of committed researchers, engineers, policy experts, and business leaders working... .... About the Role The Cloud Inference team scales and optimizes Claude to serve... ...environments Strong familiarity with LLM inference optimization, batching, caching...Full timeWork at officeVisa sponsorshipFlexible hours- ...Role At Mach9, Sensor Data Integration Engineers build the algorithms and pipelines that transform... ...up agent harnesses — orchestrating LLM-driven workflows for triage, debugging,... ...data pipelines that feed ML training and inference. Familiar with C++. About Mach9...Full timeWork at officeWork from home
$300 per month
...About the RoleAt Crusoe, our Production Engineering team ensures the reliability and scalability... ...with a focus on serving and scaling LLM workloadsDefine, measure, and improve SLIs... ...to optimize large-scale training and inference clustersAutomate observability by building...Temporary work$300 per month
...About the Role:At Crusoe, our Production Engineering team ensures the reliability and... ...services with a focus on serving and scaling LLM workloadsBuild automation and reliability... ...to support distributed AI pipelines and inference servicesDefine, measure, and improve SLIs...Temporary work$187.5k - $247.5k
...all from batteries we already have. Staff Mechanical Design Engineer, EPC Redwood Materials is hiring for a Staff Mechanical... ...personal records, professional or employment information, and inferences drawn from your PI. We collect your PI for our purposes, including...Full timeWork experience placement- ...generation code review tool powered by large language models. We are seeking a Senior Full-Stack Engineer to lead the development of a proof-of-concept prototype that integrates LLM-based analysis into existing CI/CD workflows. Responsibilities: - Design and implement...Full timeContract workRemote work
$206.4k - $379.1k
...’s Generative AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain.... ...generative models into Adobe’s flagship products.Design and architect inference infrastructure for enterprise-scale model customization,...Full timeTemporary workLocal areaWorldwide$298k - $368k
...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently... ...hardware. You will: Design VLM/LLM model architecture and drive strong... ...Proven expertise in low-latency on-device inference techniques and a deep understanding of hardware...Full timeRemote work- Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation techniques...
$190k - $265k
...insights to improve their business. Founded by engineers — and customer-obsessed — we leap at... ...models across real-time and batch inference, powering model inference at enterprise scale... ...strategy.The impact you will have:Build LLM infrastructure powering large-scale inference...Local areaWorldwide$227.2k - $417k
...About the Role:As a Software Engineer on the ML Infrastructure team, you will collaborate... ...teams to build world-class machine learning inference platforms. These platforms power... ...serving systems that support Deep Learning, LLM, and Search models. This involves building...Full timeTemporary workLocal areaFlexible hours- ...users. The Difference You Will Make: As a staff software engineer, you will lead two areas that are critical to Airbnb community support... ...on prompt engineering to fine-tune AI capabilities for AI/LLM-driven scenarios. ~ Expertise of RAG patterns, memory routing,...Work experience placementFlexible hours
$300 per month
...Crusoe.About This Role:We are seeking a Staff Hardware Systems Engineer to strengthen Crusoe’s Hardware Systems Engineering team and... ...in-depth workload characterization studies across training and inference - dense, MoE, long-context, and multimodal models to understand...Temporary work$250k - $300k
...reliably in production. That means owning the inference stack end to end: profiling where time... ...will also work directly with customer engineering teams to tailor deployments to their... ...low latency inference.Comfort with modern LLM serving frameworks such as vLLM or SGLang...Temporary work$293k - $385k
About the RoleWe are seeking a Tokens-as-a-Service (TaaS) Engineer to help build the systems that convert large-scale infrastructure... ...benchmarking, or workload optimization.Familiarity with model porting, inference/training workloads, token economics, or compute efficiency...Work at officeLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Engineer. Be the first to apply!



