Global Inference Library Engineer
$175k - $250kLeoForce
Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $250,000 per year Job Details We’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments. What We’re Looking For Strong experience building AI/ML infrastructure, inference systems, or high-performance computing software Strong programming experience with Python and C++, Rust, or similar systems languages Experience with LLM inference frameworks and model-serving infrastructure Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies Experience developing, integrating, or optimizing performance-critical compute kernels Understanding of modern transformer and LLM architectures Familiarity with inference concepts including batching, attention, KV caching, quantization, and memory management Experience benchmarking and profiling AI workloads across different hardware environments Strong understanding of GPU or accelerator architecture and performance characteristics Experience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable A bit about us: We're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures. Why join us? Well-funded by leading tech investors Cutting edge technical problems with complex solutions Lucrative Equity in a seed stage startup Competitive compensation Excellent benefits (healthcare, vision, dental)
- techservices #c #python #gpu #rust #dataflow #optimization #library #cuda #algorithm #latency #itl #tvm #llvm #multimodal #quantization #vllm #kv-cache #kernel-variants #ml-inference #systolic-arrays #ttft #tpot #tier3
- J-18808-Ljbffr LeoForce
$175k - $250k
...Competitive compensation Excellent benefits (healthcare, vision, dental) Job Details We're looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal...SuggestedLocal area- Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from...Suggested
- ...Staff, Infrastructure to own the end-to-end cloud infrastructure for a high-performance compression API. You will build global, low-latency GPU ML inference systems in the critical path of customer traffic, driving reliability, scalability, and cost-efficiency for a...Suggested
$160k - $230k
...state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). Our mission is to optimize... ....We are seeking anInference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines...SuggestedFull time$170k - $245k
...popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber,... ...+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries...SuggestedWork at office- Sail is hiring for an engineering role in San Francisco to design and implement high-performance... ...control, queuing, and fairness across a global fleet. You will also build LV routing... ...for memory/compute trade-offs in LLM inference stacks. You will contribute to deep observability...
- Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while focusing...
- An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate...
- Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-time local intelligence. You will primarily work on uzu, our inference engine, and focus on supporting new modalities and a variety of features. The ideal candidates...Local area
- ...Francisco, on‑site ABOUT THE ROLE You build and operate the inference systems that serve our models in production. The work spans serving... ...that come with running real workloads. This is an engineering role, not a research role. You'll measure, profile, debug, and...
$300k - $400k
Growth Engineer - Globalization (North America) About OpenArt OpenArt is an AI Storytelling and Visual Creation Platform used by millions worldwide. We're building the next generation of creative tools powered by cutting-edge AI, enabling anyone to create videos, visuals...WorldwideVisa sponsorship- Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open...
- Inferact is seeking an inference runtime engineer to enhance the performance and capabilities of LLM and diffusion model serving. This role requires expertise in optimizing model execution on various hardware architectures and has significant implications for AI inference...Remote work
- ...Join a small, focused team of YC and unicorn founders and senior engineers with deep expertise in 3D, generative video, developer... ...possible. About the Role We're looking for a Founding Engineer, ML Inference with deep expertise in high-performance ML engineering. This...RelocationVisa sponsorshipRelocation package
- Help build the inference stack behind the next generation of multimodal foundation models. Most inference roles are about making existing... ...production. You’ll sit between frontier research and product engineering, designing the infrastructure that allows cutting-edge models...Work at officeRelocation package
- OpenArt is seeking a Growth Engineer, Globalization to lead engineering projects that expand OpenArt's international reach. You’ll own end-to-end globalization features across localization, onboarding, pricing, checkout, and lifecycle communications. The role emphasizes...
- OpenArt is seeking a Growth Engineer focused on globalization for North America. You will own end-to-end globalization features, including multi-language infrastructure, region-specific onboarding, pricing experiments, and localization workflows. You will collaborate with...
- Together AI is building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). We seek an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support...
$220k - $320k
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques...$167.2k - $209k
...builders in the world. DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be the... ...batch size performance using AMD's AITER library for AMD MI355X - identify and tune AITER's CK (composable...Local areaRemote workWorldwideFlexible hours$220k - $320k
A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation...Local area- TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing...
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...
- ...Its APA system, built on the industry’s first Process Reasoning Engine (PRE) and specialized AI agents, combines process discovery, RPA... ...intersection of technical excellence and executive strategy, owning the global pre-sales function and serving as a key voice in shaping company...Full timeLocal areaRemote workWorldwideFlexible hours
- ...infrastructure, influencing latency, throughput, and reliability of RL and training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies, batching, and long-context workloads while collaborating with #J-18808-...
- Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm...
$293k - $385k
...seeking a Workload Porting & Performance Engineer to evaluate new hardware platforms by... ...AI/ML workloads, including training or inference systems.Familiarity with GPU or accelerator... ...requests can be made via this link.OpenAI Global Applicant Privacy PolicyAt OpenAI, we...Work at officeLocal areaRelocation packageFlexible hours$225k
Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate...$420k
...real-world industry insight, drawing on decades of experience in engineering, construction management, scheduling and delay analysis, and... ...during project execution.This is an opportunity to work on globally significant matters within a team known for its strategic thinking...For contractorsWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Global Inference Library Engineer. Be the first to apply!
