Staff Inference Systems Engineer — High-Throughput AI
Kindredventures
Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm. Ideal candidates have a track record in building efficient inference stacks, GPU-aware optimization, and deep learning frameworks like PyTorch or JAX. Collaboration with researchers and a focus on reliability are essential. #J-18808-Ljbffr Kindredventures
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...Suggested
- Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms, architectures, and scheduling for low-latency, high-throughput inference. You will implement changes in production-grade inference engines, including kernel backends...Suggested
- ...Product Manager to own LiveKit Inference—the gateway for developers to... ...access the best models for voice AI through a single integration. The PM team is small and high-leverage, and Inference is a... ...and roadmap, partner with engineering, manage model providers and deployment...SuggestedRemote job
- ...building an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting high-impact bets with the research... ...alongside Forward Deployed Engineers to deploy and tune models, while...Suggested
- Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while focusing...Suggested
- Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission... ...You will also build LV routing systems to dispatch workloads with... .../compute trade-offs in LLM inference stacks. You will contribute to...
- SemiAnalysis is seeking a highly motivated Member of Technical Staff to join our growing engineering team. You will contribute to inference benchmarks, system modeling, and the preparation of technical... ...a global team across multiple AI and hardware domains. #J-18808-Ljbffr...Daily paid
- A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving... ...candidates should have strong software engineering skills and experience with ML inference...
- OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to run... ...kernels, compilers, and model execution. You will develop high‑performance kernels, improve compiler support, and ensure...
- ...is seeking a Member of Technical Staff to own coverage of the serverless inference landscape. You will benchmark endpoints... ...and keep us ahead in frontier AI benchmarking. You will analyze metrics... ...frontiers, collaborating with engineers and industry leaders. #J-18808-Ljbffr...
- ...the da Vinci surgical system and Ion—have transformed... ...worldwide.We’re a team of engineers, clinicians, and... ...Systems GPU Engineer - AI & Robotics, you will be... ...improving and integrating high performance robotic AI... ...models to real-time onboard inference—while serving as a core...Local areaWorldwideFlexible hours
- ...Responsibilities As a senior Machine Learning Systems Engineer on the Search Platform team, you will... .... Contribute to the architecture of high-throughput, low-latency search systems that meet... ...workflows. Partner with Rovo and AI platform teams to evolve search infrastructure...Work at officeLocal area
- Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large... ..., influencing latency, throughput, and reliability of RL and... ...enabling fast inference and scalable RL iteration,...
- OpenAI in San Francisco is seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime contexts. You will design the OpenAI-hosted control plane and partner-side software, focusing on reliable long...
- ...powers mission-critical inference for the world's most dynamic AI companies, like... ...build the platform engineers turn to to ship AI... ...the global operating system for distributed,... ...converging. The massive throughput of H100, B200, and... ...experience with high-performance networking...Full timeFlexible hours
- ...building the world's most efficient software for inference and agent hosting. In this role, you'll be one of the first engineers on Sailboxes, contributing across the stack—... ...networking stack to building large-scale systems that maximize efficiency. You'll focus on distributed...Work at office
- Gimlet Labs, Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference... ...Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance...
$200.8k - $251k
A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning... ...for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch...Full time- Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is... ...enterprises who are building AI systems to power magical experiences... ...). Improve training throughput and stability on multi-node clusters... ...configurations support high-performance training. Investigate...Full timeWork at officeRemote workFlexible hours
- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production...
- ..., California | Primarily On-site We are seeking a highly technical Enterprise Agent Systems Engineer to build and deploy advanced agent systems for large... ...role combines strong software engineering with applied AI research judgment. You will work directly with...
- ...site We are seeking an Multimodal Intelligence Systems Engineer to join an early-stage technology company building AI systems for real-world industrial environments.... ...Strong software engineering fundamentals and a high level of technical ownership. Candidate...Work at officeVisa sponsorship
- ...We are seeking a Full-Stack Engineer to build the product and infrastructure... ...around an advanced wearable AI platform used in real-world... ...wearable devices with AI systems. Design systems that operate... ...-constrained applications. High-growth early-stage software products...
- ...building the observability layer for AI-assisted software development, measuring how much of an engineering org's code is actually written... ...What You'll DoBuild the core systems that detect and attribute AI-... ...to the userHigh-agency, high-responsibility individuals with...Work at officeLocal area
$124k - $280k
...people in data and analytics engineering focus on leveraging advanced technologies... ...and implementing advanced AI and ML solutions to drive... ...algorithms, models, and systems to enable intelligent decision... ...ability to develop and sustain high performing, diverse, and inclusive...Full timeH1b$293k - $385k
...operating deeply technical systems that must work reliably... ...is seeking a Security Engineer, Host Assurance to help... ...excited by ambiguous, high-impact problems at the... ...preserving deployment throughput and operational reliability... ...OpenAIOpenAI is an AI research and deployment...Work at officeLocal areaRelocation packageFlexible hours$190k - $250k
...flagship product—an AI-driven, non-... ...patients worldwide. As a Staff Application... ...writing and tuning high-performance queries... ...Own OLTP Database Systems: Take hands-on ownership... ...— for high-throughput, low-latency web services... ...; partner with engineering and SRE teams to...Work experience placementLocal areaWorldwideRelocation$120k - $155k
Simbe is building the AI powered operating system for physical retail. Our autonomous... ...is looking for a Data Engine & Annotation Systems Engineer... ...workflows, and tooling that power high quality training data for... ..., edge case handling, and throughput monitoring. Build data...$264.8k - $331k
...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAIAI... ...the development of AI applications. For 9 years, Scale... ...and optimize our training and inference framework.Post-train state of... ...decisions. Our products provide the high-quality data and full-stack...Full time- ...AI Systems EngineerTransluce is a fast-moving research lab building... ...for an exceptional AI systems engineer to lead the design and development... .... As an early member of a highly collaborative team, you will... ...checkpointsInterpretability: Inference stacks that are as performant...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Inference Systems Engineer — High-Throughput AI. Be the first to apply!
- senior staff systems engineer San Francisco, CA
- staff data engineer San Francisco, CA
- assistant engineer San Francisco, CA
- assistant electrical engineer San Francisco, CA
- engineering aide San Francisco, CA
- software engineer staff San Francisco, CA
- staff design engineer San Francisco, CA
- senior staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA



