Staff Engineer, AI Inference & Distributed Systems
Sail Research
Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while focusing on deep observability of our systems. The ideal candidate has a strong background in distributed systems and is eager to engage in complex challenges. Enjoy a vibrant work environment with excellent meals and a collaborative team atmosphere. #J-18808-Ljbffr Sail Research
- Sail is hiring for an engineering role in San Francisco to design and implement high-performance... ...fleet. You will also build LV routing systems to dispatch workloads with latency... ...caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep...Suggested
- Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm...Suggested
- ...in San Francisco is seeking a Member of Technical Staff to design and build distributed systems for AI workloads. The role involves developing scheduling... ...APIs. Ideal candidates should have strong software engineering skills and experience with distributed systems. This...Suggested
$150k - $350k
...Inc. is seeking a Member of Technical Staff to focus on distributed systems in San Francisco, California. This... ...and building the core platform for AI workloads, developing resource management... ...should have strong software engineering fundamentals and experience with distributed...Suggested- Harrison Clarke is working with a high-growth startup in San Francisco seeking a Staff Distributed Systems Engineer. This hands-on role involves designing and building core systems for a cutting-edge AI code generation product. The ideal candidate will have excellent coding...Suggested
- Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms, architectures, and scheduling for low-latency, high-throughput inference. You will implement changes in production-grade inference engines, including kernel backends...
$200k - $400k
A leading AI technology company located in San Francisco is seeking an infrastructure engineer to build distributed systems for their AI inference engine. The role involves designing systems that ensure minimal latency and maximum reliability. Candidates should have a...Visa sponsorship$150k - $250k
Asari AI in San Francisco is looking for a skilled individual to build the supercomputing infrastructure that runs AI agents,... ...performance workloads. Your role will involve designing cloud compute, distributed systems, and sandboxed tooling to ensure efficiency and scalability....$190.9k - $232.8k
A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The... ...requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong...$225k
Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate...- Eventual is seeking a Member of Technical Staff to build Eventual's core products and... ...autonomously solving problems. We value engineers who can scope tasks and implement efficient... ..., with a focus on performance and reliability across distributed #J-18808-Ljbffr Mixpeek
$350k
Mirendil in San Francisco is searching for an engineer to develop and optimize inference systems for cutting-edge AI models. You will handle the complete inference stack, enhancing performance and reliability. The role involves partnering with teams to deploy new architectures...- ...building the best way to talk to AI and humans together —... .... Member of Technical Staff is the title we use for engineers who own hard problems end... ...Representative projects Designing systems that give 3M+ AI agents... ...fine-tuning, evaluation, inference, or RAG at scale High-...
- ...Francisco is hiring Members of Technical Staff to build systems that accelerate LLM inference and own customer workloads end to... ...-performance kernels, inference engine internals, and production... ...collaborating with a fast-growing AI inference company. #J-18808-Ljbffr...
- B Capital is seeking a skilled engineer for GPU infrastructure in San Francisco. This... ...designing and operating high-performance systems for model inference, synthetic data generation, and... ...a passion for working in cutting-edge AI. Benefits include top-tier compensation...
- Modal is building an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting... ...work with customers alongside Forward Deployed Engineers to deploy and tune models, while expanding...
$220k
We build and run the inference engine behind every Perplexity query and deploy dozens of model... ...CUTLASS, or similar). Any other deep systems programming experience is a plus. You... ...You've built and operated production distributed systems under real load - ideally performance...$250k - $300k
...the only vertically integrated AI infrastructure company built... .... That means owning the inference stack end to end: profiling where... ...not good enough. This is core systems and performance work on some... ...work directly with customer engineering teams to tailor deployments to...Temporary work- Eventual is hiring a Member of Technical Staff in San Francisco to build the company\'s core products and architecture. You will ship... ...scope and solve difficult technical challenges. We seek engineers who combine strong coding fundamentals with practical experience...
- ...trusted decision-ready AI to the world's most... ...on. As a Staff Machine Learning Engineer, you’ll own AI-driven... ...works to the production system the mission depends... ...scale, drawing on real distributed-systems experience.... ...latency, high-concurrency inference (Triton, vLLM, GPU-...Full timeContract workRemote workFlexible hours
$250k - $285k
...only vertically integrated AI infrastructure company... ...This RoleWe’re seeking a Staff Product Security Engineer with deep AI/ML security... ...applications, infrastructure, and distributed AI systems. This is a highly... ...stack, including MLOps, inference architectures, vector databases...Temporary work$273k - $345k
...Atoms builds Physical AI- real-world robots for... ...mining, and transport. Our systems are designed to... ...We are roboticists, engineers, operators, and builders... ...Profile real-time inference pipelines to identify... ...undefined real-world data distributions. Why join us...Full timeInternshipWork at officeFlexible hours- ...a self-serve GPU compute platform for training and inference workloads. You will design and operate the system that lets researchers launch jobs across multi-cloud... ..., and a coherent platform roadmap for scalable AI workloads. #J-18808-Ljbffr United States Digital Space...
$350k
Mirendil, based in San Francisco, is searching for an engineer to work at the intersection of research and systems on their pretraining stack. The role involves implementing model architectures and scaling distributed training jobs across thousands of GPUs. We offer a...$300 per month
...only vertically integrated AI infrastructure company... ...This Role:We are seeking a Staff Hardware Systems Engineer to strengthen Crusoe’s Hardware... ...across training and inference - dense, MoE, long-context... ....Hands-on experience with distributed training and/or inference...Temporary work- ...The role: SoFi’s Staff AI Engineer is a hands-on AI engineering role... ...problems at scale. Distributed Agent Memory & State: Develop... ...high-throughput, low-latency inference across diverse hardware... ...methodologies for LLM and agentic systems. ~ Experience with cloud...Full time
$227.33k - $312.58k
We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier Models organization... ...and building the data systems that power frontier‑scale... ...with data‑intensive or distributed systems.Proven experience... ..., evaluation, or inference workflows.Solid understanding...Full timeWork at officeLocal areaImmediate start3 days per week- ...governing autonomous AI agents across industries... ...We are seeking a Staff Research Engineer, AI/ML & Cybersecurity... ...production-grade AI systems, secure model deployment... ...components Improve inference performance, observability... ..., and artifact distribution Contribute to SOC2-aligned...
$230k - $322k
...find them useful. As a Staff Machine Learning Engineer on Shopping Ads, you... ...through multiple systems and teams. Lead the... ...evaluation pipelines, online inference, and experimentation... ...and online/offline distribution shift. Experience at... ...intelligence (AI). You will have the...For contractorsWork experience placementFlexible hoursShift work- ...founding member of the Business Technologies team in San Francisco to design scalable internal infrastructure, not just modify existing systems. You’ll be the second hire, partnering with the TPM Lead to define architecture across identity, SaaS, and cloud platforms. The...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Engineer, AI Inference & Distributed Systems. Be the first to apply!
- software engineer staff San Francisco, CA
- assistant engineer San Francisco, CA
- engineering aide San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineering manager San Francisco, CA
- senior staff systems engineer San Francisco, CA
- technology administrator San Francisco, CA
- project engineer assistant project manager San Francisco, CA


