Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Inference Systems Engineer — High-Throughput AI

Kindredventures

Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm. Ideal candidates have a track record in building efficient inference stacks, GPU-aware optimization, and deep learning frameworks like PyTorch or JAX. Collaboration with researchers and a focus on reliability are essential. #J-18808-Ljbffr Kindredventures

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Staff Inference Systems Engineer — High-Throughput AI in San Francisco, CA vacancy
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 
    Suggested

    Causal Labs

    San Francisco, CA
    1 day ago
  • Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms, architectures, and scheduling for low-latency, high-throughput inference. You will implement changes in production-grade inference engines, including kernel backends... 
    Suggested

    Together

    San Francisco, CA
    2 days ago
  •  ...Product Manager to own LiveKit Inference—the gateway for developers to...  ...access the best models for voice AI through a single integration. The PM team is small and high-leverage, and Inference is a...  ...and roadmap, partner with engineering, manage model providers and deployment... 
    Suggested
    Remote job

    LiveKit

    San Francisco, CA
    5 hours ago
  •  ...building an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting high-impact bets with the research...  ...alongside Forward Deployed Engineers to deploy and tune models, while... 
    Suggested

    Mixpeek

    San Francisco, CA
    5 hours ago
  • Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while focusing... 
    Suggested

    Sail Research

    San Francisco, CA
    3 days ago
  • Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission...  ...You will also build LV routing systems to dispatch workloads with...  .../compute trade-offs in LLM inference stacks. You will contribute to... 

    Sail

    San Francisco, CA
    2 days ago
  • SemiAnalysis is seeking a highly motivated Member of Technical Staff to join our growing engineering team. You will contribute to inference benchmarks, system modeling, and the preparation of technical...  ...a global team across multiple AI and hardware domains. #J-18808-Ljbffr... 
    Daily paid

    S27a

    San Francisco, CA
    5 hours ago
  • A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving...  ...candidates should have strong software engineering skills and experience with ML inference... 

    Gimlet Labs

    San Francisco, CA
    1 day ago
  • OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to run...  ...kernels, compilers, and model execution. You will develop high‑performance kernels, improve compiler support, and ensure... 

    Slope

    San Francisco, CA
    5 hours ago
  •  ...is seeking a Member of Technical Staff to own coverage of the serverless inference landscape. You will benchmark endpoints...  ...and keep us ahead in frontier AI benchmarking. You will analyze metrics...  ...frontiers, collaborating with engineers and industry leaders. #J-18808-Ljbffr... 

    Artificial Analysis, Inc.

    San Francisco, CA
    2 days ago
  •  ...the da Vinci surgical system and Ion—have transformed...  ...worldwide.We’re a team of engineers, clinicians, and...  ...Systems GPU Engineer - AI & Robotics, you will be...  ...improving and integrating high performance robotic AI...  ...models to real-time onboard inference—while serving as a core... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    2 days ago
  •  ...Responsibilities As a senior Machine Learning Systems Engineer on the Search Platform team, you will...  .... Contribute to the architecture of high-throughput, low-latency search systems that meet...  ...workflows. Partner with Rovo and AI platform teams to evolve search infrastructure... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    2 days ago
  • Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large...  ..., influencing latency, throughput, and reliability of RL and...  ...enabling fast inference and scalable RL iteration,... 

    Magic AI, Inc

    San Francisco, CA
    3 days ago
  • OpenAI in San Francisco is seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime contexts. You will design the OpenAI-hosted control plane and partner-side software, focusing on reliable long... 

    Slope

    San Francisco, CA
    5 hours ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  ...build the platform engineers turn to to ship AI...  ...the global operating system for distributed,...  ...converging. The massive throughput of H100, B200, and...  ...experience with high-performance networking... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  •  ...building the world's most efficient software for inference and agent hosting. In this role, you'll be one of the first engineers on Sailboxes, contributing across the stack—...  ...networking stack to building large-scale systems that maximize efficiency. You'll focus on distributed... 
    Work at office

    Sail Research Inc.

    San Francisco, CA
    15 hours ago
  • Gimlet Labs, Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference...  ...Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance... 

    Gimlet Labs, Inc.

    San Francisco, CA
    3 days ago
  • $200.8k - $251k

    A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning...  ...for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch... 
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is...  ...enterprises who are building AI systems to power magical experiences...  ...). Improve training throughput and stability on multi-node clusters...  ...configurations support high-performance training. Investigate... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production... 

    MakerMaker.AI

    San Francisco, CA
    5 hours ago
  •  ..., California | Primarily On-site We are seeking a highly technical Enterprise Agent Systems Engineer to build and deploy advanced agent systems for large...  ...role combines strong software engineering with applied AI research judgment. You will work directly with... 

    MaxIT Consulting - Max Corporate Group

    San Francisco, CA
    2 days ago
  •  ...site We are seeking an Multimodal Intelligence Systems Engineer to join an early-stage technology company building AI systems for real-world industrial environments....  ...Strong software engineering fundamentals and a high level of technical ownership. Candidate... 
    Work at office
    Visa sponsorship

    MaxIT Consulting - Max Corporate Group

    San Francisco, CA
    2 days ago
  •  ...We are seeking a Full-Stack Engineer to build the product and infrastructure...  ...around an advanced wearable AI platform used in real-world...  ...wearable devices with AI systems. Design systems that operate...  ...-constrained applications. High-growth early-stage software products... 

    MaxIT Consulting - Max Corporate Group

    San Francisco, CA
    2 days ago
  •  ...building the observability layer for AI-assisted software development, measuring how much of an engineering org's code is actually written...  ...What You'll DoBuild the core systems that detect and attribute AI-...  ...to the userHigh-agency, high-responsibility individuals with... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    5 days ago
  • $124k - $280k

     ...people in data and analytics engineering focus on leveraging advanced technologies...  ...and implementing advanced AI and ML solutions to drive...  ...algorithms, models, and systems to enable intelligent decision...  ...ability to develop and sustain high performing, diverse, and inclusive... 
    Full time
    H1b

    PwC

    San Francisco, CA
    5 days ago
  • $293k - $385k

     ...operating deeply technical systems that must work reliably...  ...is seeking a Security Engineer, Host Assurance to help...  ...excited by ambiguous, high-impact problems at the...  ...preserving deployment throughput and operational reliability...  ...OpenAIOpenAI is an AI research and deployment... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    5 days ago
  • $190k - $250k

     ...flagship product—an AI-driven, non-...  ...patients worldwide. As a Staff Application...  ...writing and tuning high-performance queries...  ...Own OLTP Database Systems: Take hands-on ownership...  ...— for high-throughput, low-latency web services...  ...; partner with engineering and SRE teams to... 
    Work experience placement
    Local area
    Worldwide
    Relocation

    HeartFlow

    San Francisco, CA
    5 days ago
  • $120k - $155k

    Simbe is building the AI powered operating system for physical retail. Our autonomous...  ...is looking for a Data Engine & Annotation Systems Engineer...  ...workflows, and tooling that power high quality training data for...  ..., edge case handling, and throughput monitoring. Build data... 

    Simbe Robotics, Inc.

    San Francisco, CA
    2 days ago
  • $264.8k - $331k

     ...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAIAI...  ...the development of AI applications. For 9 years, Scale...  ...and optimize our training and inference framework.Post-train state of...  ...decisions. Our products provide the high-quality data and full-stack... 
    Full time

    Scale AI

    San Francisco, CA
    5 days ago
  •  ...AI Systems EngineerTransluce is a fast-moving research lab building...  ...for an exceptional AI systems engineer to lead the design and development...  .... As an early member of a highly collaborative team, you will...  ...checkpointsInterpretability: Inference stacks that are as performant... 
    Flexible hours

    Transluce

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Inference Systems Engineer — High-Throughput AI. Be the first to apply!