Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff — Inference Infrastructure

Causal Labs

Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it. To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather. Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for infrastructure engineers who are excited to tackle unsolved problems. Progress on an LPM is gated by how fast we can evaluate it: large-scale backtesting against decades of physical observations, ensemble generation, and rollout evaluation across model scales. Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations Design and implement techniques that improve latency, throughput, and efficiency for real-time inference Optimize the inference stack to fully utilize hardware FLOPs, bandwidth, and memory Extend orchestration frameworks (e.g. Kubernetes, Ray, Slurm) for distributed inference and large-batch evaluation sweeps Establish standards for reliability, observability, and reproducibility across the inference stack, so every evaluation is trustworthy and repeatable Collaborate with researchers to enable high-performance inference for novel architectures as they emerge What we\'re looking for We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains. Experience building or optimizing inference and serving systems for throughput and latency (e.g. TensorRT) Understanding of distributed compute, GPU parallelism, and hardware-aware optimization Deep familiarity with deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures Strong engineering skills: performant, maintainable code and the ability to debug complex codebases Bonus: contributions to open-source inference or systems infrastructure (e.g. vLLM, SGLang, Triton) #J-18808-Ljbffr Causal Labs

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff — Inference Infrastructure in San Francisco, CA vacancy
  • $150k - $300k

     ...spanning cloud LLM serving, LLM inference optimization and RL systems....  ...areas are: Building the infrastructure to serve LLMs efficiently at...  ...RL training stack. Core Technical Responsibilities LLM Serving...  ...development and encourage team members to contribute to the broader... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    3 days ago
  •  ...recognize parts of inputs that are unimportant, reducing inference costs for scale-ups and enterprises that integrate...  ...is 5 people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems that serve our... 
    Suggested
    Visa sponsorship

    The Token Company

    San Francisco, CA
    2 days ago
  • $250k

     ...building the next generation of agentic infrastructure for GPU-intensive workloads....  ...opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's...  ...on orchestration intelligence, inference gateways, and agentic operations tooling... 
    Suggested
    Full time
    San Francisco, CA
    9 days ago
  • $200k

    Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important...  ..., domain-specific RL, ultra-long context, and inference-time compute to achieve this goal. About the... 
    Suggested
    Relocation
    Visa sponsorship

    Magic AI, Inc

    San Francisco, CA
    1 day ago
  • # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco / Bay Area preferred. Remote exceptional for the right person.We look...  ...AI workloads cheaper and easier to own by turning inference behavior, traces, workload replay, GPU signals, and task... 
    Suggested
    Full time
    Remote work

    Touchdown Labs, Inc.

    San Francisco, CA
    2 days ago
  • $200k - $400k

     ...to an algorithm. We're building the infrastructure to understand human behavior at scale...  ...simulating a society means running inference over populations of agents, not single...  ...to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will... 
    Live in
    Flexible hours

    Simile

    San Francisco, CA
    3 days ago
  •  ...with today’s homogeneous, vertically integrated infrastructure. Gimlet addresses this by decoupling AI workloads...  ...-class AI datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build the... 

    Gimlet Labs

    San Francisco, CA
    1 day ago
  • $150k - $300k

    Building Open Superintelligence Infrastructure Prime Intellect is building the open superintelligence...  ...for GPU Infrastructure, you'll be the technical expert who transforms customer...  ...deployment strategies for LLM training, inference, and HPC workloads Present architectural... 

    Prime Intellect

    San Francisco, CA
    2 days ago
  • $350k

     ...model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments. Our team includes researchers...  ...and MIT. The Role We are looking for an engineer to own the inference systems that power our models in production and research. You... 

    Mirendil

    San Francisco, CA
    3 days ago
  • $200k - $400k

    About The Role We're looking for an inference runtime engineer to push the boundaries of...  ...Contributions to open-source ML or system infrastructure projects. Bonus points if you have:...  ...LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM... 
    Remote work
    Visa sponsorship
    Shift work

    Inferact

    San Francisco, CA
    1 day ago
  • About Us Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates...  ...the cluster infrastructure behind Gimlet’s heterogeneous inference cloud. Unlike traditional cloud platforms built around a... 

    Gimlet Labs

    San Francisco, CA
    2 days ago
  •  ...observe their code. We are responsible for designing, building, and scaling core infrastructure that powers a high-volume data platform for AI applications. We are looking for team members who love building enabling systems that empower our engineers and power our rapidly... 
    Work at office

    LlamaIndex

    San Francisco, CA
    3 days ago
  • $225k

    About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve...  ...sits at the boundary between model execution and distributed infrastructure. You will work on systems that determine inference latency,... 
    Relocation
    Visa sponsorship

    Magic

    San Francisco, CA
    2 days ago
  •  ...and other industry leaders. The Opportunity Language model inference is the fastest-moving market in AI: dozens of providers,...  ...benchmarks are the industry’s reference point, and we’re hiring a Member of Technical Staff to drive them. You’ll own coverage of the serverless... 
    Shift work

    Artificial Analysis

    San Francisco, CA
    2 days ago
  • $350k

     ...model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments. Our team includes researchers...  ...scheduling, autoscaling, and multi-tenant isolation Training and inference infrastructure - understand the resource and scheduling... 

    Mirendil

    San Francisco, CA
    3 days ago
  • $250k - $300k

     ...development? Join one of the most exciting AI infrastructure companies in the market, building a...  ...next-generation AI training and inference at scale. This role offers the opportunity...  ...Have: ~ A track record of impressive technical work you can speak to in depth, the... 
    Full time
    Remote work
    San Francisco, CA
    9 days ago
  • Member of Technical Staff - Infrastructure Security We're partnering with a frontier AI research company that is building next-generation open-weight foundation models with the mission of making advanced AI broadly accessible. Their team includes researchers, engineers... 

    Xcede

    San Francisco, CA
    2 days ago
  •  ...DeepMind, OpenAI, Google Brain, Meta, Character.AI, Anthropic and beyond. Role Overview Reflection.AI is looking for a Member of Technical Staff - Infrastructure Security to secure our geographically diverse multi-cloud Kubernetes and cloud environments. In this role, you’ll... 
    Relocation package

    Reflection

    San Francisco, CA
    4 days ago
  •  ...candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making...  ...get gamed. Who We Are Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA... 
    Full time

    Touring Capital

    San Francisco, CA
    1 day ago
  • About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure...  ..., and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at... 
    Work at office

    Mixpeek

    San Francisco, CA
    1 day ago
  •  ...people take ownership, grow together, and share both the challenges and the wins. What You'll Do Build the supercomputing infrastructure that runs our agents. Our agents tackle long-horizon, high-performance workloads, and you'll design the cloud compute,... 
    Work at office
    Remote work
    Flexible hours

    Asari AI

    San Francisco, CA
    2 hours ago
  • $120k - $300k

     ...FinTech The Role A backend-leaning Member of Technical Staff building the end-to-end systems that...  ...workflows — high-performance infrastructure, reliability, and backend architecture...  ...Tech stack: Python, AWS (distributed inference, caching, queue orchestration, self-... 
    Full time
    H1b
    Visa sponsorship

    David Joseph & Company

    San Francisco, CA
    20 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical...  ...to focus on innovation, not on infrastructure. We aim to simplify the AI development...  ...transformation, training/fine-tuning, and inference? You will also: Find opportunities... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    3 days ago
  • Sail builds the world's most efficient software for inference ("processing LLM tokens") and agent hosting (cloud VMs). Together, our technologies...  ...CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear. Meet the CEO. This is... 
    Work at office
    Immediate start

    Theory Ventures

    San Francisco, CA
    1 day ago
  •  .... Successful candidates typically come from staff or principal-level roles and are recognized for establishing technical direction, leading large-scale initiatives,...  ...teams use to right‑size space and budgets. This infrastructure already powers 16,000 workplaces and 9,000+... 
    Work at office
    Local area
    Monday to Thursday

    Envoy Inc.

    San Francisco, CA
    1 day ago
  •  ...curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet... 

    Sieve

    San Francisco, CA
    3 days ago
  •  ...and parallelism strategies, and help us squeeze every FLOP out of our hardware. What you’ll do Modify and extend state-of-the-art inference engines like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain... 

    Sail Research

    San Francisco, CA
    3 days ago
  •  ...neocloud designed for fast, efficient inference. As AI workloads become more complex and...  ...workloads at massive scale and help define the infrastructure layer for the future of AI. This gives...  ...and run this company. As an early member of the team, you will have significant... 

    The Consensus

    San Francisco, CA
    1 day ago
  • $150k - $265k

     ...human again. Mission We're building the platform for the future of voice technology. Our market edge is extensible, reliable infrastructure designed for the full complexity of voice interactions. 18 months, 150k developers, adding 1000 every day. Give it a try here... 
    Full time
    Shift work

    Vapi

    San Francisco, CA
    1 day ago
  • About Mandolin Nearly every disease will become treatable in our lifetimes. Mandolin is laying the clinical and financial infrastructure to get groundbreaking treatments to patients faster, powered by AI agents. Mandolin partners closely with the largest healthcare institutions... 
    Local area

    Mandolin

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff — Inference Infrastructure. Be the first to apply!