Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff - Inference

Sail Research

Optimize token processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new scheduling and parallelism strategies, and help us squeeze every FLOP out of our hardware. What you’ll do Modify and extend state-of-the-art inference engines like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an NSys profile. Design and implement exotic parallelism schemes to work with "interesting" hardware topologies. Write custom GPU kernels to excel in specific regimes, such as cascade attention. What we’re looking for Strong understanding of LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases. Interest in MLSys research—great ideas like speculative decoding and sparse attention come from research, that we need to follow closely. Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these! Benefits Meals are provided. Every employee receives a Studio Display. #J-18808-Ljbffr Sail Research

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff - Inference in San Francisco, CA vacancy
  • Job Description - Member of Technical Staff (Inference) Location: San Francisco (on-site at our offices) About Artificial Analysis Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities... 
    Suggested
    Shift work

    Artificial Analysis, Inc.

    San Francisco, CA
    1 day ago
  • $200k - $260k

    Member of Technical Staff, Inference Engine You will build the core of our inference engine: the runtime that takes a set of weights and serves them as a low-latency, high-throughput endpoint. This is the layer where scheduling, batching, memory, and the model meet. Your... 
    Suggested

    ATBF Labs Inc.

    San Francisco, CA
    4 days ago
  • $150k - $300k

     ...position spanning cloud LLM serving, LLM inference optimization and RL systems. You will be...  ...into our RL training stack. Core Technical Responsibilities LLM Serving Multi‑tenant...  ...in open development and encourage team members to contribute to the broader AI community... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    2 days ago
  •  ...towards financial investors, distils our deep technical research and knowledge into key insights on...  ...and government agencies. Position Overview Member of Technical Staff will play a crucial role in developing training & inference benchmarks & system modelling. You will... 
    Suggested
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    4 days ago
  •  ...power real production workloads built to scale to gigawatt-class AI datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build the inference systems that execute full models end-to-end under... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    5 days ago
  • Member of Technical Staff - ML Systems & Inference Bay Area, CA | Onsite Join a well-funded AI infrastructure startup building the orchestration layer for next-generation AI workloads This role sits at the intersection of ML systems, inference, distributed systems, and... 

    Acceler8 Talent

    San Francisco, CA
    1 day ago
  • $170k - $265k

     ...Perplexity is looking for a technical program manager to be the connective tissue between our model providers, engineering, and product teams, driving our core inference platform forward. Perplexity runs one of the highest-throughput inference stacks in the industry,... 
    Full time
    Shift work

    Perplexity

    San Francisco, CA
    6 days ago
  • $225k

    About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed... 
    Relocation
    Visa sponsorship

    Magic

    San Francisco, CA
    1 day ago
  • Description We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models. You'll work across the inference stack... 
    Visa sponsorship
    Relocation package

    Reactor

    San Francisco, CA
    3 days ago
  •  ...physical observations, ensemble generation, and rollout evaluation across model scales. Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations Design and implement techniques... 

    Causal Labs

    San Francisco, CA
    2 days ago
  • $200k - $400k

    About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving...  ...OpenRLHF, Unsloth, LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM inference. Logistics... 
    Remote work
    Visa sponsorship
    Shift work

    Inferact

    San Francisco, CA
    5 days ago
  •  .... They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led... 
    Work at office

    Mixpeek

    San Francisco, CA
    3 days ago
  •  ...engineers and scientists to design, optimize, and scale the systems that power our diffusion LLMs in production. Your work will make inference faster, more cost-effective, and more reliable. Key Responsibilities Build and optimize high-performance model serving systems for... 

    Inception

    San Francisco, CA
    2 days ago
  •  ...‑shaping articles are: InferenceMAX: The world first open inference benchmark that continuous benchmarks performance of popular...  ...Overview We are seeking a highly motivated & skilled Member of Technical Staff to join our growing engineering team. Member of Technical... 
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    4 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering...  ...ingestion, transformation, training/fine-tuning, and inference? You will also: Find opportunities to go deep into a wide... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    2 days ago
  •  ...volume data platform for AI applications. We are looking for team members who love building enabling systems that empower our engineers...  ...for production LLM applications, including model integrations, inference workloads, evaluation pipelines, observability, and the... 
    Work at office

    LlamaIndex

    San Francisco, CA
    6 hours ago
  • $150k - $350k

     ...Member of Technical Staff - Distributed Systems San Francisco, CA- 5 days per week onsite $150,000–$350,000 + equity The Opportunity Join a rapidly...  ...a multi-silicon cloud platform for fast, efficient inference. The future of AI inference will not run exclusively on GPUs... 

    Acceler8 Talent

    San Francisco, CA
    5 days ago
  • $150k - $350k

     ...Member of Technical Staff | Distributed Systems San Francisco - Onsite $150k-$350k base + equity I'm working with a small, deeply technical $...  ...million Series A AI infrastructure company , building an inference cloud for agentic workloads that can partition and orchestrate... 

    Acceler8 Talent

    San Francisco, CA
    5 days ago
  •  ...financial investors, distils our deep technical research and knowledge into key...  ...We are looking for a highly motivated member of technical staff to join our engineering team to work...  ...across both frontier LLM training & inference models Implement modern parallelism &... 
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    4 days ago
  • $240k - $280k

     ...Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how...  ...Referrals increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    1 day ago
  •  ...to recognize parts of inputs that are redundant, reducing inference costs for scale-ups and enterprises that integrate LLMs into...  ...with a research and product focus. Role Overview As a Member of Technical Staff on our research team, you'll own the model training stack... 
    Visa sponsorship

    The Token Company

    San Francisco, CA
    5 days ago
  •  ...Job Title Member of Technical Staff: Infrastructure Salary Not Disclosed Company Description Observable Intuition is an early‑stage AI infrastructure...  ...first infrastructure hire, you will build a production inference platform from the ground up. You’ll design portable, multi... 

    Jack & Jill

    San Francisco, CA
    4 days ago
  •  ...uses Shapes every single day, and everyone talks to users. Member of Technical Staff is the title we use for engineers who own hard problems...  ...have experience with LLM training, fine-tuning, evaluation, inference, or RAG at scale High-performance Python backends at scale... 

    Shapes

    San Francisco, CA
    2 days ago
  •  ...design and the responsibility to defend. About the role As a Member of Technical Staff, ML Product Engineer, this role owns the layer between the...  ...APIs, batch and compute systems, and services that make inference fast, reliable, and cheap at genome scale. You would build... 
    Local area

    Radical Numerics

    San Francisco, CA
    6 hours ago
  • $285k - $315k

     ...benchmark across hundreds of production kernels. We're hiring a Member of Technical Staff for GPU Kernel Engineering to establish and push the...  ...kernels that execute during pre-training, post-training and inference, across NVIDIA, AMD, TPU and Trainium. You'll then take... 
    Full time
    Work at office
    Immediate start
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    2 days ago
  • $200k - $400k

     ...Adam D'Angelo, and Guillermo Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement...  ..., uncertainty quantification, Bayesian modeling, causal inference, psychometrics, polling, or measurement theory. Behavioral... 
    Flexible hours

    Simile

    San Francisco, CA
    3 days ago
  •  ...a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers...  ...learn) large-scale datasets and distributed training or inference pipelines. Understanding of LLM architectures, tuning... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Visa Hunt

    San Francisco, CA
    2 days ago
  • $200k - $300k

     ...startup founders to help them hire for high-priority roles. We're partnering with a fast-growing AI inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems that make LLM inference fast, and own the customers running on them... 
    H1b
    Work at office

    Simplify

    San Francisco, CA
    5 days ago
  •  ...the first multi-silicon neocloud designed for fast, efficient inference. As AI workloads become more complex and new hardware architectures...  ...every part of how we build and run this company. As an early member of the team, you will have significant ownership over your work... 

    The Consensus

    San Francisco, CA
    5 days ago
  • Member of Technical Staff (Infra) We're looking for an experienced Backend Engineer to join Fearn as we build up a scalable infrastructure. Role...  ...in architecting and building robust infrastructure and inference systems that power our platform, while mentoring junior engineers... 
    Full time
    Work at office

    Kindredventures

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff - Inference. Be the first to apply!