Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff - Inference

Sail Research

Optimize token processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new scheduling and parallelism strategies, and help us squeeze every FLOP out of our hardware. What you’ll do Modify and extend state-of-the-art inference engines like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an NSys profile. Design and implement exotic parallelism schemes to work with "interesting" hardware topologies. Write custom GPU kernels to excel in specific regimes, such as cascade attention. What we’re looking for Strong understanding of LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases. Interest in MLSys research—great ideas like speculative decoding and sparse attention come from research, that we need to follow closely. Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these! Benefits Meals are provided. Every employee receives a Studio Display. #J-18808-Ljbffr Sail Research

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff - Inference in San Francisco, CA vacancy
  • $150k - $300k

     ...position spanning cloud LLM serving, LLM inference optimization and RL systems. You will be...  ...into our RL training stack. Core Technical Responsibilities LLM Serving Multi‑tenant...  ...in open development and encourage team members to contribute to the broader AI community... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    4 days ago
  • Job Description - Member of Technical Staff (Inference) Location: San Francisco (on-site at our offices) About Artificial Analysis Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities... 
    Suggested
    Shift work

    Artificial Analysis, Inc.

    San Francisco, CA
    3 days ago
  •  ...production workloads built to scale to gigawatt‑class AI datacenters. Mission Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build inference systems that execute full models end‑to‑end under real... 
    Suggested

    Gimlet Labs, Inc.

    San Francisco, CA
    4 days ago
  • $350k

     ...engineers from Anthropic, Google DeepMind, xAI, OpenAI, Microsoft, Apple, and MIT. The Role We are looking for an engineer to own the inference systems that power our models in production and research. You'll work across the full inference stack, from serving infrastructure... 
    Suggested

    Mirendil

    San Francisco, CA
    4 days ago
  • Sail builds the world's most efficient software for inference ("processing LLM tokens") and agent hosting (cloud VMs). Together, our technologies...  ...CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear. Meet the CEO. This is... 
    Suggested
    Work at office
    Immediate start

    Theory Ventures

    San Francisco, CA
    2 days ago
  • $225k

    About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed... 
    Relocation
    Visa sponsorship

    Magic

    San Francisco, CA
    3 days ago
  • $200k - $400k

    About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving...  ...OpenRLHF, Unsloth, LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM inference. Logistics... 
    Remote work
    Visa sponsorship
    Shift work

    Inferact

    San Francisco, CA
    2 days ago
  •  ...physical observations, ensemble generation, and rollout evaluation across model scales. Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations Design and implement techniques... 

    Causal Labs

    San Francisco, CA
    4 days ago
  • What You\'ll Work On The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual... 
    Full time

    Touring Capital

    San Francisco, CA
    2 days ago
  •  .... They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led... 
    Work at office

    Mixpeek

    San Francisco, CA
    2 days ago
  •  ...week in our SF Mission district office. Your Role As a Member of Technical Staff, you will be responsible for building Eventual's core...  ...for distributed training experience. Familiarity with inference optimization (batching, GPU utilization, TensorRT/ONNX Runtime... 
    Work at office
    Immediate start
    Flexible hours
    Night shift

    Eventual

    San Francisco, CA
    10 hours ago
  • $120k - $300k

     ...(Seed) · Industry: FinTech The Role A backend-leaning Member of Technical Staff building the end-to-end systems that power the client's production...  ...automation. Tech stack: Python, AWS (distributed inference, caching, queue orchestration, self-healing pipelines);... 
    Full time
    H1b
    Visa sponsorship

    David Joseph & Company

    San Francisco, CA
    21 days ago
  •  ...to recognize parts of inputs that are redundant, reducing inference costs for scale-ups and enterprises that integrate LLMs into...  ...with a research and product focus. Role Overview As a Member of Technical Staff on our research team, you'll own the model training stack... 
    Visa sponsorship

    The Token Company

    San Francisco, CA
    2 days ago
  • $240k - $280k

     ...Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how...  ...Referrals increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    3 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering...  ...ingestion, transformation, training/fine-tuning, and inference? You will also: Find opportunities to go deep into a wide... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    4 days ago
  •  ...uses Shapes every single day, and everyone talks to users. Member of Technical Staff is the title we use for engineers who own hard problems...  ...have experience with LLM training, fine-tuning, evaluation, inference, or RAG at scale High-performance Python backends at scale... 

    Shapes

    San Francisco, CA
    4 days ago
  •  ...the first multi-silicon neocloud designed for fast, efficient inference. As AI workloads become more complex and new hardware architectures...  ...every part of how we build and run this company. As an early member of the team, you will have significant ownership over your work... 

    The Consensus

    San Francisco, CA
    2 days ago
  • $200k - $400k

     ...Adam D'Angelo, and Guillermo Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement...  ..., uncertainty quantification, Bayesian modeling, causal inference, psychometrics, polling, or measurement theory. Behavioral... 
    Flexible hours

    Simile

    San Francisco, CA
    1 day ago
  • $250k

    Eragon — Member of Technical Staff Type: Full-time | On-site | San Francisco, CA Compensation: $250,000-$450,000 + 0.75%-2% equity Hiring count...  ...engineering: Design scalable pipelines for training, inference, and data processing Performance optimization: Improve latency... 
    Full time
    H1b
    Work at office
    Local area
    Visa sponsorship

    davidjoseph-co

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...startup founders to help them hire for high-priority roles. We're partnering with a fast-growing AI inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems that make LLM inference fast, and own the customers running on them... 
    H1b
    Work at office

    Simplify

    San Francisco, CA
    2 days ago
  •  ...a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers...  ...learn) large-scale datasets and distributed training or inference pipelines. Understanding of LLM architectures, tuning... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Visa Hunt

    San Francisco, CA
    4 days ago
  •  ...Job Description Job Description Job Description Member of Technical Staff, Machine Learning, Artificial Intelligence (AI) Required,...  ...improve ML components across data, training, evaluation, and inference. - Fine-tune and adapt models as part of larger... 
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    27 days ago
  • $250k

     ...scale. This opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth. You'll help build...  ...AI infrastructure, working on orchestration intelligence, inference gateways, and agentic operations tooling that underpin a... 
    Full time
    San Francisco, CA
    10 days ago
  •  ...record in AI and company-building, with each member having previously steered AI startups to...  ...and practical systems. Act as the technical lead for this research direction. Collaborate...  ...plus: experience optimizing large-scale inference systems, including latency, throughput,... 

    Enam, Inc.

    San Francisco, CA
    3 days ago
  • $150k - $350k

    Sieve — Member of Technical Staff, Applied Research Type: Full-time | On-site | San Francisco, CA Compensation: $150,000-$350,000 + 0.05%-0.4%...  ...+ APIs via pre/post-processing, parallelism, pipelining, inference optimization, and occasional fine-tuning Work across computer... 
    Full time
    Work experience placement
    H1b
    Work at office
    Visa sponsorship

    davidjoseph-co

    San Francisco, CA
    4 days ago
  • # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco / Bay Area preferred. Remote exceptional for the right person...  ...make AI workloads cheaper and easier to own by turning inference behavior, traces, workload replay, GPU signals, and task-path... 
    Full time
    Remote work

    Touchdown Labs, Inc.

    San Francisco, CA
    3 days ago
  • $200k - $400k

     ...stakes decisions. Second, simulating a society means running inference over populations of agents, not single requests. A single...  ...research is even possible to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will build the platform our... 
    Live in
    Flexible hours

    Simile

    San Francisco, CA
    4 days ago
  • $200k

    Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe AGI that accelerates humanity’s progress...  ...pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal. About the role As an engineer... 
    Relocation
    Visa sponsorship

    Magic AI, Inc

    San Francisco, CA
    2 days ago
  • Careers / Member of Technical Staff (AI research) Member of Technical Staff (AI research) You will build and ship core parts of Fearn’s platform...  ...'ll play a crucial role in training models and designing inference pipelines, pushing the frontier of VLMs for the patent... 
    Full time
    Work at office

    Kindredventures

    San Francisco, CA
    1 day ago
  • Member of Technical Staff - Post‑Training Join to apply for the Member of Technical Staff - Post‑Training role at Reflection AI . Our Mission...  ...pipelines, reward models, reinforcement learning algorithms, and inference‑time scaling techniques. Collaborate across pre‑training... 
    Full time
    Relocation package

    Reflection AI

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff - Inference. Be the first to apply!