Member of Technical Staff - Inference
Sail Research
Optimize token processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new scheduling and parallelism strategies, and help us squeeze every FLOP out of our hardware. What you’ll do Modify and extend state-of-the-art inference engines like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an NSys profile. Design and implement exotic parallelism schemes to work with "interesting" hardware topologies. Write custom GPU kernels to excel in specific regimes, such as cascade attention. What we’re looking for Strong understanding of LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases. Interest in MLSys research—great ideas like speculative decoding and sparse attention come from research, that we need to follow closely. Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these! Benefits Meals are provided. Every employee receives a Studio Display. #J-18808-Ljbffr Sail Research
- Job Description - Member of Technical Staff (Inference) Location: San Francisco (on-site at our offices) About Artificial Analysis Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities...SuggestedShift work
$200k - $260k
Member of Technical Staff, Inference Engine You will build the core of our inference engine: the runtime that takes a set of weights and serves them as a low-latency, high-throughput endpoint. This is the layer where scheduling, batching, memory, and the model meet. Your...Suggested$150k - $300k
...position spanning cloud LLM serving, LLM inference optimization and RL systems. You will be... ...into our RL training stack. Core Technical Responsibilities LLM Serving Multi‑tenant... ...in open development and encourage team members to contribute to the broader AI community...SuggestedWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work- ...towards financial investors, distils our deep technical research and knowledge into key insights on... ...and government agencies. Position Overview Member of Technical Staff will play a crucial role in developing training & inference benchmarks & system modelling. You will...SuggestedFull timeWork at officeRemote workWorldwide
- ...power real production workloads built to scale to gigawatt-class AI datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build the inference systems that execute full models end-to-end under...Suggested
- Member of Technical Staff - ML Systems & Inference Bay Area, CA | Onsite Join a well-funded AI infrastructure startup building the orchestration layer for next-generation AI workloads This role sits at the intersection of ML systems, inference, distributed systems, and...
$170k - $265k
...Perplexity is looking for a technical program manager to be the connective tissue between our model providers, engineering, and product teams, driving our core inference platform forward. Perplexity runs one of the highest-throughput inference stacks in the industry,...Full timeShift work$225k
About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed...RelocationVisa sponsorship- Description We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models. You'll work across the inference stack...Visa sponsorshipRelocation package
- ...physical observations, ensemble generation, and rollout evaluation across model scales. Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations Design and implement techniques...
$200k - $400k
About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving... ...OpenRLHF, Unsloth, LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM inference. Logistics...Remote workVisa sponsorshipShift work- .... They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led...Work at office
- ...engineers and scientists to design, optimize, and scale the systems that power our diffusion LLMs in production. Your work will make inference faster, more cost-effective, and more reliable. Key Responsibilities Build and optimize high-performance model serving systems for...
- ...‑shaping articles are: InferenceMAX: The world first open inference benchmark that continuous benchmarks performance of popular... ...Overview We are seeking a highly motivated & skilled Member of Technical Staff to join our growing engineering team. Member of Technical...Full timeWork at officeRemote workWorldwide
- ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering... ...ingestion, transformation, training/fine-tuning, and inference? You will also: Find opportunities to go deep into a wide...Full timePart timeWork at officeWork from homeFlexible hours2 days per week
- ...volume data platform for AI applications. We are looking for team members who love building enabling systems that empower our engineers... ...for production LLM applications, including model integrations, inference workloads, evaluation pipelines, observability, and the...Work at office
$150k - $350k
...Member of Technical Staff - Distributed Systems San Francisco, CA- 5 days per week onsite $150,000–$350,000 + equity The Opportunity Join a rapidly... ...a multi-silicon cloud platform for fast, efficient inference. The future of AI inference will not run exclusively on GPUs...$150k - $350k
...Member of Technical Staff | Distributed Systems San Francisco - Onsite $150k-$350k base + equity I'm working with a small, deeply technical $... ...million Series A AI infrastructure company , building an inference cloud for agentic workloads that can partition and orchestrate...- ...financial investors, distils our deep technical research and knowledge into key... ...We are looking for a highly motivated member of technical staff to join our engineering team to work... ...across both frontier LLM training & inference models Implement modern parallelism &...Full timeWork at officeRemote workWorldwide
$240k - $280k
...Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how... ...Referrals increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified...Full timeRemote workWorldwideRelocation- ...to recognize parts of inputs that are redundant, reducing inference costs for scale-ups and enterprises that integrate LLMs into... ...with a research and product focus. Role Overview As a Member of Technical Staff on our research team, you'll own the model training stack...Visa sponsorship
- ...Job Title Member of Technical Staff: Infrastructure Salary Not Disclosed Company Description Observable Intuition is an early‑stage AI infrastructure... ...first infrastructure hire, you will build a production inference platform from the ground up. You’ll design portable, multi...
- ...uses Shapes every single day, and everyone talks to users. Member of Technical Staff is the title we use for engineers who own hard problems... ...have experience with LLM training, fine-tuning, evaluation, inference, or RAG at scale High-performance Python backends at scale...
- ...design and the responsibility to defend. About the role As a Member of Technical Staff, ML Product Engineer, this role owns the layer between the... ...APIs, batch and compute systems, and services that make inference fast, reliable, and cheap at genome scale. You would build...Local area
$285k - $315k
...benchmark across hundreds of production kernels. We're hiring a Member of Technical Staff for GPU Kernel Engineering to establish and push the... ...kernels that execute during pre-training, post-training and inference, across NVIDIA, AMD, TPU and Trainium. You'll then take...Full timeWork at officeImmediate startRelocation package$200k - $400k
...Adam D'Angelo, and Guillermo Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement... ..., uncertainty quantification, Bayesian modeling, causal inference, psychometrics, polling, or measurement theory. Behavioral...Flexible hours- ...a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers... ...learn) large-scale datasets and distributed training or inference pipelines. Understanding of LLM architectures, tuning...Full timeWork at officeLocal areaRemote workHome office
$200k - $300k
...startup founders to help them hire for high-priority roles. We're partnering with a fast-growing AI inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems that make LLM inference fast, and own the customers running on them...H1bWork at office- ...the first multi-silicon neocloud designed for fast, efficient inference. As AI workloads become more complex and new hardware architectures... ...every part of how we build and run this company. As an early member of the team, you will have significant ownership over your work...
- Member of Technical Staff (Infra) We're looking for an experienced Backend Engineer to join Fearn as we build up a scalable infrastructure. Role... ...in architecting and building robust infrastructure and inference systems that power our platform, while mentoring junior engineers...Full timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff - Inference. Be the first to apply!
- mri tech aide San Francisco, CA
- salesforce technical analyst San Francisco, CA
- service desk assistant San Francisco, CA
- end user support technician San Francisco, CA
- operations support technician San Francisco, CA
- help desk technical support San Francisco, CA
- technical assistant San Francisco, CA
- support analyst San Francisco, CA
- technical associate San Francisco, CA
- life support technician San Francisco, CA

