Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff (Inference)

Artificial Analysis

Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. We are the go-to authority for understanding AI, from AI labs and enterprises to media, investors, and policymakers. Our benchmarks don't just measure the cutting edge of AI, they are actively shaping the frontier. Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist. We are a team of 40+, on track to triple by year end, backed by Nat Friedman (Github, Meta), Daniel Gross (SSI), Andrew Ng (Google Brain, DeepLearning.ai, Amazon), Adam D'Angelo (Quora, Poe, OpenAI), Clem Delangue (Hugging Face) and other industry leaders. The Opportunity Language model inference is the fastest-moving market in AI: dozens of providers, constant price and performance shifts, and billions of dollars of deployment decisions riding on independent data. Our inference benchmarks are the industry’s reference point, and we’re hiring a Member of Technical Staff to drive them. You’ll own coverage of the serverless inference landscape: benchmarking endpoints across quality, speed and price, extending our methodology to new dimensions like cached pricing, endpoint accuracy and agentic performance, and working directly with the inference providers and neoclouds who ship on our numbers. What You’ll Do Benchmark the Inference Landscape: Own performance and price coverage of serverless API inference across the provider ecosystem, from frontier labs to specialist providers Extend Our Methodology: Drive new benchmarking dimensions including cached pricing, endpoint accuracy and agentic performance, keeping our measurement ahead of how the industry deploys Partner with Providers: Work directly with inference providers and neoclouds to benchmark their endpoints, resolve methodology questions and shape how the market measures serving performance Analyze the Market: Produce the analysis the industry uses to understand inference performance and economics, from throughput and time-to-first-token to price-performance frontiers Drive Product Direction: Shape the roadmap of our inference benchmarking platform together with our engineers and pillar lead Become AI-Native: Embrace an AI-native workflow, using cutting-edge AI tools to generate leverage in a fast-changing industry and maintain our competitive edge in AI benchmarking What We’re Looking For You should know the inference market from the inside. Backgrounds include: engineering, product, developer relations or technical GTM roles at inference providers and neoclouds (e.g. Together AI, Fireworks, Baseten, Cerebras, Novita, Parasail, DeepInfra, Modal, CoreWeave, Lambda, Nebius, Crusoe or similar), or teams serving models in production at scale. Required: 3+ years of professional experience, including at least 2 years at, or working closely with, inference providers, neoclouds or teams serving models in production Strong analytical and critical thinking skills Proficiency in Python and data analysis Hands-on familiarity with the model serving stack (e.g. vLLM, SGLang, TensorRT-LLM) and inference APIs across providers Fluency in inference performance metrics and economics: tokens per second, time to first token, throughput versus latency trade-offs, cost per token Genuine, demonstrable interest and knowledge of Frontier AI. We want people who have informed opinions about where AI is heading, not just people who use AI tools Why Artificial Analysis? Shape how AI gets built: The leading AI labs track our benchmarks and use them to guide their development priorities. Your work will directly influence the direction of AI. Become a world expert in AI: You will evaluate every major model, across every major capability, as they are released. Very few roles offer this breadth of exposure to frontier AI. Work with the most important players in AI: You’ll manage relationships with teams at the leading AI labs and major enterprises as a trusted, independent voice. Join at a defining moment: We’re 40+ people, on track to double by end of year, backed by some of the most connected investors in AI. The people who join now will shape the product, the team, and the strategy as we scale. Competitive compensation including equity Interested in this role? #J-18808-Ljbffr Artificial Analysis

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff (Inference) in San Francisco, CA vacancy
  • $150k - $300k

     ...position spanning cloud LLM serving, LLM inference optimization and RL systems. You will be...  ...into our RL training stack. Core Technical Responsibilities LLM Serving Multi‑tenant...  ...in open development and encourage team members to contribute to the broader AI community... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    3 days ago
  •  ...power real production workloads built to scale to gigawatt-class AI datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build the inference systems that execute full models end-to-end under... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    1 day ago
  • $350k

     ...engineers from Anthropic, Google DeepMind, xAI, OpenAI, Microsoft, Apple, and MIT. The Role We are looking for an engineer to own the inference systems that power our models in production and research. You'll work across the full inference stack, from serving infrastructure... 
    Suggested

    Mirendil

    San Francisco, CA
    3 days ago
  • Sail builds the world's most efficient software for inference ("processing LLM tokens") and agent hosting (cloud VMs). Together, our technologies...  ...CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear. Meet the CEO. This is... 
    Suggested
    Work at office
    Immediate start

    Theory Ventures

    San Francisco, CA
    1 day ago
  •  ...physical observations, ensemble generation, and rollout evaluation across model scales. Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations Design and implement techniques... 
    Suggested

    Causal Labs

    San Francisco, CA
    3 days ago
  • $200k - $400k

    About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving...  ...OpenRLHF, Unsloth, LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM inference. Logistics... 
    Remote work
    Visa sponsorship
    Shift work

    Inferact

    San Francisco, CA
    1 day ago
  • Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs). Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work. In this role, you'll own token processing... 

    Sail

    San Francisco, CA
    2 days ago
  • $225k

    About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed... 
    Relocation
    Visa sponsorship

    Magic

    San Francisco, CA
    2 days ago
  • What You\'ll Work On The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual... 
    Full time

    Touring Capital

    San Francisco, CA
    1 day ago
  •  .... They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led... 
    Work at office

    Mixpeek

    San Francisco, CA
    1 day ago
  • $120k - $300k

     ...(Seed) · Industry: FinTech The Role A backend-leaning Member of Technical Staff building the end-to-end systems that power the client's production...  ...automation. Tech stack: Python, AWS (distributed inference, caching, queue orchestration, self-healing pipelines);... 
    Full time
    H1b
    Visa sponsorship

    David Joseph & Company

    San Francisco, CA
    20 days ago
  • $240k - $280k

     ...Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how...  ...Referrals increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    2 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering...  ...ingestion, transformation, training/fine-tuning, and inference? You will also: Find opportunities to go deep into a wide... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    3 days ago
  •  ...the first multi-silicon neocloud designed for fast, efficient inference. As AI workloads become more complex and new hardware architectures...  ...every part of how we build and run this company. As an early member of the team, you will have significant ownership over your work... 

    The Consensus

    San Francisco, CA
    1 day ago
  •  ...recognize parts of inputs that are unimportant, reducing inference costs for scale-ups and enterprises that integrate LLMs into...  ...team is 5 people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems... 
    Visa sponsorship

    The Token Company

    San Francisco, CA
    2 days ago
  • $200k - $400k

     ...Adam D'Angelo, and Guillermo Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement...  ..., uncertainty quantification, Bayesian modeling, causal inference, psychometrics, polling, or measurement theory. Behavioral... 
    Flexible hours

    Simile

    San Francisco, CA
    6 hours ago
  • $250k

    Eragon — Member of Technical Staff Type: Full-time | On-site | San Francisco, CA Compensation: $250,000-$450,000 + 0.75%-2% equity Hiring count...  ...engineering: Design scalable pipelines for training, inference, and data processing Performance optimization: Improve latency... 
    Full time
    H1b
    Work at office
    Local area
    Visa sponsorship

    davidjoseph-co

    San Francisco, CA
    3 days ago
  • $200k - $300k

     ...startup founders to help them hire for high-priority roles. We're partnering with a fast-growing AI inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems that make LLM inference fast, and own the customers running on them... 
    H1b
    Work at office

    Simplify

    San Francisco, CA
    1 day ago
  • $250k

     ...scale. This opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth. You'll help build...  ...AI infrastructure, working on orchestration intelligence, inference gateways, and agentic operations tooling that underpin a... 
    Full time
    San Francisco, CA
    9 days ago
  •  ...Job Description Job Description Job Description Member of Technical Staff, Machine Learning, Artificial Intelligence (AI) Required,...  ...improve ML components across data, training, evaluation, and inference. - Fine-tune and adapt models as part of larger... 
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    26 days ago
  •  ...a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers...  ...with large‑scale datasets and distributed training or inference pipelines. Understanding of LLM architectures, tuning... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    San Francisco, CA
    4 days ago
  • Member of Technical Staff, Applied AI The opportunity We are looking for a Member of Technical Staff with deep expertise in generative modelling...  ...of generative model architectures, training dynamics and inference behaviour. You are a skilful ML developer. You write ML... 
    Flexible hours

    Latent Labs

    San Francisco, CA
    4 days ago
  •  ...design and the responsibility to defend. About the Role As a Member of Technical Staff focused on Protein Structure Modeling, you will develop...  ...Improve the efficiency and reliability of model training and inference on large‑scale compute systems. Collaborate with... 
    Local area

    Radical Numerics

    San Francisco, CA
    2 days ago
  • Careers / Member of Technical Staff (AI research) Member of Technical Staff (AI research) You will build and ship core parts of Fearn’s platform...  ...'ll play a crucial role in training models and designing inference pipelines, pushing the frontier of VLMs for the patent... 
    Full time
    Work at office

    Kindredventures

    San Francisco, CA
    15 hours ago
  • Member of Technical Staff - Post‑Training Join to apply for the Member of Technical Staff - Post‑Training role at Reflection AI . Our Mission...  ...pipelines, reward models, reinforcement learning algorithms, and inference‑time scaling techniques. Collaborate across pre‑training... 
    Full time
    Relocation package

    Reflection AI

    San Francisco, CA
    3 days ago
  • $200k

    Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe AGI that accelerates humanity’s progress...  ...pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal. About the role As an engineer... 
    Relocation
    Visa sponsorship

    Magic AI, Inc

    San Francisco, CA
    1 day ago
  • $150k - $350k

    Sieve — Member of Technical Staff, Applied Research Type: Full-time | On-site | San Francisco, CA Compensation: $150,000-$350,000 + 0.05%-0.4%...  ...+ APIs via pre/post-processing, parallelism, pipelining, inference optimization, and occasional fine-tuning Work across computer... 
    Full time
    Work experience placement
    H1b
    Work at office
    Visa sponsorship

    davidjoseph-co

    San Francisco, CA
    3 days ago
  • # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco / Bay Area preferred. Remote exceptional for the right person...  ...make AI workloads cheaper and easier to own by turning inference behavior, traces, workload replay, GPU signals, and task-path... 
    Full time
    Remote work

    Touchdown Labs, Inc.

    San Francisco, CA
    2 days ago
  • $200k - $400k

     ...stakes decisions. Second, simulating a society means running inference over populations of agents, not single requests. A single...  ...research is even possible to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will build the platform our... 
    Live in
    Flexible hours

    Simile

    San Francisco, CA
    3 days ago
  •  ...record in AI and company-building, with each member having previously steered AI startups to...  ...and practical systems. Act as the technical lead for this research direction. Collaborate...  ...plus: experience optimizing large-scale inference systems, including latency, throughput,... 

    Enam, Inc.

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff (Inference). Be the first to apply!