Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff (AI Inference Engineer)

$220k

Perplexity

We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us. What you will work on Examples Of Real Work The Team Does New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic. Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving. Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. Who we're looking for Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus. You understand modern LLM architectures and are able to bring them up reliably in a production environment. You've built and operated production distributed systems under real load - ideally performance-critical ones. Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels. You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday. Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. Good if you touched any of ML compilers and framework internals: PyTorch internals, torch.compile, custom operators. Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism. Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving. Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis. Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. Qualifications 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow). Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores). Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation). Compensation Range: $220K - $485K #J-18808-Ljbffr Perplexity

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff (AI Inference Engineer) in San Francisco, CA vacancy
  • $150k - $300k

     ...cloud LLM serving, LLM inference optimization and RL systems...  ...training stack. Core Technical Responsibilities LLM...  ...PyTorch: LLM Inference engine development and integration...  ...to shape decentralized AI and RL at Prime...  ...development and encourage team members to contribute to the... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    3 days ago
  •  ...first heterogeneous neocloud for AI workloads. As AI systems scale,...  ...datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design...  .... This role is ideal for engineers who deeply understand how modern... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    1 day ago
  •  ...Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand...  ...Opportunity Language model inference is the fastest-moving market...  ...reference point, and we’re hiring a Member of Technical Staff to drive them. You’ll own... 
    Suggested
    Shift work

    Artificial Analysis

    San Francisco, CA
    2 days ago
  • $150k - $250k

     ...servicing with the industry’s most advanced AI credit-servicing agents. We are backed...  ...Product Hunt), Charlie Songhurst (Board Member, Meta), and Michael Jones (Former Chair,...  ...the United Nations, UChicago, and Oxford engineers and researchers. Our omnichannel... 
    Suggested
    Full time
    Internship
    Worldwide

    Krew

    San Francisco, CA
    more than 2 months ago
  • $100k - $300k

    About Cogent Cogent is an Applied AI Lab building the next generation of AI agents for...  ...looking for talented, ambitious AI/ML Engineers who are excited to build in the Applied AI...  ...Onboard, support and uplevel future team members Mentor and grow future junior team members... 
    Suggested

    Cogent

    San Francisco, CA
    3 days ago
  • Member of Technical Staff - Applied AI Engineer Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Valthos Inc. Valthos is an applied biological intelligence company. We build and deploy software and biological AI systems to safeguard humanity. Applied... 
    Full time
    Work at office

    Valthos

    San Francisco, CA
    1 day ago
  • $350k

     ...Our first goal is to democratize frontier AI R&D across scientific disciplines. We...  ...experiments. Our team includes researchers and engineers from Anthropic, Google DeepMind, xAI,...  ...We are looking for an engineer to own the inference systems that power our models in production... 

    Mirendil

    San Francisco, CA
    3 days ago
  •  ...s most efficient software for inference ("processing LLM tokens") and...  ...allow our customers to deploy AI agents at large scale to do the...  ...parallelism strategies at the engine level, and help us use a heterogenous...  ...experience, and share as much technical detail about Sail as you want... 
    Work at office
    Immediate start

    Theory Ventures

    San Francisco, CA
    1 day ago
  •  ...evaluating it against the current best on real inference workloads, and keeping it only when it...  .... What We\'re Looking For Performance engineering on real inference or accelerator code....  .... Who We Are Infinity is an early-stage AI infrastructure research company building... 
    Full time

    Touring Capital

    San Francisco, CA
    1 day ago
  • $225k

    About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed... 
    Relocation
    Visa sponsorship

    Magic

    San Francisco, CA
    2 days ago
  • $200k - $400k

    About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible...  ...directly impact how the world runs AI inference. Skills And Qualifications...  ...LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM... 
    Remote work
    Visa sponsorship
    Shift work

    Inferact

    San Francisco, CA
    1 day ago
  •  ...mission is general causal intelligence; AI that is capable of (1) predicting the future...  ..., and CERN. We look for infrastructure engineers who are excited to tackle unsolved...  ...Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting... 

    Causal Labs

    San Francisco, CA
    3 days ago
  •  ...the world's most efficient software for inference (processing LLM tokens) and agent hosting...  ...technologies allow our customers to deploy AI agents at large scale to do the most...  ...scheduling and parallelism strategies at the engine level, and help us use a heterogenous mix... 

    Sail

    San Francisco, CA
    2 days ago
  • About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era...  ..., so it's simple to serve low-latency inference, fine-tune models, and access production...  ...international olympiad medalists, and experienced engineering and product leaders with decades of... 
    Work at office

    Mixpeek

    San Francisco, CA
    1 day ago
  • $300 per month

     ...us Edison Scientific builds and deploys AI scientist agents to accelerate science and...  ...an ambitious team run by scientists and engineers from leading institutions across biology,...  ...Mathematics, Physics, Data Science, or a related technical field. Proficiency in Python and/or... 
    Full time
    Work at office

    Edison Scientific Inc.

    San Francisco, CA
    16 hours ago
  • $256k - $276k

    Overview Member of Technical Staff, AI Reliability & Monitoring Engineering Lead — Postman Join to apply for the Member of Technical Staff, AI Reliability & Monitoring Engineering Lead role at Postman. What You’ll Do Develop and manage reliability metrics (SLOs) for AI... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    4 days ago
  •  ...every one of them fans out into multiple AI inference requests running in real time. Behind...  ...cloud providers. Today, our inference engineers and researchers build models while also...  ...shifts without human intervention. Set technical direction across teams. Partner with inference... 
    Shift work

    United States Digital Space LLC

    San Francisco, CA
    3 days ago
  • Artificial Analysis, the leading AI benchmarking company, is hiring a Member of Technical Staff to own serverless inference coverage and drive performance benchmarks across provider...  ...serving performance. You will work with engineers and leadership to direct the roadmap of... 

    Artificial Analysis

    San Francisco, CA
    2 days ago
  •  ...Perplexity is seeking energetic engineers to join our highly driven Agents engineering team. The Agents team consists of backend, full-stack, and AI/ML engineers who collaborate to build harnesses and AI systems powering delightful agentic experiences. These experiences... 
    Full time
    Flexible hours

    Perplexity

    San Francisco, CA
    2 days ago
  •  ...is hiring builders to join our Multimodal AI group, an industry-leading team defining...  ...modalities we have yet to invent. As an engineer on the Multimodal AI team, you will work...  ...products end‑to‑end, from problem definition to technical design, implementation, and launch. Hill... 

    Perplexity AI Inc.

    San Francisco, CA
    2 days ago
  • $250k

     ...career? Join a fast-growing AI compute platform building...  ...the chance to join as a Member of Technical Staff at a pivotal stage in the company...  ...intelligence, inference gateways, and agentic operations...  ...without hiding it from the engineers who need to debug it You... 
    Full time
    San Francisco, CA
    9 days ago
  • $120k - $300k

     ...About the Company Our client builds AI compliance analysts that work inside browsers...  ...The Role A backend-leaning Member of Technical Staff building the end-to-end systems that power...  ...Tech stack: Python, AWS (distributed inference, caching, queue orchestration, self-healing... 
    Full time
    H1b
    Visa sponsorship

    David Joseph & Company

    San Francisco, CA
    20 days ago
  •  ...Description Job Description Member of Technical Staff, Machine Learning, Artificial Intelligence (AI) Required, Work From Home...  ...Technical Staff role is for engineers who want to develop strong systems...  ..., training, evaluation, and inference. - Fine-tune and adapt... 
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    26 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for...  ...As a founding member of the engineering team, you will impact the design...  ...is revolutionizing the AI development landscape with...  ...training/fine-tuning, and inference? You will also: Find opportunities... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    3 days ago
  • $240k - $280k

     ...0/yr Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how breakthrough...  ...increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    2 days ago
  •  ...security-first enterprise AI company. We build cutting-edge...  ...is a team of researchers, engineers, designers, and more, who...  ...or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work...  ...distributed training or inference pipelines. Understanding of... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    San Francisco, CA
    4 days ago
  •  ...re at a pivotal moment for AI and energy. Demand for compute...  ...at . About the Role As a Member of Technical Staff, you will help invent and...  ...skills with hands‑on software engineering experience and are excited...  ..., distributed training/inference frameworks, or large‑scale... 
    Work from home
    Flexible hours
    2 days per week

    Emerald AI

    San Francisco, CA
    1 day ago
  • $150k - $300k

     ...Chief Scientist, Together AI), Dylan Patel (...  ...runs the jobs. Core Technical Responsibilities Hosted...  ...Kubernetes-based training and inference orchestration across...  ...We're looking for engineers who are fluent across...  ...development and encourage team members to contribute to the... 
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Kubelt

    San Francisco, CA
    4 days ago
  • $200k - $400k

     ...that. We have built the first AI simulation of society,...  ...Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile...  ...model training, and research engineering. Others may bring...  ...Bayesian modeling, causal inference, psychometrics, polling, or... 
    Flexible hours

    Simile

    San Francisco, CA
    7 hours ago
  •  ...us at the frontier of AI, acquisitions, and transformation...  ...-building, with each member having previously...  ...is our transformation engine. Each company we bring...  ...practical systems. Act as the technical lead for this research...  ...optimizing large-scale inference systems, including... 

    Enam, Inc.

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff (AI Inference Engineer). Be the first to apply!