Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff (AI Inference Engineer)

$220k

Perplexity

We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us. What you will work on Examples Of Real Work The Team Does New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic. Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving. Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. Who we're looking for Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus. You understand modern LLM architectures and are able to bring them up reliably in a production environment. You've built and operated production distributed systems under real load - ideally performance-critical ones. Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels. You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday. Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. Good if you touched any of ML compilers and framework internals: PyTorch internals, torch.compile, custom operators. Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism. Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving. Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis. Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. Qualifications 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow). Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores). Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation). Compensation Range: $220K - $485K #J-18808-Ljbffr Perplexity

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff (AI Inference Engineer) in San Francisco, CA vacancy
  • Member of Technical Staff - ML Systems & Inference Bay Area, CA | Onsite Join a well-funded AI infrastructure startup building the orchestration layer for next-generation AI workloads...  ..., distributed systems, and performance engineering. You'll build production inference... 
    Suggested

    Acceler8 Talent

    San Francisco, CA
    2 days ago
  • $150k - $300k

     ...cloud LLM serving, LLM inference optimization and RL systems...  ...training stack. Core Technical Responsibilities LLM...  ...PyTorch: LLM Inference engine development and integration...  ...to shape decentralized AI and RL at Prime...  ...development and encourage team members to contribute to the... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    3 days ago
  •  ...the Semiconductor and AI industries. Our in-...  ..., distils our deep technical research and knowledge...  .... Position Overview Member of Technical Staff will play a crucial...  ...training & inference benchmarks & system...  ...in Computer Science, Engineering or other relevant technical... 
    Suggested
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    22 hours ago
  • Job Description - Member of Technical Staff (Inference) Location: San Francisco (on-site at our offices) About Artificial Analysis...  ...Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities... 
    Suggested
    Shift work

    Artificial Analysis, Inc.

    San Francisco, CA
    2 days ago
  •  ...first heterogeneous neocloud for AI workloads. As AI systems scale,...  ...datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design...  .... This role is ideal for engineers who deeply understand how modern... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    1 day ago
  • $150k - $250k

     ...servicing with the industry’s most advanced AI credit-servicing agents. We are backed...  ...Product Hunt), Charlie Songhurst (Board Member, Meta), and Michael Jones (Former Chair,...  ...the United Nations, UChicago, and Oxford engineers and researchers. Our omnichannel... 
    Full time
    Internship
    Worldwide

    Krew

    San Francisco, CA
    9 days ago
  • Member of Technical Staff - Applied AI Engineer Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Valthos Inc. Valthos is an applied biological intelligence company. We build and deploy software and biological AI systems to safeguard humanity. Applied... 
    Full time
    Work at office

    Valthos

    San Francisco, CA
    1 day ago
  • BURNT Member of Technical Staff AI/ML Engineer (MLOps-Focused) Location On-site, San Francisco · Experience 5-7 years · Compensation $150,000-$275,000 + equity About Burnt Burnt isn't building software on top of ERPs. We don't believe ERPs will exist in the long run. They... 
    Seasonal work
    Live in

    Burnt Group

    San Francisco, CA
    4 days ago
  • About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era...  ..., so it's simple to serve low-latency inference, fine-tune models, and access production...  ...international olympiad medalists, and experienced engineering and product leaders with decades of... 
    Work at office

    Mixpeek

    San Francisco, CA
    1 day ago
  •  ...mission is general causal intelligence; AI that is capable of (1) predicting the future...  ..., and CERN. We look for infrastructure engineers who are excited to tackle unsolved...  ...Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting... 

    Causal Labs

    San Francisco, CA
    3 days ago
  • $200k - $400k

    About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible...  ...directly impact how the world runs AI inference. Skills And Qualifications...  ...LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM... 
    Remote work
    Visa sponsorship
    Shift work

    Inferact

    San Francisco, CA
    1 day ago
  •  ...parallelism strategies, and help us squeeze every FLOP out of our hardware. What you’ll do Modify and extend state-of-the-art inference engines like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every... 

    Sail Research

    San Francisco, CA
    3 days ago
  • Overview We\'re looking for engineers and scientists to design, optimize, and scale the systems that power our diffusion LLMs in production. Your work will make inference faster, more cost-effective, and more reliable. Key Responsibilities Build and optimize high-performance... 

    Inception

    San Francisco, CA
    3 days ago
  • $225k

    About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed... 
    Relocation
    Visa sponsorship

    Magic

    San Francisco, CA
    2 days ago
  • Description We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models. You'll work across the inference stack... 
    Visa sponsorship
    Relocation package

    Reactor

    San Francisco, CA
    12 hours ago
  • $256k - $276k

    Overview Member of Technical Staff, AI Reliability & Monitoring Engineering Lead — Postman Join to apply for the Member of Technical Staff, AI Reliability & Monitoring Engineering Lead role at Postman. What You’ll Do Develop and manage reliability metrics (SLOs) for AI... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    4 days ago
  • $300 per month

     ...us Edison Scientific builds and deploys AI scientist agents to accelerate science and...  ...an ambitious team run by scientists and engineers from leading institutions across biology,...  ...Mathematics, Physics, Data Science, or a related technical field. Proficiency in Python and/or... 
    Full time
    Work at office

    Edison Scientific Inc.

    San Francisco, CA
    22 hours ago
  •  ...every one of them fans out into multiple AI inference requests running in real time. Behind...  ...cloud providers. Today, our inference engineers and researchers build models while also...  ...shifts without human intervention. Set technical direction across teams. Partner with inference... 
    Shift work

    Neura Market

    San Francisco, CA
    2 days ago
  •  ...is hiring builders to join our Multimodal AI group, an industry-leading team defining...  ...modalities we have yet to invent. As an engineer on the Multimodal AI team, you will work...  ...products end‑to‑end, from problem definition to technical design, implementation, and launch. Hill... 

    Perplexity AI Inc.

    San Francisco, CA
    2 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for...  ...As a founding member of the engineering team, you will impact the design...  ...is revolutionizing the AI development landscape with...  ...training/fine-tuning, and inference? You will also: Find opportunities... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    3 days ago
  •  ...building the best way to talk to AI and humans together — where AI...  ...day, and everyone talks to users. Member of Technical Staff is the title we use for engineers who own hard problems end to end...  ..., fine-tuning, evaluation, inference, or RAG at scale High-performance... 

    Shapes

    San Francisco, CA
    3 days ago
  • $240k - $280k

     ...0/yr Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how breakthrough...  ...increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    2 days ago
  • $200k - $300k

    Location San Francisco Employment Type Full time Department AI Compensation $200K - $300K • Offers Equity U.S. Benefits Full‑time...  ...the amounts listed above. Perplexity is seeking an energetic engineer to join our highly driven Comet Agents engineering team. The... 
    Full time
    Flexible hours

    B Capital

    San Francisco, CA
    4 days ago
  • $250k

     ...career? Join a fast-growing AI compute platform building...  ...the chance to join as a Member of Technical Staff at a pivotal stage in the company...  ...intelligence, inference gateways, and agentic operations...  ...without hiding it from the engineers who need to debug it You... 
    Full time
    San Francisco, CA
    29 days ago
  • $200k - $350k

    Member of Technical Staff — LLM Research & Training About the Role We are looking...  ...to join an early-stage AI company building and training...  ...motivated researchers and engineers who want to contribute directly...  ...CUDA and Triton. Improve inference and training kernels when necessary... 
    H1b
    Visa sponsorship

    Pragmatike

    San Francisco, CA
    4 days ago
  • $285k - $315k

     ...believe that the future of AI depends on the unglamorous:...  ...re looking for researchers, engineers and organizations who agree...  ...kernels. We're hiring a Member of Technical Staff for GPU Compiler Engineering...  ...passes to optimize training and inference workloads You'll work in... 
    Full time
    Work at office
    Immediate start
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    3 days ago
  •  ...platform for evaluating how AI models perform in the...  ...ML researchers and engineers to design experiments,...  ...posts) clearly to technical and non-technical partners...  ...modeling, causal inference, and experimental design...  ...markets where our team members are based. The base salary... 
    Permanent employment
    Work at office
    Shift work

    Arena Intelligence, Inc.

    San Francisco, CA
    1 day ago
  •  ...in the Semiconductor and AI industries. Our in-depth...  ..., distils our deep technical research and knowledge into...  ...for a highly motivated member of technical staff to join our engineering team to work on system modelling...  ...frontier LLM training & inference models Implement modern... 
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    22 hours ago
  •  ...recognize parts of inputs that are unimportant, reducing inference costs for scale-ups and enterprises that integrate LLMs into...  ...team is 5 people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems... 
    Visa sponsorship

    The Token Company

    San Francisco, CA
    2 days ago
  •  ...re at a pivotal moment for AI and energy. Demand for compute...  ...at . About the Role As a Member of Technical Staff, you will help invent and...  ...skills with hands‑on software engineering experience and are excited...  ..., distributed training/inference frameworks, or large‑scale... 
    Work from home
    Flexible hours
    2 days per week

    Emerald AI

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff (AI Inference Engineer). Be the first to apply!