Member of Technical Staff (AI Inference Engineer)
$220kPerplexity
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us. What you will work on Examples Of Real Work The Team Does New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic. Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving. Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. Who we're looking for Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus. You understand modern LLM architectures and are able to bring them up reliably in a production environment. You've built and operated production distributed systems under real load - ideally performance-critical ones. Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels. You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday. Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. Good if you touched any of ML compilers and framework internals: PyTorch internals, torch.compile, custom operators. Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism. Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving. Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis. Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. Qualifications 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow). Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores). Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation). Compensation Range: $220K - $485K #J-18808-Ljbffr Perplexity
- Member of Technical Staff - ML Systems & Inference Bay Area, CA | Onsite Join a well-funded AI infrastructure startup building the orchestration layer for next-generation AI workloads... ..., distributed systems, and performance engineering. You'll build production inference...Suggested
$150k - $300k
...cloud LLM serving, LLM inference optimization and RL systems... ...training stack. Core Technical Responsibilities LLM... ...PyTorch: LLM Inference engine development and integration... ...to shape decentralized AI and RL at Prime... ...development and encourage team members to contribute to the...SuggestedWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work- ...the Semiconductor and AI industries. Our in-... ..., distils our deep technical research and knowledge... .... Position Overview Member of Technical Staff will play a crucial... ...training & inference benchmarks & system... ...in Computer Science, Engineering or other relevant technical...SuggestedFull timeWork at officeRemote workWorldwide
- Job Description - Member of Technical Staff (Inference) Location: San Francisco (on-site at our offices) About Artificial Analysis... ...Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities...SuggestedShift work
- ...first heterogeneous neocloud for AI workloads. As AI systems scale,... ...datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design... .... This role is ideal for engineers who deeply understand how modern...Suggested
$150k - $250k
...servicing with the industry’s most advanced AI credit-servicing agents. We are backed... ...Product Hunt), Charlie Songhurst (Board Member, Meta), and Michael Jones (Former Chair,... ...the United Nations, UChicago, and Oxford engineers and researchers. Our omnichannel...Full timeInternshipWorldwide- Member of Technical Staff - Applied AI Engineer Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Valthos Inc. Valthos is an applied biological intelligence company. We build and deploy software and biological AI systems to safeguard humanity. Applied...Full timeWork at office
- BURNT Member of Technical Staff AI/ML Engineer (MLOps-Focused) Location On-site, San Francisco · Experience 5-7 years · Compensation $150,000-$275,000 + equity About Burnt Burnt isn't building software on top of ERPs. We don't believe ERPs will exist in the long run. They...Seasonal workLive in
- About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era... ..., so it's simple to serve low-latency inference, fine-tune models, and access production... ...international olympiad medalists, and experienced engineering and product leaders with decades of...Work at office
- ...mission is general causal intelligence; AI that is capable of (1) predicting the future... ..., and CERN. We look for infrastructure engineers who are excited to tackle unsolved... ...Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting...
$200k - $400k
About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible... ...directly impact how the world runs AI inference. Skills And Qualifications... ...LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM...Remote workVisa sponsorshipShift work- ...parallelism strategies, and help us squeeze every FLOP out of our hardware. What you’ll do Modify and extend state-of-the-art inference engines like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every...
- Overview We\'re looking for engineers and scientists to design, optimize, and scale the systems that power our diffusion LLMs in production. Your work will make inference faster, more cost-effective, and more reliable. Key Responsibilities Build and optimize high-performance...
$225k
About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed...RelocationVisa sponsorship- Description We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models. You'll work across the inference stack...Visa sponsorshipRelocation package
$256k - $276k
Overview Member of Technical Staff, AI Reliability & Monitoring Engineering Lead — Postman Join to apply for the Member of Technical Staff, AI Reliability & Monitoring Engineering Lead role at Postman. What You’ll Do Develop and manage reliability metrics (SLOs) for AI...Full timeWork at officeFlexible hours3 days per week$300 per month
...us Edison Scientific builds and deploys AI scientist agents to accelerate science and... ...an ambitious team run by scientists and engineers from leading institutions across biology,... ...Mathematics, Physics, Data Science, or a related technical field. Proficiency in Python and/or...Full timeWork at office- ...every one of them fans out into multiple AI inference requests running in real time. Behind... ...cloud providers. Today, our inference engineers and researchers build models while also... ...shifts without human intervention. Set technical direction across teams. Partner with inference...Shift work
- ...is hiring builders to join our Multimodal AI group, an industry-leading team defining... ...modalities we have yet to invent. As an engineer on the Multimodal AI team, you will work... ...products end‑to‑end, from problem definition to technical design, implementation, and launch. Hill...
- ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for... ...As a founding member of the engineering team, you will impact the design... ...is revolutionizing the AI development landscape with... ...training/fine-tuning, and inference? You will also: Find opportunities...Full timePart timeWork at officeWork from homeFlexible hours2 days per week
- ...building the best way to talk to AI and humans together — where AI... ...day, and everyone talks to users. Member of Technical Staff is the title we use for engineers who own hard problems end to end... ..., fine-tuning, evaluation, inference, or RAG at scale High-performance...
$240k - $280k
...0/yr Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how breakthrough... ...increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified...Full timeRemote workWorldwideRelocation$200k - $300k
Location San Francisco Employment Type Full time Department AI Compensation $200K - $300K • Offers Equity U.S. Benefits Full‑time... ...the amounts listed above. Perplexity is seeking an energetic engineer to join our highly driven Comet Agents engineering team. The...Full timeFlexible hours$250k
...career? Join a fast-growing AI compute platform building... ...the chance to join as a Member of Technical Staff at a pivotal stage in the company... ...intelligence, inference gateways, and agentic operations... ...without hiding it from the engineers who need to debug it You...Full time$200k - $350k
Member of Technical Staff — LLM Research & Training About the Role We are looking... ...to join an early-stage AI company building and training... ...motivated researchers and engineers who want to contribute directly... ...CUDA and Triton. Improve inference and training kernels when necessary...H1bVisa sponsorship$285k - $315k
...believe that the future of AI depends on the unglamorous:... ...re looking for researchers, engineers and organizations who agree... ...kernels. We're hiring a Member of Technical Staff for GPU Compiler Engineering... ...passes to optimize training and inference workloads You'll work in...Full timeWork at officeImmediate startRelocation package- ...platform for evaluating how AI models perform in the... ...ML researchers and engineers to design experiments,... ...posts) clearly to technical and non-technical partners... ...modeling, causal inference, and experimental design... ...markets where our team members are based. The base salary...Permanent employmentWork at officeShift work
- ...in the Semiconductor and AI industries. Our in-depth... ..., distils our deep technical research and knowledge into... ...for a highly motivated member of technical staff to join our engineering team to work on system modelling... ...frontier LLM training & inference models Implement modern...Full timeWork at officeRemote workWorldwide
- ...recognize parts of inputs that are unimportant, reducing inference costs for scale-ups and enterprises that integrate LLMs into... ...team is 5 people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems...Visa sponsorship
- ...re at a pivotal moment for AI and energy. Demand for compute... ...at . About the Role As a Member of Technical Staff, you will help invent and... ...skills with hands‑on software engineering experience and are excited... ..., distributed training/inference frameworks, or large‑scale...Work from homeFlexible hours2 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff (AI Inference Engineer). Be the first to apply!
- mri tech aide San Francisco, CA
- salesforce technical analyst San Francisco, CA
- service desk assistant San Francisco, CA
- end user support technician San Francisco, CA
- operations support technician San Francisco, CA
- help desk technical support San Francisco, CA
- technical assistant San Francisco, CA
- support analyst San Francisco, CA
- technical associate San Francisco, CA
- life support technician San Francisco, CA



