Member of Technical Staff (AI Inference Engineer)
$220kPerplexity
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us. What you will work on Examples Of Real Work The Team Does New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic. Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving. Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. Who we're looking for Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus. You understand modern LLM architectures and are able to bring them up reliably in a production environment. You've built and operated production distributed systems under real load - ideally performance-critical ones. Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels. You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday. Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. Good if you touched any of ML compilers and framework internals: PyTorch internals, torch.compile, custom operators. Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism. Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving. Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis. Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. Qualifications 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow). Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores). Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation). Compensation Range: $220K - $485K #J-18808-Ljbffr Perplexity
$150k - $300k
...cloud LLM serving, LLM inference optimization and RL systems... ...training stack. Core Technical Responsibilities LLM... ...PyTorch: LLM Inference engine development and integration... ...to shape decentralized AI and RL at Prime... ...development and encourage team members to contribute to the...SuggestedWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work- ...first heterogeneous neocloud for AI workloads. As AI systems scale,... ...datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design... .... This role is ideal for engineers who deeply understand how modern...Suggested
- ...Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand... ...Opportunity Language model inference is the fastest-moving market... ...reference point, and we’re hiring a Member of Technical Staff to drive them. You’ll own...SuggestedShift work
$150k - $250k
...servicing with the industry’s most advanced AI credit-servicing agents. We are backed... ...Product Hunt), Charlie Songhurst (Board Member, Meta), and Michael Jones (Former Chair,... ...the United Nations, UChicago, and Oxford engineers and researchers. Our omnichannel...SuggestedFull timeInternshipWorldwide$100k - $300k
About Cogent Cogent is an Applied AI Lab building the next generation of AI agents for... ...looking for talented, ambitious AI/ML Engineers who are excited to build in the Applied AI... ...Onboard, support and uplevel future team members Mentor and grow future junior team members...Suggested- Member of Technical Staff - Applied AI Engineer Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Valthos Inc. Valthos is an applied biological intelligence company. We build and deploy software and biological AI systems to safeguard humanity. Applied...Full timeWork at office
- ...s most efficient software for inference ("processing LLM tokens") and... ...allow our customers to deploy AI agents at large scale to do the... ...parallelism strategies at the engine level, and help us use a heterogenous... ...experience, and share as much technical detail about Sail as you want...Work at officeImmediate start
$350k
...Our first goal is to democratize frontier AI R&D across scientific disciplines. We... ...experiments. Our team includes researchers and engineers from Anthropic, Google DeepMind, xAI,... ...We are looking for an engineer to own the inference systems that power our models in production...- ...evaluating it against the current best on real inference workloads, and keeping it only when it... .... What We\'re Looking For Performance engineering on real inference or accelerator code.... .... Who We Are Infinity is an early-stage AI infrastructure research company building...Full time
$200k - $400k
About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible... ...directly impact how the world runs AI inference. Skills And Qualifications... ...LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM...Remote workVisa sponsorshipShift work- ...the world's most efficient software for inference (processing LLM tokens) and agent hosting... ...technologies allow our customers to deploy AI agents at large scale to do the most... ...scheduling and parallelism strategies at the engine level, and help us use a heterogenous mix...
- ...mission is general causal intelligence; AI that is capable of (1) predicting the future... ..., and CERN. We look for infrastructure engineers who are excited to tackle unsolved... ...Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting...
$225k
About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed...RelocationVisa sponsorship- About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era... ..., so it's simple to serve low-latency inference, fine-tune models, and access production... ...international olympiad medalists, and experienced engineering and product leaders with decades of...Work at office
$300 per month
...us Edison Scientific builds and deploys AI scientist agents to accelerate science and... ...an ambitious team run by scientists and engineers from leading institutions across biology,... ...Mathematics, Physics, Data Science, or a related technical field. Proficiency in Python and/or...Full timeWork at office$256k - $276k
Overview Member of Technical Staff, AI Reliability & Monitoring Engineering Lead — Postman Join to apply for the Member of Technical Staff, AI Reliability & Monitoring Engineering Lead role at Postman. What You’ll Do Develop and manage reliability metrics (SLOs) for AI...Full timeWork at officeFlexible hours3 days per week- ...every one of them fans out into multiple AI inference requests running in real time. Behind... ...cloud providers. Today, our inference engineers and researchers build models while also... ...shifts without human intervention. Set technical direction across teams. Partner with inference...Shift work
- Artificial Analysis, the leading AI benchmarking company, is hiring a Member of Technical Staff to own serverless inference coverage and drive performance benchmarks across provider... ...serving performance. You will work with engineers and leadership to direct the roadmap of...
- ...Perplexity is seeking energetic engineers to join our highly driven Agents engineering team. The Agents team consists of backend, full-stack, and AI/ML engineers who collaborate to build harnesses and AI systems powering delightful agentic experiences. These experiences...Full timeFlexible hours
$220k - $405k
...is hiring builders to join our Multimodal AI group, an industry-leading team defining... ...modalities we have yet to invent. As an engineer on the Multimodal AI team, you will work... ...products end-to-end, from problem definition to technical design, implementation, and launch. Hill...$250k
...career? Join a fast-growing AI compute platform building... ...the chance to join as a Member of Technical Staff at a pivotal stage in the company... ...intelligence, inference gateways, and agentic operations... ...without hiding it from the engineers who need to debug it You...Full time- ...Description Job Description Member of Technical Staff, Machine Learning, Artificial Intelligence (AI) Required, Work From Home... ...Technical Staff role is for engineers who want to develop strong systems... ..., training, evaluation, and inference. - Fine-tune and adapt...Remote workWork from home
$120k - $300k
...About the Company Our client builds AI compliance analysts that work inside browsers... ...The Role A backend-leaning Member of Technical Staff building the end-to-end systems that power... ...Tech stack: Python, AWS (distributed inference, caching, queue orchestration, self-healing...Full timeH1bVisa sponsorship- ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for... ...As a founding member of the engineering team, you will impact the design... ...is revolutionizing the AI development landscape with... ...training/fine-tuning, and inference? You will also: Find opportunities...Full timePart timeWork at officeWork from homeFlexible hours2 days per week
$240k - $280k
...0/yr Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how breakthrough... ...increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified...Full timeRemote workWorldwideRelocation- ...security-first enterprise AI company. We build cutting-edge... ...is a team of researchers, engineers, designers, and more, who... ...or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work... ...distributed training or inference pipelines. Understanding of...Full timeWork at officeLocal areaRemote workHome office
- ...re at a pivotal moment for AI and energy. Demand for compute... ...at . About the Role As a Member of Technical Staff, you will help invent and... ...skills with hands‑on software engineering experience and are excited... ..., distributed training/inference frameworks, or large‑scale...Work from homeFlexible hours2 days per week
$150k - $300k
...Chief Scientist, Together AI), Dylan Patel (... ...runs the jobs. Core Technical Responsibilities Hosted... ...Kubernetes-based training and inference orchestration across... ...We're looking for engineers who are fluent across... ...development and encourage team members to contribute to the...Work at officeLocal areaRemote workVisa sponsorshipRelocation packageFlexible hours$200k - $400k
...that. We have built the first AI simulation of society,... ...Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile... ...model training, and research engineering. Others may bring... ...Bayesian modeling, causal inference, psychometrics, polling, or...Flexible hours$200k
Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission... ..., ultra-long context, and inference-time compute to achieve... ...goal. About the role As an engineer on the Supercomputing Platform... ...to schedule and manage AI workloads Develop modular,...RelocationVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff (AI Inference Engineer). Be the first to apply!
- work from home technical support specialist San Francisco, CA
- product support technician San Francisco, CA
- helpdesk support technician San Francisco, CA
- help desk assistant San Francisco, CA
- senior technical associate San Francisco, CA
- IT help desk technician San Francisco, CA
- technical solutions specialist San Francisco, CA
- desktop support analyst San Francisco, CA
- trade support analyst San Francisco, CA
- senior IT support technician San Francisco, CA



