Member of Technical Staff (AI Inference Engineer)
$220kPerplexity
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us. What you will work on Examples Of Real Work The Team Does New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic. Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving. Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. Who we're looking for Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus. You understand modern LLM architectures and are able to bring them up reliably in a production environment. You've built and operated production distributed systems under real load - ideally performance-critical ones. Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels. You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday. Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. Good if you touched any of ML compilers and framework internals: PyTorch internals, torch.compile, custom operators. Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism. Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving. Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis. Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. Qualifications 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow). Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores). Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation). Compensation Range: $220K - $485K #J-18808-Ljbffr Perplexity
$125k - $200k
Founding AI Engineer / Member of Technical Staff YC - Startup New York City or San Francisco Bay Area $125,000.00 - 200,000.00 (US Dollar) Ability to travel will be critical. Please apply only if the location is suitable for you and you are willing to travel! Thank you!...SuggestedTemporary workWork at office$100k - $300k
About Ataraxis AI Ataraxis is a clinical AI research lab working... ...structure, where every team member is empowered to actively contribute... ...a multidisciplinary team of engineers and scientists. Co-mentor... ...learning, domain adaptation, causal inference, model interpretability and...SuggestedWorldwide- ...in the Semiconductor and AI industries. Our in-depth... ..., distils our deep technical research and knowledge into... ...for a highly motivated member of technical staff to join our engineering team to work on system modelling... ...frontier LLM training & inference models Implement modern...SuggestedFull timeWork at officeRemote workWorldwide
- A leading software development company in New York is seeking an entry-level Python engineer to join their team in the Brooklyn office. The role involves working on the AI inference pipeline that powers sophisticated OCR and computer vision products. Candidates should have...SuggestedFull timeWork at office
- ...Overview Sr. Lead AI Engineer (FM Hosting, LLM Inference). At Capital One, we are creating responsible and reliable AI systems that are changing... ...cross‑functional teams of engineers, research scientists, technical program managers, and product managers to deliver AI‑...SuggestedLocal area
$229.9k - $286.2k
AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years,... ...cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-...Full timePart timeLocal area$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking... ...cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-...Full timePart timeLocal area- Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight...
$180k - $250k
About Crosby AI Crosby is an AI-first legal platform reimagining corporate legal services from the ground up. We are a team of technologists... ...the next generation of fast-growing companies. The Team The Engineering team builds the core systems, infrastructure, and developer...- ...Opportunity OffDeal is the world's first AI-native investment bank for small... ...sale of their lives… Until now. Our engineers built software to automate 80%+ of... ...getting started. The Role As a Member of Technical Staff, you'll work directly with the CTO to...Work experience placementRelocation package
- ...As a Member of Technical Staff at Quadrillion Labs, you'll build the systems that power Qualia, our research agent. This is a mix of exciting... ...Build effective systems for agent orchestration . You'll engineer the core system that drives Qualia, improving its ability...Work at officeLocal area
- ...Member of Technical Staff Shared Context is building adaptive personal AI that understands the texture of real life: our relationships, routines, responsibilities,... ...and intention. We're looking for a senior engineer who is hands-on and cares deeply about craft....
- ...Member of Technical Staff Location: NYC (onsite only – not remote) Alliance is the leading accelerator for crypto & AI founders. Since 2020 we've backed 300+ startups (Rain, Pump, Synthetix... ...Staff to join our in-house engineering team. You'll report directly to Carter...Temporary workRelocation
- ...largest corpus of action-labeled gaming data in the world. Member of Technical Staff is the title everyone in our technical team holds. Each... ...help you find the most important problem across research, engineering, and infrastructure that aligns with our team's ambitious...
- ...Member Of Technical Staff, Machine Learning Drug discovery is a prediction problem. Scientists design... .... At Inductive Bio, we're using AI to build in silico models that more accurately... ...closely with chemists and software engineers to integrate models into our software...
- ...when expert operators and purpose-built AI work together – which is why we... ...building the software to fix it. As a Member of the Technical Staff at Finch, you'll own critical features... ...the product is evolving quickly and engineers have real ownership – expect the scope...Work at officeRemote workFlexible hours1 day per week
- ...Modal Growth Engineer Opportunity AI needs a new infrastructure layer. We're building it at Modal... ...so it's simple to serve low-latency inference, fine-tune models, and access production... ...for a Growth Engineer to own the technical foundation of Modal's marketing and developer...Work at office
- .... We're creating a new category of AI-native creative tooling: a collaborative... ...The Role We're looking for a Member of Technical Staff to help build the core product and infrastructure... ...that ties everything together. All engineers at Melius own features end-to-end ,...Work at office
- ...Stripe, DoorDash, and Ramp. About the Role Members of Technical Staff (MTS) are the senior engineers who build the platform that everything else at Beacon... ...of the world - compounding growth. How We Use AI in Our Hiring Process: To ensure transparency, we want...
$300 per month
...Delangue and many other operators/technical leaders. _"Basis is on the... ...." — Prashant Mital, Applied AI Lead, OpenAI_ The Work Being a Member of Technical Staff at Basis means you'll face... ...team expands. It's common to see engineers do core infra work one quarter...Work at officeShift work$180k - $250k
...Physical AI will decide the balance of power for the next... ...yourself and through the other technical staff you coordinate on-site and... ...in time. Run the process-engineering side of deployment — sequencing... ...and sensor data, causal inference or econometrics, optimization...Full timeWork at office$200k - $300k
...Why you should join us At Solstice, we're building AI software that helps life sciences teams turn complex scientific and brand... ...help customers direct, inspect, and use their output. Take on engineering problems with depth. Make long-running agents reliable, stream...H1bWork at officeRelocation package- ...the field and shape what comes next. Member of the Technical Staff, Molecular Generation Location Employment... ...reach. The hardest problems in both AI and biology are being solved here,... ...with 5+ years of hands‑on research and engineering experience in generative modeling...Full time
$180k - $280k
Join to apply for the Full Stack Engineer role at OffDeal . Base Pay Range $180,000 - $280... ...per year. OffDeal is the world’s first AI-native investment bank for small businesses... .... Opportunity to be a foundational team member at a well‑funded, high‑growth startup....Relocation package- ...enabling companies to build, train, and serve AI models tailored to their own data,... .... The Role As a Training Infrastructure Engineer, you'll design, build, and optimize the infrastructure... ...of AI infrastructure, from low-latency inference to scalable model serving. Build What’s...
- ...companies to build, train, and serve AI models tailored to their own... ...As a Training Infrastructure Engineer, you'll design, develop, and... ...machine learning training, inference, and data processing... ...backend infrastructure, lead technical design discussions, mentor engineers...
- Member of Technical Staff: Backend Monk is an AI-native accounts receivable (AR) platform for B2B companies helping businesses get paid fast. The AR stack... ...is ahead of us. The role: We're looking for a backend engineer to join us in person in Flatiron, NYC. You'll own the...Work at officeFlexible hours
- About Decagon Decagon is the leading conversational AI platform empowering every brand to deliver... ...how we work and grow as a team. About The Role As a Member of Technical Staff, you'll join one of four engineering teams building the systems behind Decagon's AI agents...InternshipWork at officeLocal area
$140k - $270k
...focus on delivering care. We’ve built an AI-powered platform designed by... ...rest will follow. The Team At Anterior, engineers share a strong "sense of product" and... ...in your application. About The Role Members of Technical Staff at Anterior own problems end-to-end —...ApprenticeshipFlexible hours- ...companies to build, train, and serve AI models tailored to their own... ...runs on one of the busiest inference platforms in the world —... ...new grads into teams across Engineering, and we match you to a team based... ..., Engineering, or a related technical field, completed within the...Summer workInternship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff (AI Inference Engineer). Be the first to apply!
- senior IT support technician New York, NY
- remote support technician New York, NY
- tech aide New York, NY
- senior technical associate New York, NY
- IT help desk technician New York, NY
- customer support analyst New York, NY
- it technical specialist New York, NY
- work from home technical support specialist New York, NY
- product support technician New York, NY
- mri tech aide New York, NY

