Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff (AI Inference Engineer)

$220k

Perplexity

We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us. What you will work on Examples Of Real Work The Team Does New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic. Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving. Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. Who we're looking for Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus. You understand modern LLM architectures and are able to bring them up reliably in a production environment. You've built and operated production distributed systems under real load - ideally performance-critical ones. Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels. You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday. Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. Good if you touched any of ML compilers and framework internals: PyTorch internals, torch.compile, custom operators. Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism. Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving. Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis. Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. Qualifications 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow). Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores). Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation). Compensation Range: $220K - $485K #J-18808-Ljbffr Perplexity

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff (AI Inference Engineer) in New York, NY vacancy
  • $125k - $200k

    Founding AI Engineer / Member of Technical Staff YC - Startup New York City or San Francisco Bay Area $125,000.00 - 200,000.00 (US Dollar) Ability to travel will be critical. Please apply only if the location is suitable for you and you are willing to travel! Thank you!... 
    Suggested
    Temporary work
    Work at office

    Butterfly Recruitment

    New York, NY
    19 hours ago
  • $100k - $300k

    About Ataraxis AI Ataraxis is a clinical AI research lab working...  ...structure, where every team member is empowered to actively contribute...  ...a multidisciplinary team of engineers and scientists. Co-mentor...  ...learning, domain adaptation, causal inference, model interpretability and... 
    Suggested
    Worldwide

    Ataraxis AI

    New York, NY
    3 days ago
  •  ...in the Semiconductor and AI industries. Our in-depth...  ..., distils our deep technical research and knowledge into...  ...for a highly motivated member of technical staff to join our engineering team to work on system modelling...  ...frontier LLM training & inference models Implement modern... 
    Suggested
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    New York, NY
    3 days ago
  • A leading software development company in New York is seeking an entry-level Python engineer to join their team in the Brooklyn office. The role involves working on the AI inference pipeline that powers sophisticated OCR and computer vision products. Candidates should have... 
    Suggested
    Full time
    Work at office

    Mathpix

    New York, NY
    4 days ago
  •  ...Overview Sr. Lead AI Engineer (FM Hosting, LLM Inference). At Capital One, we are creating responsible and reliable AI systems that are changing...  ...cross‑functional teams of engineers, research scientists, technical program managers, and product managers to deliver AI‑... 
    Suggested
    Local area

    Capital One

    New York, NY
    2 days ago
  • $229.9k - $286.2k

    AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years,...  ...cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    3 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    1 day ago
  • Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight... 

    Perplexity

    New York, NY
    4 days ago
  • $180k - $250k

    About Crosby AI Crosby is an AI-first legal platform reimagining corporate legal services from the ground up. We are a team of technologists...  ...the next generation of fast-growing companies. The Team The Engineering team builds the core systems, infrastructure, and developer... 

    SwiftCruit

    New York, NY
    3 days ago
  •  ...Opportunity OffDeal is the world's first AI-native investment bank for small...  ...sale of their lives… Until now. Our engineers built software to automate 80%+ of...  ...getting started. The Role As a Member of Technical Staff, you'll work directly with the CTO to... 
    Work experience placement
    Relocation package

    The Ask

    New York, NY
    19 hours ago
  •  ...As a Member of Technical Staff at Quadrillion Labs, you'll build the systems that power Qualia, our research agent. This is a mix of exciting...  ...Build effective systems for agent orchestration . You'll engineer the core system that drives Qualia, improving its ability... 
    Work at office
    Local area

    Quadrillion Labs, Inc

    New York, NY
    3 days ago
  •  ...Member of Technical Staff Shared Context is building adaptive personal AI that understands the texture of real life: our relationships, routines, responsibilities,...  ...and intention. We're looking for a senior engineer who is hands-on and cares deeply about craft.... 

    Shared Context Lab

    New York, NY
    1 day ago
  •  ...Member of Technical Staff Location: NYC (onsite only – not remote) Alliance is the leading accelerator for crypto & AI founders. Since 2020 we've backed 300+ startups (Rain, Pump, Synthetix...  ...Staff to join our in-house engineering team. You'll report directly to Carter... 
    Temporary work
    Relocation

    Alliance

    New York, NY
    2 days ago
  •  ...largest corpus of action-labeled gaming data in the world. Member of Technical Staff is the title everyone in our technical team holds. Each...  ...help you find the most important problem across research, engineering, and infrastructure that aligns with our team's ambitious... 

    Medal

    New York, NY
    3 days ago
  •  ...Member Of Technical Staff, Machine Learning Drug discovery is a prediction problem. Scientists design...  .... At Inductive Bio, we're using AI to build in silico models that more accurately...  ...closely with chemists and software engineers to integrate models into our software... 

    Inductive Bio

    New York, NY
    19 hours ago
  •  ...when expert operators and purpose-built AI work together – which is why we...  ...building the software to fix it. As a Member of the Technical Staff at Finch, you'll own critical features...  ...the product is evolving quickly and engineers have real ownership – expect the scope... 
    Work at office
    Remote work
    Flexible hours
    1 day per week

    Finch Services

    New York, NY
    19 hours ago
  •  ...Modal Growth Engineer Opportunity AI needs a new infrastructure layer. We're building it at Modal...  ...so it's simple to serve low-latency inference, fine-tune models, and access production...  ...for a Growth Engineer to own the technical foundation of Modal's marketing and developer... 
    Work at office

    Modal

    New York, NY
    19 hours ago
  •  .... We're creating a new category of AI-native creative tooling: a collaborative...  ...The Role We're looking for a Member of Technical Staff to help build the core product and infrastructure...  ...that ties everything together. All engineers at Melius own features end-to-end ,... 
    Work at office

    Melius AI, Inc

    New York, NY
    3 days ago
  •  ...Stripe, DoorDash, and Ramp. About the Role Members of Technical Staff (MTS) are the senior engineers who build the platform that everything else at Beacon...  ...of the world - compounding growth. How We Use AI in Our Hiring Process: To ensure transparency, we want... 

    BEACON SOFTWARE COMPANY

    New York, NY
    4 days ago
  • $300 per month

     ...Delangue and many other operators/technical leaders. _"Basis is on the...  ...." — Prashant Mital, Applied AI Lead, OpenAI_ The Work Being a Member of Technical Staff at Basis means you'll face...  ...team expands. It's common to see engineers do core infra work one quarter... 
    Work at office
    Shift work

    Basis

    New York, NY
    1 day ago
  • $180k - $250k

     ...Physical AI will decide the balance of power for the next...  ...yourself and through the other technical staff you coordinate on-site and...  ...in time. Run the process-engineering side of deployment — sequencing...  ...and sensor data, causal inference or econometrics, optimization... 
    Full time
    Work at office

    AIC

    New York, NY
    3 days ago
  • $200k - $300k

     ...Why you should join us At Solstice, we're building AI software that helps life sciences teams turn complex scientific and brand...  ...help customers direct, inspect, and use their output. Take on engineering problems with depth. Make long-running agents reliable, stream... 
    H1b
    Work at office
    Relocation package

    Solstice Corp

    New York, NY
    2 days ago
  •  ...the field and shape what comes next. Member of the Technical Staff, Molecular Generation Location Employment...  ...reach. The hardest problems in both AI and biology are being solved here,...  ...with 5+ years of hands‑on research and engineering experience in generative modeling... 
    Full time

    Outputbiosciences

    New York, NY
    3 days ago
  • $180k - $280k

    Join to apply for the Full Stack Engineer role at OffDeal . Base Pay Range $180,000 - $280...  ...per year. OffDeal is the world’s first AI-native investment bank for small businesses...  .... Opportunity to be a foundational team member at a well‑funded, high‑growth startup.... 
    Relocation package

    OffDeal

    New York, NY
    2 days ago
  •  ...enabling companies to build, train, and serve AI models tailored to their own data,...  .... The Role As a Training Infrastructure Engineer, you'll design, build, and optimize the infrastructure...  ...of AI infrastructure, from low-latency inference to scalable model serving. Build What’s... 

    Fireworks AI

    New York, NY
    3 days ago
  •  ...companies to build, train, and serve AI models tailored to their own...  ...As a Training Infrastructure Engineer, you'll design, develop, and...  ...machine learning training, inference, and data processing...  ...backend infrastructure, lead technical design discussions, mentor engineers... 

    Fireworks AI

    New York, NY
    19 hours ago
  • Member of Technical Staff: Backend Monk is an AI-native accounts receivable (AR) platform for B2B companies helping businesses get paid fast. The AR stack...  ...is ahead of us. The role: We're looking for a backend engineer to join us in person in Flatiron, NYC. You'll own the... 
    Work at office
    Flexible hours

    Monk, Inc.

    New York, NY
    4 days ago
  • About Decagon Decagon is the leading conversational AI platform empowering every brand to deliver...  ...how we work and grow as a team. About The Role As a Member of Technical Staff, you'll join one of four engineering teams building the systems behind Decagon's AI agents... 
    Internship
    Work at office
    Local area

    Decagon

    New York, NY
    1 day ago
  • $140k - $270k

     ...focus on delivering care. We’ve built an AI-powered platform designed by...  ...rest will follow. The Team At Anterior, engineers share a strong "sense of product" and...  ...in your application. About The Role Members of Technical Staff at Anterior own problems end-to-end —... 
    Apprenticeship
    Flexible hours

    SupportFinity™

    New York, NY
    2 days ago
  •  ...companies to build, train, and serve AI models tailored to their own...  ...runs on one of the busiest inference platforms in the world —...  ...new grads into teams across Engineering, and we match you to a team based...  ..., Engineering, or a related technical field, completed within the... 
    Summer work
    Internship

    Fireworks AI

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff (AI Inference Engineer). Be the first to apply!