Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff, Inference

Full-time

Genesis

Role Description

  • Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics.
  • Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization.
  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks.
  • Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation).
  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks.

Qualifications

  • Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years).
  • Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go).
  • Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling.
  • Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments.
  • System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness.
Vacancy posted 8 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff, Inference in Remote vacancy
  • $150k - $300k

     ...position spanning cloud LLM serving, LLM inference optimization and RL systems. You will be...  ...into our RL training stack. Core Technical Responsibilities LLM Serving Multi‑tenant...  ...in open development and encourage team members to contribute to the broader AI community... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime-Intellect

    San Francisco, CA
    2 days ago
  • $200k - $400k

    About The Role We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving...  ...OpenRLHF, Unsloth, LlamaFactory, etc). Written widely-shared technical blogs or side projects on vLLM or LLM inference. Logistics... 
    Suggested
    Remote work
    Visa sponsorship
    Shift work

    Inferact

    San Francisco, CA
    1 day ago
  •  ...and reliable machine learning systems? Do you want to set technical direction and help shape the next generation of AI platforms...  ...advanced NLP applications? We are looking for a Lead Member of Technical Staff to join the Model Serving team at Cohere. The team is responsible... 
    Suggested
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Remote
    1 day ago
  • $240k - $280k

     ...Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how...  ...Referrals increase your chances of interviewing at Cabana by 2x Inferred from the description for this job 401(k) Get notified... 
    Suggested
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    2 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering...  ...ingestion, transformation, training/fine-tuning, and inference? You will also: Find opportunities to go deep into a wide... 
    Suggested
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    3 days ago
  • $160k - $190k

    Role Description As a Member of Technical Staff, Machine Learning, you will build core ML components. The Member of Technical Staff will work...  ...improve ML components across data, training, evaluation, and inference. ~Fine-tune and adapt models as part of larger... 
    Full time
    Remote work

    Ginas Tech Jobs

    Remote
    11 days ago
  • $125k - $200k

     ...Washington Post, TIME, CNN, and many others. About the role As a Member of Technical Staff, you will research, build, and refine the software...  ..., from the look and feel of the UI to the speed of the AI inference. Have an opinionated aesthetic sense of what makes a good... 
    Full time
    Internship
    Work at office
    Remote work
    Visa sponsorship
    Weekday work

    CivAI

    Berkeley, CA
    2 days ago
  •  ...York City, Montreal, Seoul, Germany and Paris. Join us! Member of Technical Staff, Search Why this role? We are looking for talented individuals...  ...Work closely with the model serving team to ensure that inference is fast and stable. Collaborate with product teams to... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  • $300k

     ...the ground up. About the Role We’re looking for a deeply technical Member of Technical Staff to own RL and post‑training for large‑scale omni models....  ...such as vLLM. Experience with large‑scale training or inference systems, including rollout generation, model serving,... 
    H1b
    Work at office
    Visa sponsorship
    Shift work

    Nuance Labs

    Seattle, WA
    4 days ago
  •  ...a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers...  ...with large‑scale datasets and distributed training or inference pipelines. Understanding of LLM architectures, tuning... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    4 days ago
  •  ...Activant, 1984 Ventures and Page One. The Role We’re hiring a Member of Technical Staff - AI/ML to design, build, and deploy AI-powered systems...  ...the needle Build robust AI pipelines from ingestion to inference — reliable, maintainable, and cost‑efficient through smart... 
    Full time
    Flexible hours

    Stuut

    New York, NY
    4 days ago
  • # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco / Bay Area preferred. Remote exceptional for the right person...  ...make AI workloads cheaper and easier to own by turning inference behavior, traces, workload replay, GPU signals, and task-path... 
    Full time
    Remote work

    Touchdown Labs, Inc.

    San Francisco, CA
    2 days ago
  • $250k

     ...queries a month, and every one of them fans out into multiple AI inference requests running in real time. Behind that sits a large GPU...  ...hiccups, and capacity shifts without human intervention. ~Set technical direction across teams. ~Partner with inference and cloud... 
    Full time
    Shift work

    Perplexity

    Remote
    5 days ago
  • $150k - $300k

     ...infrastructure that runs the jobs. Core Technical Responsibilities Hosted Training...  ...and operate Kubernetes-based training and inference orchestration across multi-cluster, multi...  ...in open development and encourage team members to contribute to the broader AI community... 
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Prime-Intellect

    San Francisco, CA
    4 days ago
  • $230k

     ...allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning users to effortlessly...  ...services. Cerebras Systems Inc. has multiple openings for Sr. Member of Technical Staff. Title: Sr. Member of Technical Staff Job... 
    Remote work

    Cerebras Systems

    Sunnyvale, CA
    5 days ago
  • $250k - $350k

    Member of Technical Staff - Quantitative Research New York City (Remote possible for exceptional candidates) About Uncharted/Udio Udio builds...  ...apply your findings to our pretraining, post-training and inference systems as applicable. Drive product & research roadmap You... 
    Work experience placement
    Remote work
    Flexible hours

    Udio

    New York, NY
    1 day ago
  •  ...enterprises, and we’re looking for a senior member for the Agent Code team. You’ll be at...  ...collaboration. As a Member of Technical Staff on the Agent Code team, you will: Stay...  ...models, and deploy agent frameworks for inference and sampling. You will be collaborating... 
    Full time
    Work at office
    Local area
    Remote work
    Home office
    Flexible hours

    Cohere

    New York, NY
    2 days ago
  • Member of Technical Staff - Agents at Prime Intellect - San Francisco Building the Future of Open Source + Decentralized AI At Prime Intellect...  ...Infrastructure : Understanding how to optimize agent training or inference on GPUs. Advanced AI/ML Knowledge : Familiarity with... 
    Remote work
    Flexible hours

    Victrays

    San Francisco, CA
    5 days ago
  • $200k - $300.09k

    Role Description The OpenClaw Foundation is seeking exceptional Members of Technical Staff (MTS) to serve as full-time maintainers, builders, and stewards of the OpenClaw ecosystem. This is not a traditional software engineering role. ~Operate as both an open-source... 
    Full time

    OpenClaw Foundation

    Remote
    5 days ago
  • $175k - $220k

     ...the highest-quality models with the fastest and most scalable inference in the industry. We've been independently benchmarked as the leader...  ...ll architect scalable, resilient backend infrastructure, lead technical design discussions, mentor engineers, and establish best... 

    Fireworks AI

    New York, NY
    3 days ago
  •  ...Career Launch is hiring for this role and related opportunities. Career Launch is hiring candidates for Member of Technical Staff roles and similar opportunities with fast-moving teams. This is a remote-friendly opportunity for candidates who are analytical, practical,... 
    Remote work

    Career-Launch.net

    San Francisco, CA
    7 hours ago
  •  ...our customers to insights in seconds and to business impact in minutes using our products. We are looking for a Founding Member of Technical Staff to join our team guiding the development and deployment of complex ML systems reporting to Dan Wald, Co-Founder & CAIO.... 
    Work at office
    Remote work

    Sciemo

    New York, NY
    4 days ago
  • $324k - $396k

     ...About the Role Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques and analyses to the AI field to facilitate breakthroughs in quantitative reasoning and language understanding. Stabilize large language model training, pipeline parallelism... 
    Remote work

    Xai

    Palo Alto, CA
    3 days ago
  •  ...products. We power the next generation of AI experiences that will reshape how people discover and buy online. Role As a Member of Technical Staff, you will ship core systems, set engineering culture, and move the mission from prototype to platform. You will work across... 
    Work at office

    Getcatalog

    San Francisco, CA
    3 days ago
  •  ...architectures, implement dynamic worker provisioning based on queue depth, optimize connection pooling and caching layers Deploy and optimize inference for vision‑language models powering our agents – low latency, high throughput, cost‑efficient GPU utilization Expand our... 
    Immediate start
    Remote work

    CloudCruise

    San Francisco, CA
    2 days ago
  •  ...looking for an Infrastructure Platform Engineer to design, build, and operate the cluster infrastructure behind Gimlet’s heterogeneous inference cloud. Unlike traditional cloud platforms built around a single hardware ecosystem, Gimlet's infrastructure spans multiple... 

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  • $324k - $396k

     ...to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques and analyses to the AI field to facilitate breakthroughs in... 
    Remote work

    Pantera Capital

    Palo Alto, CA
    3 days ago
  •  ...systems. You are a fast prototyper and hacker who can also write beautiful production code. You have experience running training and inference on cloud hardware, parallelizing data and models across accelerators. You are a data engineer. You have experience building... 
    Remote work
    Flexible hours

    Latent Labs

    United States
    5 days ago
  •  ...Job Role: Member of Technical Staff Job Location: Austin, TX (Remote) Job Description: ~ The Member of Technical Staff (MTS) - Systems is a senior individual contributor role that provides technical leadership and expertise for complex systems engineering... 
    Remote work

    ACL Digital

    United States
    5 days ago
  • $250k

     ...the creation of maintainable, scalable systems and make sound technical decisions. You lead large projects from ideation to delivery, balancing...  ...excellent, and you actively invest in the growth of your team members. You can hold a team to high standards while being comfortable... 
    H1b
    Work at office
    Work from home
    Home office
    Relocation package
    Flexible hours
    3 days per week

    METR

    Berkeley, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff, Inference. Be the first to apply!