Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer- GPU Fabric Observability

Full-time

Baseten

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE
Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not.


We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding.


This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next.


EXAMPLE INITIATIVES

  • Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fabric, host, GPU, and workload telemetry.

  • Service-aware fabric diagnosis — Build collectors, probes, and topology-aware correlation to detect latency, drops, stalls, congestion, bad paths, and degradation.

  • Software-aware RDMA diagnosis — Tie network behavior to RDMA operations, GPUDirect paths, KV cache transfers, prefill/decode disaggregation, and request latency.

RESPONSIBILITIES

  • Own Baseten’s GPU fabric observability and root-cause analysis architecture.

  • Build telemetry pipelines across switches, NICs, hosts, GPUs, Kubernetes, and inference services.

  • Model topology, flow paths, service ownership, and failure domains.

  • Separate true fabric faults from host, NIC, GPU, kernel, driver, RDMA, scheduler, and workload failures.

  • Create clear operator workflows for triage, remediation, and post-incident learning.

REQUIREMENTS

  • Staff-level or senior staff-level experience building production infrastructure software.

  • Strong distributed systems background, especially streaming systems, telemetry pipelines, diagnostics, or control-plane software.

  • Experience building systems that process high-volume, high-cardinality, noisy operational data.

  • Understanding of networking fundamentals and high-performance networks

  • Ability to work with low-level infrastructure signals and build practical correlation, anomaly detection, or root-cause analysis systems.

BENEFITS

  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Software Engineer- GPU Fabric Observability in San Francisco, CA vacancy
  •  ...and help build the platform engineers turn to to ship AI products....  ...foundational engineers to lead our GPU Networking efforts, making...  ...to architect the software fabric that unifies thousands of GPUs...  ...and minimal latency. Build Observability: You will design the tools... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    3 days ago
  •  ...teams spanning hardware and software. Speed and scale are our key...  ...forward. The Production Engineering Team Examples of key exciting...  ...0s of GWs: at our scale, a GPU failure isn't a ticket. It's...  ..., at any scale: build the observability and orchestration layer that... 
    Suggested
    Local area

    FluidStack

    San Francisco, CA
    5 days ago
  • $200k - $280k

     ...lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud...  ...generative AI platform. The storage and observability team is crucial for designing, implementing...  ...critical insights into system performance and GPU utilization, and proactively identifying... 
    Suggested
    Full time
    Remote work

    Together AI

    San Francisco, CA
    1 day ago
  •  ...speed, security, and exceptional developer experience.Now, software is entering a new era, and the next generation of...  ...comes next.About the Role:We are looking for a Software Engineer to join our Observability team. Vercel users rely on Observability to monitor and... 
    Suggested
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Monday to Friday
    Flexible hours

    Vercel

    San Francisco, CA
    3 days ago
  • $170k - $240k

    Senior Software Engineer - Observability and ReliabilityAbout the RoleWe are growing the engineering team and looking for engineers who have the chops to build and deliver world-class technology. You will be part of a talented team of engineers with a shared mission to... 
    Suggested
    Full time
    Work at office

    Sigma Computing

    San Francisco, CA
    1 day ago
  •  ...GPU Kernel Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable... 
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $139k - $257.55k

     ...experiences using Adobe Express.We are seeking an experienced Software Development Engineer to help build the infrastructure and tooling that powers...  ...is built and operated, you will help shape intelligent observability capabilities that enable engineering teams to... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    4 days ago
  •  ...Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering...  ...About the Role We’re building the observability product for OpenAI—from scalable...  ...through notebook-like UIs. We’re hiring software engineers across the stack—infra,... 
    Full time

    OpenAI

    San Francisco, CA
    3 days ago
  •  ...Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    3 days ago
  •  ...well as accelerating research progression via model inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack - from low-level kernel performance to high-level... 
    Full time

    OpenAI

    San Francisco, CA
    3 days ago
  • $230k - $405k

     ...compute into a reliable engine for frontier AI. We...  ...centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience...  ...protocols, RDMA, NCCL, GPU hardware behavior,...  ...performance networking fabrics, protocols, and observability... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    4 days ago
  • $293k - $325k

     ...preparing for future transformative technologies, and engaging a robust security culture. About the RoleWe are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that... 
    Work at office
    Local area
    Remote work
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  • $202k - $237k

     ...distributed computing and make it accessible to software developers of all skill levels. We’re...  ...We are seeking a Backend Software Engineer to join our team focused on building...  ...About the team The Workspace & Observability Team is dedicated to empowering clients... 
    Full time
    Work at office
    Flexible hours

    Anyscale

    San Francisco, CA
    3 days ago
  • $250k

     ...infrastructure provider building a next-generation GPU platform designed for AI training,...  ...for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and...  ...to improve reliability, automation, and observability across distributed compute environments... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...of PositionAs a Senior Software Engineer - Simulation & ML Platform...  ...execution.Develop secure, GPU-enabled containerized...  ...trajectories, control commands, sensor observations, timing, and safety status.... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    1 day ago
  • $180k - $250k

     ...Fal Engineer Position Fal is the generative media ecosystem powering...  ...performance inference, orchestration, and observability come together to unlock new...  ...-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive.... 
    Local area
    Relocation package

    fal

    San Francisco, CA
    4 days ago
  • $197.3k - $313.7k

     ..., and how the team evolves.How We Build: AI-FirstUnified Observability is being built with an AI-first engineering strategy. AI assistants are core to how we design, write, review, test, and operate our software. Engineers on this team are expected to:Use AI coding assistants... 
    Full time

    Salesforce

    San Francisco, CA
    4 days ago
  • $210k

     ...never before possible. About the RoleThe Engineering Acceleration team designs, builds and...  ...engineering teams on best practices for ensuring observable, scalable systems.Like all other teams,...  ...is a large-scale deployment of GPU nodes running in dozens of Kubernetes clusters... 
    Work at office
    Local area
    Immediate start
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  • $258k - $376k

     ...real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us!Figma’s Observability engineering team builds and operates the systems that give us deep visibility into the health, performance, and efficiency of our platform... 
    Minimum wage
    Full time
    Local area
    Remote work
    Flexible hours

    Figma

    San Francisco, CA
    1 day ago
  • A global open source software provider is seeking a Junior Software Developer for their Observability team. This remote position requires expertise in Python and a working knowledge of Go. The successful candidate will develop a cloud-native monitoring stack, collaborating... 
    Remote job

    Canonical

    San Francisco, CA
    3 days ago
  • $179k - $218k

     ...bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical authority...  ...operational and engineering standards.Technical RequirementsSilicon & Fabric Mastery: Expert-level knowledge of NVIDIA (Hopper/Blackwell/... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $145k - $180k

     ...ubiquitous. We build the foundation for agent engineering in the real world, helping developers...  ...work on LangSmith, our commercial AI observability and evals platform product. In this...  ...this role ~2+ years of experience in software engineering working on complex platform... 
    Full time
    Work at office
    Flexible hours

    LangChain

    San Francisco, CA
    3 days ago
  • $300 per month

     ...About the Role: We are seeking Senior Software Engineers to design and develop internal datacenter...  ...-Deployment Automation readies every GPU serverHealthy by design: How Crusoe...  ...environmentsBackground in infrastructure observability or monitoring systemsBenefits:Industry... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • $300 per month

     ...Crusoe.About This RoleWe’re looking for a Senior Streaming Software Engineer to join the Observability team within our Cloud Infrastructure organization. This...  ...massive volumes of telemetry data generated across our GPU cloud and global data centers. Your work will help... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $166k - $201k

     ...at Crusoe.About This Role:As a Senior Software Engineer on our storage team, you'll be joining...  ...most demanding customer workloadsBuilding observability, metrics and tooling for our services...  ...with modern storage technologies (e.g GPU Direct Storage, F2FS, SPDK etc)Prior experience... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • $50 per hour

     ...seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure...  ...for the management of a fleet of GPU servers as well as the data centers that...  ...and implementing advanced diagnostic, observability, automation and repair tooling for high... 
    Temporary work

    Crusoe

    San Francisco, CA
    5 days ago
  •  ...fault recovery across large multi-node GPU clusters. Own the reliability, latency...  ...network partitions. Improve platform observability through metrics, tracing, and alerting....  ...decisions for core infrastructure and mentor engineers on distributed-systems practices.... 
    Full time
    Visa sponsorship
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    8 days ago
  •  ...About the role We're looking for a Software Engineer to help move data from AI models to actuators...  ...behavior across controls and vision observable and debuggable Minimum...  ...experience: camera drivers, image transport, GPU inference, or sensor synchronization... 
    Internship
    Immediate start

    Gradient Robotics

    San Francisco, CA
    4 days ago
  •  ...musicians, designers, visual artists, and engineers. This role We're looking for a...  ...occasional reach into Python ML inference on GPU clusters and our multi-cloud k8s setup...  ...Own cross-service metrics, tracing, and observability for the features you ship What we're... 
    Full time
    H1b
    Work at office
    Visa sponsorship
    Flexible hours

    Krea.ai, Inc

    San Francisco, CA
    1 day ago
  •  ...Junior Software Developer – Observability Join to apply for the Junior Software Developer – Observability role at Canonical Canonical is a leading...  ...initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's... 
    Work at office
    Remote work
    Work from home

    Canonical

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer- GPU Fabric Observability. Be the first to apply!