Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer- GPU Fabric Observability

Full-time

Baseten

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE
Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not.


We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding.


This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next.


EXAMPLE INITIATIVES

  • Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fabric, host, GPU, and workload telemetry.

  • Service-aware fabric diagnosis — Build collectors, probes, and topology-aware correlation to detect latency, drops, stalls, congestion, bad paths, and degradation.

  • Software-aware RDMA diagnosis — Tie network behavior to RDMA operations, GPUDirect paths, KV cache transfers, prefill/decode disaggregation, and request latency.

RESPONSIBILITIES

  • Own Baseten’s GPU fabric observability and root-cause analysis architecture.

  • Build telemetry pipelines across switches, NICs, hosts, GPUs, Kubernetes, and inference services.

  • Model topology, flow paths, service ownership, and failure domains.

  • Separate true fabric faults from host, NIC, GPU, kernel, driver, RDMA, scheduler, and workload failures.

  • Create clear operator workflows for triage, remediation, and post-incident learning.

REQUIREMENTS

  • Staff-level or senior staff-level experience building production infrastructure software.

  • Strong distributed systems background, especially streaming systems, telemetry pipelines, diagnostics, or control-plane software.

  • Experience building systems that process high-volume, high-cardinality, noisy operational data.

  • Understanding of networking fundamentals and high-performance networks

  • Ability to work with low-level infrastructure signals and build practical correlation, anomaly detection, or root-cause analysis systems.

BENEFITS

  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer- GPU Fabric Observability in San Francisco, CA vacancy
  •  ...and help build the platform engineers turn to to ship AI products....  ...foundational engineers to lead our GPU Networking efforts, making...  ...to architect the software fabric that unifies thousands of GPUs...  ...and minimal latency. Build Observability: You will design the tools... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  • $200k - $280k

     ...lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud...  ...generative AI platform. The storage and observability team is crucial for designing, implementing...  ...critical insights into system performance and GPU utilization, and proactively identifying... 
    Suggested
    Full time
    Remote work

    Together AI

    San Francisco, CA
    3 days ago
  •  ...well as accelerating research progression via model inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack - from low-level kernel performance to high-level... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering...  ...About the Role We’re building the observability product for OpenAI—from scalable...  ...through notebook-like UIs. We’re hiring software engineers across the stack—infra,... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...including BOND, IVP, Spark Capital, Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  •  ...Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering...  ...play a crucial role in ensuring the observability, reliability, scalability, and...  ...with cross-functional teams, including software engineers, product managers, and data scientists... 
    Full time
    Work experience placement
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  • $202k - $237k

     ...distributed computing and make it accessible to software developers of all skill levels. We’re...  ...We are seeking a Backend Software Engineer to join our team focused on building...  ...About the team The Workspace & Observability Team is dedicated to empowering clients... 
    Full time
    Work at office
    Flexible hours

    Anyscale

    San Francisco, CA
    8 hours ago
  •  ...will build, integrate, and evangelize observability platforms and solutions for our products...  ...product that improves productivity of engineers across the globe by several orders of magnitude...  ...development of scalable, distributed software systems that support globally... 

    Retool

    San Francisco, CA
    1 day ago
  •  ...Team The Core Network Engineering team owns the end-to-...  ...networking, datacenter fabrics, or global WAN...  ...span low-level systems software, distributed infrastructure...  ..., protocol readiness, observability, performance engineering...  ...and high-performance GPU interconnects Define... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...future transformative technologies, and engaging a robust security culture.    About the Role We are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that... 
    Full time
    Remote work
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment...  ...to hardware caching  Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  • $140k - $175k

     ...Founded in 2023, LangChain powers top engineering teams at companies like Replit, Lovable...  ...Engineer to work on LangSmith, our commercial observability and evals platform product. In this...  ...this role ~2+ years of experience in software engineering working on complex platform... 
    Full time

    LangChain

    San Francisco, CA
    8 hours ago
  • $175k - $240k

     ...ubiquitous. We build the foundation for agent engineering in the real world, helping developers...  ...our commercial product LangSmith, an observability and evals platform. In this role, you'...  ...Bring ~5+ years of experience in software engineering working on complex platform... 
    Full time
    Work at office
    Flexible hours

    LangChain

    San Francisco, CA
    8 hours ago
  • $145k - $180k

     ...ubiquitous. We build the foundation for agent engineering in the real world, helping developers...  ...work on LangSmith, our commercial AI observability and evals platform product. In this...  ...this role ~2+ years of experience in software engineering working on complex platform... 
    Full time
    Work at office
    Flexible hours

    LangChain

    San Francisco, CA
    8 hours ago
  •  ...A global open source software provider is seeking a Junior Software Developer for their Observability team. This remote position requires expertise in Python and a working knowledge of Go. The successful candidate will develop a cloud-native monitoring stack, collaborating... 
    Remote work

    Canonical

    San Francisco, CA
    1 day ago
  • $200k - $250k

     ...leading cloud provider, is looking for a Software Engineer, Infrastructure Platform to build the...  ...asset management, DCIM, monitoring and observability, security, and operational automation—that...  ...platforms for rack operations, server/GPU deployment, OS installation, quality... 
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    8 hours ago
  •  ...us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer on the Capacity team,...  ...Instrument your work with observability and monitoring so issues surface...  ...infrastructure; familiarity with GPU infrastructure, or capacity... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  •  ...and help build the platform engineers turn to to ship AI products....  ...reliability, and ease of use. As a Software Engineer on the Inference...  ..., autoscaling, scheduling, observability, and runtime management...  ...distributed runtimes, networking, and GPU workloads Make thoughtful... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  •  ...build-out in history. When people finance GPU clusters, the datacenters housing them,...  ..., fast incremental feedback for every engineer, and a credible roadmap for what this looks...  ...unlimited paid time off as well as 10+ observed holidays Parental leave We offer... 
    Long term contract
    Full time
    Contract work
    Fixed term contract
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Shift work

    San Francisco Compute Company

    San Francisco, CA
    8 hours ago
  •  ...build-out in history. When people finance GPU clusters, the datacenters housing them,...  ...offtake. About the Role As a Product Engineer at SFC, you’ll build products which...  ...offer unlimited paid time off as well as 10+ observed holidays Parental leave We offer... 
    Long term contract
    Full time
    Contract work
    Fixed term contract
    Work at office
    Local area
    Visa sponsorship
    Shift work

    The San Francisco Compute Company

    San Francisco, CA
    8 hours ago
  •  ...We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams...  ...occur. You’ll also work on improving observability, rollout safety, release automation, and...  ...failures caused by infrastructure instability, GPU scheduling, or test environment issues... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  .... In this role, you’ll lead engineering efforts to ensure our largest...  .... Build tooling and observability to detect bottlenecks, guide...  ...teams. Mentor engineers on GPU performance, CUDA development...  ...issues across hardware and software layers. Have strong familiarity... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...before possible.  About the Role The Engineering Acceleration team designs, builds and...  ...engineering teams on best practices for ensuring observable, scalable systems. Help create a...  ...is a large-scale deployment of GPU nodes running in dozens of Kubernetes clusters... 
    Full time
    Immediate start
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...very different workload profiles, from GPU-attached systems to dedicated storage hardware...  ...is a hands-on infrastructure role for engineers who want to work on deeply technical...  ...Grafana, or similar infrastructure and observability tooling About OpenAI OpenAI is an... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...before possible.  About the Role The Engineering Acceleration team designs, builds and...  ...teams on best practices for ensuring observable, scalable systems. Like all other teams...  ...infrastructure is a large-scale deployment of GPU nodes running in dozens of Kubernetes... 
    Full time
    Immediate start
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...talented minds in research, engineering, and operations to create products...  ..., test, ship, and debug software across device and cloud...  ...engineering teams on pragmatic observability, reliability, and scalability...  ...-scale deployment of CPU/GPU nodes running in Kubernetes... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  • $250k - $350k

     ...Manager, Software Engineering - Observability Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating... 
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office
    Remote work
    Work from home

    Figma

    San Francisco, CA
    2 hours ago
  •  ...distributed systems. We do the engineering and research required to...  ...build our own model training software, and focus on the lower layers...  ..., fault tolerance, and observability. The models we train are...  ...distribute work across massive GPU clusters efficiently. Design... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...a Tokens-as-a-Service (TaaS) Engineer to help build the systems that...  ...stack, ensuring GPU capacity can be onboarded, measured...  ...into OpenAI’s internal compute, observability, and workload management systems...  ...across hardware, networking, software, and workload enablement that... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  • $160k - $230k

    Senior Software Engineer - Together Cloud Infrastructure Together AI is building the AI Acceleration...  ...for pretraining. Build advanced observability stacks for our customers with automated...  ...Experience with DPUs/SmartNICs a plus GPU programming, NCCL, CUDA knowledge a plus... 
    Full time
    Work at office
    Remote work

    Together AI

    San Francisco, CA
    2 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer- GPU Fabric Observability. Be the first to apply!