Software Engineer- GPU Fabric Observability
Baseten
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.
THE ROLE
Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not.
We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding.
This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next.
EXAMPLE INITIATIVES
Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fabric, host, GPU, and workload telemetry.
Service-aware fabric diagnosis — Build collectors, probes, and topology-aware correlation to detect latency, drops, stalls, congestion, bad paths, and degradation.
Software-aware RDMA diagnosis — Tie network behavior to RDMA operations, GPUDirect paths, KV cache transfers, prefill/decode disaggregation, and request latency.
RESPONSIBILITIES
Own Baseten’s GPU fabric observability and root-cause analysis architecture.
Build telemetry pipelines across switches, NICs, hosts, GPUs, Kubernetes, and inference services.
Model topology, flow paths, service ownership, and failure domains.
Separate true fabric faults from host, NIC, GPU, kernel, driver, RDMA, scheduler, and workload failures.
Create clear operator workflows for triage, remediation, and post-incident learning.
REQUIREMENTS
Staff-level or senior staff-level experience building production infrastructure software.
Strong distributed systems background, especially streaming systems, telemetry pipelines, diagnostics, or control-plane software.
Experience building systems that process high-volume, high-cardinality, noisy operational data.
Understanding of networking fundamentals and high-performance networks
Ability to work with low-level infrastructure signals and build practical correlation, anomaly detection, or root-cause analysis systems.
BENEFITS
Competitive compensation, including meaningful equity.
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
Paid parental leave
Fertility and family-building stipend through Carrot
Company-facilitated 401(k)
Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).
- ...and help build the platform engineers turn to to ship AI products.... ...foundational engineers to lead our GPU Networking efforts, making... ...to architect the software fabric that unifies thousands of GPUs... ...and minimal latency. Build Observability: You will design the tools...SuggestedFull timeFlexible hours
$175k - $300k
...teams spanning hardware and software. Speed and scale are our key... ...forward. The Production Engineering Team Examples of key exciting... ...GW fleet: at our scale, a GPU failure isn't a ticket. It's... ..., at any scale: build the observability and orchestration layer that...SuggestedFull timeLocal area$200k - $280k
...lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud... ...generative AI platform. The storage and observability team is crucial for designing, implementing... ...critical insights into system performance and GPU utilization, and proactively identifying...SuggestedFull timeRemote work- ...energy and intelligence. We are seeking a Software Engineer to join Crusoe’s Data Center... ...focusing on software for managing a fleet of GPU servers and the data centers that house... .... You will build advanced diagnostics, observability, automation and repair tooling for high...Suggested
$215k - $285k
Senior Software Engineer, GPU Sandboxes Location: San Francisco, CA Company Stage of Funding: Series A Office Type: In Person Salary: $... ...initial architecture and implementation through deployment, observability, on-call, and ongoing reliability. Work closely with the...SuggestedFull timeWork at office- ...Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering... ...About the Role We’re building the observability product for OpenAI—from scalable... ...through notebook-like UIs. We’re hiring software engineers across the stack—infra,...Full time
- ...Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early...Full timeFlexible hours
- ...including BOND, IVP, Spark Capital, Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code...Full timeFlexible hours
- ...well as accelerating research progression via model inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack - from low-level kernel performance to high-level...Full time
$170k - $240k
...Senior Software Engineer - Observability and Reliability About the Role We are growing the engineering team and looking for engineers who have the chops to build and deliver world-class technology. You will be part of a talented team of engineers with a...Full timeWork at officeFlexible hours- ...speed, security, and exceptional developer experience.Now, software is entering a new era, and the next generation of... ...comes next.About the Role:We are looking for a Software Engineer to join our Observability team. Vercel users rely on Observability to monitor and...Work experience placementWork at officeRemote workWork from homeWorldwideMonday to FridayFlexible hours
$202k - $237k
...distributed computing and make it accessible to software developers of all skill levels. We’re... ...We are seeking a Backend Software Engineer to join our team focused on building... ...About the team The Workspace & Observability Team is dedicated to empowering clients...Full timeWork at officeFlexible hours- Thinking Machines Lab Inc. is seeking a network engineer to own the lowest layers of the network stack for large-scale training and inference. You will ensure interconnect reliability across GPU fabrics, debugging NICs, and building instrumentation for faster troubleshooting...
- ...superintelligence. One person, one GPU. If you'd like to build... ...We are seeking a Senior Software Engineer to join our Managed... ..., Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and... ...clusters Experience with observability at scale: Prometheus, Grafana...Full timeWork at officeLocal areaWork from homeFlexible hours
- ...Team The Core Network Engineering team owns the end-to-... ...networking, datacenter fabrics, or global WAN... ...span low-level systems software, distributed infrastructure... ..., protocol readiness, observability, performance engineering... ...and high-performance GPU interconnects Define...Full time
- Crusoe is seeking a Senior Streaming Software Engineer to join the Observability team in our Cloud Infrastructure group. You will design, build, and operate... ...systems that power real‑time telemetry across our GPU cloud and global data centers. You’ll work with Kafka,...
$139k - $257.55k
...experiences using Adobe Express.We are seeking an experienced Software Development Engineer to help build the infrastructure and tooling that powers... ...is built and operated, you will help shape intelligent observability capabilities that enable engineering teams to...Full timeTemporary workLocal areaWorldwide- NomadicML is seeking a Backend / Infrastructure Engineer to build and scale the video... ...from secure cloud ingestion to distributed GPU inference. You’ll collaborate with ML researchers... ...systems to frontend workflows with robust observability and reliability. #J-18808-Ljbffr SpurWorldwide
$145k - $180k
...ubiquitous. We build the foundation for agent engineering in the real world, helping developers... ...work on LangSmith, our commercial AI observability and evals platform product. In this... ...this role ~2+ years of experience in software engineering working on complex platform...Full timeWork at officeFlexible hours- ...to improve their business. Founded by engineers — and customer obsessed — we leap at every... ...'re only getting started. As a Senior Software Engineer on the Customer Foresight Team... ...a reliable and timely interface for observability Make it easy for frameworks teams across...Worldwide
- Apple is seeking engineers to help build observability infrastructure for CloudKit and iCloud services. You will design, develop, and operate distributed backend services that ingest vast volumes of operational data and run on Kubernetes at scale. You will collaborate across...
$293k - $325k
...preparing for future transformative technologies, and engaging a robust security culture. About the RoleWe are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that...Work at officeLocal areaRemote workRelocation packageFlexible hours$240k - $280k
...About the Role We're looking for a Software Engineer to build the systems that treat infrastructure... ...from discovery, inference bring-up to GPU driver/CUDA stack, health validation,... ...or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure...Full time- Snowflake is seeking a Senior Software Engineer to design and implement OTel Collector components and instrumentation libraries. You will collaborate... ...and contribute to upstream specifications, while owning the Observe Agent architecture and release strategy. You will shape...
- Apple Inc. is seeking a Senior Software Engineer for Apple Services Engineering in San Francisco. You will design and build fault-tolerant distributed systems powering observability across CloudKit and iCloud services, operating within Kubernetes and across large-scale...
- ...superintelligence. One person, one GPU. If you'd like to... ...is currently Tuesday. Engineering at Lambda is... ...looking for an experienced Software Engineer to help... ...network services through observability, failover, and redundancy... ..., Cisco ACI or Nexus Fabric Controller, Arista CVP...Work at officeLocal areaWork from homeFlexible hours
$160k - $195k
...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to... ...the real world. Today, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open...Full timeWork at officeFlexible hours$250k
...infrastructure provider building a next-generation GPU platform designed for AI training,... ...for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and... ...to improve reliability, automation, and observability across distributed compute environments...Full timeRemote work- ...We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams... ...occur. You’ll also work on improving observability, rollout safety, release automation, and... ...failures caused by infrastructure instability, GPU scheduling, or test environment issues...Full time
- ...and help build the platform engineers turn to to ship AI products.... ...reliability, and ease of use. As a Software Engineer on the Inference... ..., autoscaling, scheduling, observability, and runtime management... ...distributed runtimes, networking, and GPU workloads Make thoughtful...Full timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer- GPU Fabric Observability. Be the first to apply!
- software engineer internship San Francisco, CA
- software development engineer aws San Francisco, CA
- software developer internship no experience San Francisco, CA
- real time software engineer San Francisco, CA
- financial software developer San Francisco, CA
- part time software developer San Francisco, CA
- graduate software developer San Francisco, CA
- software engineer travel San Francisco, CA
- experienced software developer San Francisco, CA
- remote entry level software developer San Francisco, CA



