Software Engineer- GPU Fabric Observability
$200k - $380kBaseten
About Baseten Baseten powers mission‑critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting‑edge models into production. We’re growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. The Role Baseten is building its own GPU infrastructure for large‑scale inference. As we move into large‑scale, high‑density NVIDIA systems, the hardest failures are intermittent, cross‑layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis‑tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first‑class observability and root‑cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high‑volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real‑time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. Example Initiatives Real‑time telemetry engine — Build the ingestion, reduction, storage, and query path for high‑cardinality fabric, host, GPU, and workload telemetry. Service‑aware fabric diagnosis — Build collectors, probes, and topology‑aware correlation to detect latency, drops, stalls, congestion, bad paths, and degradation. Software‑aware RDMA diagnosis — Tie network behavior to RDMA operations, GPUDirect paths, KV cache transfers, prefill/decode disaggregation, and request latency. Responsibilities Own Baseten’s GPU fabric observability and root‑cause analysis architecture. Build telemetry pipelines across switches, NICs, hosts, GPUs, Kubernetes, and inference services. Model topology, flow paths, service ownership, and failure domains. Separate true fabric faults from host, NIC, GPU, kernel, driver, RDMA, scheduler, and workload failures. Create clear operator workflows for triage, remediation, and post‑incident learning. Requirements Staff‑level or senior staff‑level experience building production infrastructure software. Strong distributed systems background, especially streaming systems, telemetry pipelines, diagnostics, or control‑plane software. Experience building systems that process high‑volume, high‑cardinality, noisy operational data. Understanding of networking fundamentals and high‑performance networks. Ability to work with low‑level infrastructure signals and build practical correlation, anomaly detection, or root‑cause analysis systems. Benefits Competitive compensation, including meaningful equity. 100% coverage of medical, dental, and vision insurance for employee and dependents. Flexible PTO policy including company‑wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!). Paid parental leave. Fertility and family‑building stipend through Carrot. Company‑facilitated 401(k). Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities. Compensation Range
$200K – $380K
Apply Now Embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward‑thinking team, we would love to hear from you. At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status. We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable). #J-18808-Ljbffr Baseten- Baseten is seeking a Lead Software Engineer to own GPU fabric observability and RCA at scale in San Francisco. You will drive real-time telemetry, cross-layer diagnosis, and actionable guidance for operators during incidents. The role sits at the intersection of networking...Suggested
$180k - $200k
Lightning AI seeks an Observability Infrastructure Engineer to enhance observability systems across large-scale, GPU-enabled infrastructure. This position focuses on designing telemetry pipelines, improving data insights, and creating monitoring solutions. With a commitment...SuggestedRemote job- ...and help build the platform engineers turn to to ship AI products.... ...foundational engineers to lead our GPU Networking efforts, making... ...to architect the software fabric that unifies thousands of GPUs... ...and minimal latency. Build Observability: You will design the tools that...SuggestedFlexible hours
$150k - $215k
Nscale is looking for a Principal Observability Platform Engineer in San Francisco, California. You will own the technical strategy for Nscale's observability platform, driving decisions that impact infrastructure and operations at scale. The ideal candidate will have over...Suggested- Nscale is seeking a Staff Observability Platform Engineer based in San Francisco, California, to design and implement observability solutions for GPUs and AI workloads. The role entails partnering with SRE and engineering teams to enhance system visibility and reliability...Suggested
$200k - $280k
...lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud... ...generative AI platform. The storage and observability team is crucial for designing, implementing... ...critical insights into system performance and GPU utilization, and proactively identifying...Full timeRemote work$150k - $215k
Principal Observability Platform Engineer - Nscale About Nscale Nscale is the GPU cloud engineered for AI. We provide cost‑effective, high‑performance infrastructure for AI start‑ups and large enterprise customers. Nscale simplifies AI development while enabling superior...Flexible hours$180k - $360k
...our $300M SeriesE, backed by investors including BOND, IVP, Spark Capital, Greylock, and Conviction. The Role We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration. In this role you will craft the foundation that powers modern AI...Flexible hours- Harrison Clarke is seeking a Senior Platform Engineer to take ownership of their core platform in San... ...multi-region Kubernetes clusters, managing GPU infrastructure, overseeing networking systems, and developing observability across metrics and logs. The ideal candidate...
- ...lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. The storage and observability team is responsible for designing,... ...critical insights into system performance and GPU utilization, and proactively resolving issues...
$166k - $201k
...energy and intelligence. We’re crafting the engine that powers a world where people can... ...deep expertise in building and operating observability platforms at scale. You will design,... ...-volume workloads (AI/ML, HPC clusters, GPU infrastructure) Embedding security best...Temporary work- ...will build, integrate, and evangelize observability platforms and solutions for our products... ...product that improves productivity of engineers across the globe by several orders of magnitude... ...development of scalable, distributed software systems that support globally...
- About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance... ...future. About The Role As a Staff Observability Platform Engineer, you'll play a... ...observability is embedded throughout the software and infrastructure lifecycle. Drive...
$172.5k - $260.1k
...to ensure you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI CRM,... ...tools that produce telemetry, provide insights, and improve observability in Slack production services with a focus on performance...$170k - $240k
Senior Software Engineer - Observability and Reliability About the Role We are growing the engineering team and looking for engineers who have the chops to build and deliver world‑class technology. You will be part of a talented team of engineers with a shared mission to...Full timeWork at officeFlexible hours$172.5k - $260.1k
...telemetry, provide insights, and improve observability in Slack production services with a... ...cloud providers such as AWS, and develop software using a combination of Go, Python, or Java... ...and work closely with other teams in engineering, product development, and customer...- ...in San Francisco. We seek an ML Platforms Engineer to build and run the orchestration layer that... ...fleet running efficiently. You'll own the GPU orchestrator end-to-end, implement multi-tenant primitives, and drive observability with metrics, logs, and traces at scale. #J...
$190k - $290k
Senior Software Engineer - Customer Developer Observability Adyen provides payments, data, and financial products in a single solution for customers like Meta, Uber, H&M, and Microsoft - making us the financial technology platform of choice. At Adyen, everything we do...Work at officeFlexible hoursShift work$120k - $290k
PlanetScale is seeking a Software Engineer for its Insights team to build a customer-facing database observability product. This role involves developing APIs and dashboards for performance metrics and collaborating across teams to enhance database observability. Ideal...- Salesforce's Monitoring Infrastructure team focuses on log pipelines, observability, and data-driven insights to improve reliability at scale.... ...data routes, tooling, and interfaces, collaborating across engineering and customer experience to ensure timely delivery and robust...
- Slack is seeking a Software Engineer to join the Monitoring Infrastructure team within the Service Delivery Platform & Reliability group... ...log pipelines, develop tooling for data routing, and enhance observability across Slack production services. You will collaborate with...
$190k - $290k
Adyen is seeking a Senior Software Engineer for their Customer Developer Observability Team in San Francisco. The ideal candidate will lead technical projects, collaborate across teams, and enhance integration experiences for customers. Requirements include 7+ years of...Work at office$230k - $342k
...Team The Core Network Engineering team owns the end‑to‑end... ..., datacenter fabrics, or global WAN infrastructure... ...span low‑level systems software, distributed... ..., protocol readiness, observability, performance engineering... ...and high‑performance GPU interconnects Define...Full timeWork at officeLocal areaRelocation packageFlexible hours- A global open source software provider is seeking a Junior Software Developer for their Observability team. This remote position requires expertise in Python and a working knowledge of Go. The successful candidate will develop a cloud-native monitoring stack, collaborating...Remote job
- ...Production Lifecycle team, alongside Observability and Deploy Platform. This... ...and repetitive tasks. We use software and agents to keep the lights... ...About The Role As a Software Engineer on the Reliability Platform... ...Kafka topics, Databases, CPU/GPU pools, Service Scaffolding,...Hourly payWork at officeLocal areaRemote workFlexible hours
- ...construction veterans and world‑class engineers to solve physical‑world... ...team builds the software foundation Bedrock’s autonomous... ...connectivity Optimize CPU and GPU performance and scheduling for... ...Establish the diagnostics and observability that let a small team keep a...Work at officeFlexible hours
- USA Tech Recruit is looking for a Software Engineer (C++ Systems) to work in San Francisco. This full-time onsite role focuses on optimizing high-performance GPU virtualization technology, ideal for those passionate about low-level GPU infrastructure. The position requires...Full timeRelocation
- ...infrastructure powering a heterogeneous AI cloud. You will manage large CPU/GPU/accelerator clusters, bare‑metal provisioning, and production... ...and other schedulers. This role focuses on reliability, observability, and automation to scale deployment of AI workloads with day‑1...
- 10X Business Consulting is seeking a highly skilled Software Engineer (C++ Systems) to join our client’s team in San Francisco. You will own production systems from day one, optimize a GPU virtualization platform, and tackle challenging performance issues in a fast-growing...
- A leading consulting firm is seeking a Software Engineer (C++ Systems) in San Francisco to optimize microsecond-level performance in GPU virtualization software. Ideal candidates will have elite C++ expertise, with at least 2 years of experience in low-level systems engineering...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer- GPU Fabric Observability. Be the first to apply!
- ngo software engineer San Francisco, CA
- software developer San Francisco, CA
- software developer internship no experience San Francisco, CA
- junior software developer San Francisco, CA
- part time software developer remote San Francisco, CA
- financial software developer San Francisco, CA
- senior software engineer ruby on rails San Francisco, CA
- software engineer amazon San Francisco, CA
- senior software design engineer San Francisco, CA
- software engineer part time San Francisco, CA

