Senior Site Reliability Engineer
Andromeda Cluster, Inc
Senior Site Reliability Engineer Location: Global Remote / San Francisco • Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers. We began with a single managed cluster - but it filled almost instantly. Since then, we've been quietly building the systems, network, and orchestration layer that makes the world's AI infrastructure more accessible. Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it's needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth. Our long-term vision is to build the liquidity layer for global AI compute - a marketplace that moves the infrastructure and workloads powering AGI not dissimilar to the flows of capital in the world's financial markets. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering. The Role This is not a generalist SRE role. You will design, operate, and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems. We're looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework. What You'll Own
- GPU Cluster Architecture: Design and evolve multi-provider, multi-region GPU compute clusters optimized for large-scale training. Make topology-aware scheduling, networking, and storage decisions that directly impact training throughput and cost efficiency.
- Customer Technical Partnership: Serve as the primary technical point of contact for customers running large-scale training workloads. Onboard, troubleshoot, and optimize, often in real time.
- Reliability & Performance Engineering: Define SLOs and error budgets that account for the unique failure modes of GPU infrastructure (ECC errors, NVLink degradation, NCCL timeouts). Own capacity planning across heterogeneous GPU fleets optimized for training throughput.
- Networking & Fabric Health: Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations.
- Observability: Build deep visibility into GPU utilization, memory pressure, interconnect throughput, training job performance, and hardware health. Go well beyond standard infrastructure metrics.
- Automation & Tooling: Build production-grade automation for cluster provisioning, GPU health checks, job scheduling, self-healing, and firmware/driver lifecycle management.
- Incident Leadership: Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Drive blameless postmortems and systemic fixes.
- GPU Systems Expertise: Deep, hands-on experience operating large-scale GPU clusters (NVIDIA A100/H100/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience not documentation.
- High-Performance Networking: Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale.
- Distributed Training & ML Frameworks: Working knowledge of how large training jobs actually run - NCCL, CUDA, PyTorch distributed, DeepSpeed, Megatron, FSDP, or similar. You don't need to write the models, but you need to understand what's happening at the systems level when a 1,000-GPU training run stalls.
- Linux & Systems Internals: Expert-level Linux knowledge: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, performance profiling at the syscall and hardware level.
- Kubernetes & Orchestration: Strong experience running Kubernetes in production with GPU workloads, including device plugins, topology-aware scheduling, multi-cluster federation, and custom operators. Experience with Slurm or other HPC schedulers is equally valued.
- Automation & Software Engineering: Strong engineering skills in Python, Go, or Bash. You build production-grade tools and services, not just scripts. Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).
- Observability & Monitoring: Hands-on experience building monitoring and alerting for GPU infrastructure, not just Prometheus/Grafana basics, but GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.
- Incident Management: Proven track record leading incident response for complex distributed systems where the failure could be in hardware, firmware, networking, drivers, orchestration, or application code and you need to narrow it down fast.
- Distributed Storage: Experience with high-performance parallel file systems (VAST, Weka, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs.
- Training Optimization: Experience profiling and optimizing distributed training performance: identifying stragglers, tuning collective communication strategies, improving MFU (Model FLOPs Utilization), and reducing idle GPU time across large runs.
- Cluster Buildout & Hardware: Experience involved in physical cluster design - rack layout, power/cooling constraints, network topology design, and hardware validation/burn-in at scale.
- Team Leadership: Experience leading or mentoring a team of infrastructure engineers. We're growing and need people who raise the bar for everyone around them.
Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
- ...our Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to-... ...expect us to hit our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team to: Own uptime,...SeniorContract workWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours
$232k - $319k
...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...with self-service Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and...SeniorPermanent employmentFull timeLocal areaWorldwideFlexible hours$210k - $240k
...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $2...SeniorFull time$175k - $250k
...000.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance... ...scalability, performance, and reliability across environments. What You’ll Do Design...SeniorFull timeRemote workRelocationRelocation package$148.5k - $223.9k
...duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce... ...future of Salesforce. Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...SeniorWorldwideWeekend work$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...SeniorTemporary workImmediate startFlexible hoursShift work$166.9k - $225.9k
...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team... ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building...SeniorWork at officeImmediate startWorldwideMonday to FridayFlexible hours$220k - $235k
...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move faster, insights surface instantly, and agents push work forward, all with you in control. Whether you're buying or selling...SeniorFull timeContract workWork at office$210.8k - $272.8k
About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user...SeniorLocal area$181k - $263k
...and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a...SeniorWork from homeFlexible hoursNight shift- ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will...Senior
$300k
...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring...SeniorPermanent employment$250k
...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves...SeniorPermanent employmentRemote work$175k - $215k
...lifecycle. The Role & Your Mission We’re seeking a hands-on Senior AI Engineer who can bridge the gap between data and biomedical expertise... ...models consistently produce accurate, grounded, useful, and reliable outputs under real-world load. The ideal candidate thrives in...SeniorFull timeWork at officeRemote work- ...gets stuff built. You'll work closely with engineers, designers, and end-users to accelerate our product development. As a senior-level software engineer at Pulley, you will... ...regulations, and jurisdiction workflows into fast, reliable product experiences, with AI at the core of...SeniorFull time
$121.4k - $173.3k
...work remotely on the remaining days. On-site expectations may evolve over time to... ...seeking a highly skilled and motivated Senior Software Engineer to join our BaaS team, supporting... ...root cause analysis to improve system reliability. Draft architectural and design documentation...SeniorFull timeWork at officeRemote workFlexible hours3 days per week$260k - $275k
...SENIOR PRINCIPAL SOFTWARE ENGINEER Saviynt is an identity platform built to power and protect the world at work. With the rise of AI and Agents... ...in engineering processes, tooling, and operational reliability. Collaborate with internal teams to produce software...Senior$130k - $250k
...internationally. Our Team As an engineering team, we believe strongly that empathy... ...serves. We're looking for a Senior Software Engineer to join our core engineering... ...product requests into strong and reliable software components What we look for...SeniorWork at officeLocal area- ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems... ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes...Senior
$140k - $240k
...radar systems. Working alongside software, hardware, and systems engineers, you’ll design and operate the backend services and cloud... ...-critical sensor data at scale. The systems you build must be reliable, resilient, secure, and performant—enabling operators to depend...SeniorPermanent employmentWork experience placementCasual workRelocation package- ...and 7Wire Ventures. About the Role We are looking for a Senior Software Engineer — an individual contributor who owns a product surface end-... ...governed by HIPAA, PCI-DSS, and SOC 2. Responsibilities Deliver Reliable Craft on Your Surface Deliver reliable, high-quality craft...Senior
$170k - $230k
...history of writing scalable, performant and maintainable code. We strongly believe languages can be learned and care more about your engineering skill over frameworks * Excitement about shipping customer centric software * Experience with Java >= 11, JPA ORM mapping,...SeniorFull timeWork at officeLocal areaHome officeFlexible hours$130k - $196.5k
...from the ground up and migrate existing complex use cases into that system. Work with a team of supportive and passionate software engineers. Architect and implement systems that materialize our platform vision. Provide operational support for our production systems...SeniorFull timeWork from homeFlexible hoursNight shift$165k - $247k
...identity resolution, data import and export connectors, privacy and compliance, and the operational reliability of everything in between. As a Senior Software Engineer, you'll take on complex infrastructure challenges: designing for extreme throughput, optimizing for...SeniorFull timeHome officeFlexible hours$215k - $323k
...excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Solution Sales Engineer - Position Description: We believe Solutions Engineers at Okta are involved in all stages of the customer's...SeniorFull timeWork at officeLocal areaWorldwideFlexible hours$180k - $230k
...Sift Sift is the data infrastructure platform for hardware engineering teams. Sift turns high-frequency telemetry into engineering insights... ...software engineers to optimize application performance and reliability. Implement monitoring, alerting, and logging systems to...SeniorFull timeWork at officeRelocation- ...mission-critical industries, helping partners move more quickly and reliably from algorithm to silicon. Our platform accelerates deployment... .... The Roles We are looking for an experienced software engineer to help us build a new generation of transpilation tools...SeniorFull timeRemote workRelocation packageFlexible hours
- ...company valued at $10 billion. We work in‑person five days a week in our new SanFrancisco headquarters. About the Role As a Site Reliability Engineer (SRE) at Mercor, you’ll own production reliability across our most critical systems, partnering directly with...
- ...Open Source LLM Gateway Engineer LiteLLM is an open-source LLM Gateway with 34K+ stars on GitHub and trusted by companies like NASA... ...expanding and seeking our 6th Engineer focused on owning reliability, performance, and infrastructure stability for the LiteLLM proxy...
$152.5k - $219.2k
...global cloud platform. As a team of six engineers distributed across the US, Canada, and the... ...with a strong focus on automation, reliability, and operational excellence. We are one... ...Qualifications ~2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer San Francisco, CA
- site reliability engineer sre San Francisco, CA
- sr hr business partner San Francisco, CA
- senior lighting artist San Francisco, CA
- senior planner San Francisco, CA
- senior hvac project manager San Francisco, CA
- home instead senior care San Francisco, CA
- research associate senior research associate San Francisco, CA
- senior technical product manager San Francisco, CA
- senior cloud infrastructure engineer San Francisco, CA




