Senior Site Reliability Engineer
Luma AI
About Luma AI Luma’s mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. This requires a massive, reliable, and performant GPU infrastructure that pushes the boundaries of scale. Our SRE team is the foundation of our research and product velocity, responsible for the thousands of NVIDIA and AMD GPUs across multiple providers that power our work. Where You Come In We are looking for a hands‑on, first‑principles engineer who is fluent in Linux, comfortable operating close to the metal, and capable of architecting systems for the next generation of AI infrastructure. You will build, maintain, and scale Luma’s infrastructure across on‑prem and multi‑vendor clouds (AWS & OCI), serving as the bridge between hardware vendors, cloud providers, and our research teams. What You’ll Do Architect for Reliability & Scale: Participate in critical re‑architecture sessions to redesign our systems for higher efficiency and scale. You won't just maintain existing clusters; you will help define how our next‑generation infrastructure operates. Own Multi‑Cloud GPU Clusters: Take end‑to‑end ownership of our production clusters for training and inference across AWS and OCI, ensuring high availability and peak performance. Drive Security & Compliance: Assist in achieving and maintaining security certifications (SOC 2 Type 1 & 2, ISO standards) by implementing robust infrastructure security practices in a fast‑moving AI startup environment. Deep Linux Performance Tuning: Use your mastery of Linux systems to troubleshoot and optimize performance at the OS and kernel level. Build Robust Automation: Write high‑quality tools and automation in Python, Go, or Bash to manage, monitor, and heal our infrastructure without relying on heavy operational toil. Debug Complex Hardware/Software Failures: Serve as the final escalation point for the most challenging GPU, networking (InfiniBand/RDMA), and system‑level issues, often collaborating directly with hardware vendors like NVIDIA. Who You Are 5+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast‑paced, large‑scale environment. Deep Linux Mastery: You possess deep, hands‑on expertise in Linux, containerized systems, and debugging low‑level system performance. Expert in Technologies: You have working experience with Terraform, Airflow, and Ray. Cloud Infrastructure Expert: You have strong experience with providers like AWS or OCI. Tenacious Troubleshooter: You thrive on solving complex, low‑level problems where hardware and software intersect. Startup DNA: You are energetic and thrive in a less structured, fast‑paced environment. Security‑Minded: You possess a working knowledge of security best practices and familiarity with compliance frameworks, such as SOC 2 and ISO. Expert in High‑Performance Networking: You have practical experience with InfiniBand, RDMA, or RoCE and understand how to optimize throughput for massive distributed training jobs. What Sets You Apart (Bonus Points) Deep expertise with GPU tooling for NVIDIA and AMD GPUs like DCGM or ROCm. Experience managing large‑scale GPU clusters for AI/ML workloads (training or inference). Familiarity with job management systems based on Kubernetes or orchestration frameworks like Ray. Deep expertise in Data Pipeline and Infrastructure #J-18808-Ljbffr Luma AI
- ...Senior SRE Unify is building the first AI-native outbound platform where agents and... ...Senior SRE, you'll tackle the scaling and reliability challenges that come with adding... ...tracing, metrics, and alerting that give engineers clear visibility into system behavior and...Senior
- ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate...Senior
$181.69k - $213.75k
...Senior Site Reliability Engineer San Francisco, California; Santa Clara, California; Seattle, WA The Company You'll Join Carta connects founders, investors, and limited partners through world-class software, purpose-built for everyone in venture capital, private...SeniorFull timeWork at office$127k - $249k
THE TEAM Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ..., alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$195k - $240k
...Senior Site Reliability Engineer San Francisco (Hybrid) At You.com, we are building the AI Search Infrastructure that powers modern AI systems. Our goal is to create the trusted knowledge layer that agents, applications, and enterprises rely on to retrieve real-...SeniorFull timeImmediate startRemote workWork from homeFlexible hours$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...SeniorTemporary workImmediate startFlexible hoursShift work$148.5k - $223.9k
...duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce... ...future of Salesforce. Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely...SeniorWorldwideWeekend work$117k - $209.33k
...Job Requisition ID # 26WD99273 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products. As part of a...SeniorFor contractors$159.2k - $301.6k
...running Graphs on the cloud. In this reliability-focused role, you will own the availability... .... You'll partner with the backend engineers building these APIs to make sure the system... ...Science. ~5-10 years of experience in site reliability engineering, infrastructure,...SeniorTemporary workLocal areaWorldwide$166.9k - $225.9k
...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team... ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building...SeniorWork at officeImmediate startWorldwideMonday to FridayFlexible hours- ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and...SeniorImmediate startRemote workWorldwide
$170k - $190k
Medrio is seeking a Senior Site Reliability Engineer in San Francisco, California. The role involves maintaining and supporting the Medrio Platform, troubleshooting issues, and implementing solutions. Candidates should have experience in Kubernetes, cloud services, and...SeniorFlexible hours- A tech company specializing in AI is seeking a Site Reliability Engineer to ensure the reliability and observability of their production services. You will instrument services, develop SRE standards, and manage incident response. The ideal candidate has experience in AWS...SeniorRemote workFlexible hours
$180k - $200k
Parabola is seeking a Senior Site Reliability Engineer to join their team in San Francisco. In this role, you will monitor and improve software performance, maintain infrastructure, and collaborate with engineering teams. Candidates should have over 5 years of experience...Senior$287k
...Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was... ...expect us to hit our SLAs. What? We're looking for a Senior or Staff Site level Reliability Engineer as part of Infrastructure team to: Own...SeniorContract workWork at officeRemote work$160k - $300k
Hebbia is looking for a Site Reliability Engineer to manage and improve critical production systems. This role requires writing production-quality code and collaborating with product engineering teams to enhance system reliability. Applicants should have 5+ years in software...Senior$220k - $235k
...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move faster, insights surface instantly, and agents push work forward, all with you in control. Whether you're buying or selling...SeniorFull timeContract workWork at office$210.8k - $272.8k
Thumbtack is hiring for a Site Reliability Engineer to enhance the reliability and scalability of our services in San Francisco, CA. You'll design and support resilient systems, ensuring a smooth user experience. The ideal candidate will have extensive experience with AWS...Senior- Early Warning is seeking a Staff Site Reliability Engineer to enhance application performance and resiliency while guiding development teams. This role involves designing automation and monitoring systems, improving scalability and availability, and participating in a 2...Senior
$181k - $263k
...and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a...SeniorWork from homeFlexible hoursNight shift$175k - $250k
...00.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance... ...scalability, performance, and reliability across environments. What You’ll Do...SeniorFull timeRemote workRelocationRelocation package$210.8k - $272.8k
About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user...SeniorLocal area$200k - $260k
Senior Software Engineer, Site Reliability Engineer (SRE) Why Harvey At Harvey, we’re transforming how legal and professional services operate — not incrementally, but end‑to‑end. By combining frontier agentic AI, an enterprise‑grade platform, and deep domain expertise...SeniorRelocation package- ...systems. It's designed so Stellar's ecosystem can make a real-world, lasting impact. About the Role SDF is looking for a Senior Site Reliability Engineer to help build and operate the foundation that powers our engineering teams. You'll ensure the reliability and...Senior
- Location San Francisco, CA Employment Type Full time Department Engineering Who We Are Hyperbolic Labs is on a mission to democratize AI... ...to redefine computing. About the Role We\'re seeking a Site Reliability Engineer to ensure Hyperbolic\'s GPU marketplace and AI infrastructure...SeniorFull time
- ...respond when things go wrong, helping every organization be more reliable. We do this by building an industry‑leading incident... ...Build tools and automation to eliminate manual toil, improve engineering velocity and developer experience, and improve system reliability...SeniorHome office
- CloudDevs works with fast-moving, venture-backed startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will either be placed directly into one of our partner startups or added to our...SeniorLocal area
$170k - $190k
Position Medrio Senior Site Reliability Engineer Responsibilities Build, maintain, and support all environments which host the Medrio Platform Monitor environments for issues, configuring and building alerting/self‑healing of issues Update and maintain documentation...SeniorTemporary workFlexible hours- ...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...of accountability A strong desire to perform and grow as an engineer 5+ years of software development experience Technologies We...SeniorFlexible hours
- ...is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the...SeniorHome officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer San Francisco, CA
- site reliability engineer remote San Francisco, CA
- site reliability engineer sre San Francisco, CA
- senior trade analyst San Francisco, CA
- senior app developer San Francisco, CA
- senior customer service advisor San Francisco, CA
- senior international account manager San Francisco, CA
- senior product manager mobile San Francisco, CA
- senior magento developer San Francisco, CA
- senior quantitative risk analyst San Francisco, CA

