Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$170k - $290k

Luma AI

About Luma AI

Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. This requires a massive, reliable, and performant GPU infrastructure that pushes the boundaries of scale. Our SRE team is the foundation of our research and product velocity, responsible for the thousands of NVIDIA and AMD GPUs across multiple providers that power our work.

Where You Come In

We are looking for a hands-on, first-principles engineer who is fluent in Linux, comfortable operating close to the metal, and capable of architecting systems for the next generation of AI infrastructure.

You will build, maintain, and scale Luma's infrastructure across on-prem and multi-vendor clouds (AWS & OCI), serving as the bridge between hardware vendors, cloud providers, and our research teams.

What You'll Do

  • Architect for Reliability & Scale: Participate in critical re-architecture sessions to redesign our systems for higher efficiency and scale. You won't just maintain existing clusters; you will help define how our next-generation infrastructure operates.
  • Own Multi-Cloud GPU Clusters: Take end-to-end ownership of our production clusters for training and inference across AWS and OCI, ensuring high availability and peak performance.
  • Drive Security & Compliance: Assist in achieving and maintaining security certifications (SOC 2 Type 1 & 2, ISO standards) by implementing robust infrastructure security practices in a fast-moving AI startup environment.
  • Deep Linux Performance Tuning: Use your mastery of Linux systems to troubleshoot and optimize performance at the OS and kernel level.
  • Build Robust Automation: Write high-quality tools and automation in Python, Go, or Bash to manage, monitor, and heal our infrastructure without relying on heavy operational toil.
  • Debug Complex Hardware/Software Failures: Serve as the final escalation point for the most challenging GPU, networking (InfiniBand/RDMA), and system-level issues, often collaborating directly with hardware vendors like NVIDIA.
Who You Are
  • 5+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep Linux Mastery: You possess deep, hands-on expertise in Linux, containerized systems, and debugging low-level system performance.
  • Expert in Technologies: You have working experiencewith Terraform, Airflow, and Ray
  • Cloud Infrastructure Expert: You have strong experience with providers like AWS or OCI.
  • Tenacious Troubleshooter: You thrive on solving complex, low-level problems where hardware and software intersect.
  • Startup DNA: You are energetic and thrive in a less structured, fast-paced environment.
  • Security-Minded: You possess a working knowledge of security best practices and familiarity with compliance frameworks, such as SOC 2 and ISO.
  • Expert in High-Performance Networking: You have practical experience with InfiniBand, RDMA, or RoCE and understand how to optimize throughput for massive distributed training jobs.
What Sets You Apart (Bonus Points)
  • Deep expertise with GPU tooling for NVIDIA and AMD GPUs like DCGM or ROCm.
  • Experience managing large-scale GPU clusters for AI/ML workloads (training or inference).
  • Familiarity with job management systems based on Kubernetes or orchestration frameworks like Ray.
  • Deep expertise in Data Pipeline and Infrastructure

Compensation

The base pay range for this role is $170,000 - $290,000 per year.

About Luma

Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world.

We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  •  ...professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of...  ...and position yourself among the top echelon in site reliability.  As a Senior Lead Site Reliability Engineer at JPMorgan Chase within... 
    Senior

    J.P. Morgan

    Palo Alto, CA
    7 days ago
  •  ...Guide and shape the future of technology at a globally recognized firm, driven by pride in ownership. As a Senior Manager of Site Reliability Engineering at JPMorgan Chase within the Corporate Investment Bank, Markets team, you are the non-functional requirement owner... 
    Senior
    Bank staff
    Shift work

    J.P. Morgan

    New York, NY
    19 days ago
  •  ...looking for people just like you. Join our team and help us develop game-changing, high-quality solutions. As a Senior Lead Site Reliability Engineer at JPMorganChase within the Core Engineering Solutions team of Consumer and Community Banking , you are an integral... 
    Senior

    J.P. Morgan

    Jersey City, NJ
    4 days ago
  • Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments and predicting.
    Senior

    Software Technology Inc

    Washington DC
    3 days ago
  • $112.92k - $180k

     ...Senior Site Reliability Engineer Wal-Mart Dallas, TX Senior Site Reliability Engineer professional opening available at Wal-Martin Dallas, TX. Master's or equiv in CS, Comp Engg, Comp Info. Systs,SW Engg, Electrical Engg, or rel. area & 1 yr of exp in site reliability... 
    Senior
    Temporary work

    Walmart

    Dallas, TX
    17 hours ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Senior
    Local area

    E-Solutions

    New York, NY
    4 days ago
  •  ...infra has to match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi...  ...-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions... 
    Senior
    Remote work

    Satsuma

    United States
    1 day ago
  • $400k

     ...Salary : Up to $400,000 Total Compensation Senior Site Reliability Engineer We are working with a leading trading technology firm building a high-performance infrastructure engineering team focused on reliability, automation, and large-scale platform resilience. This... 
    Senior

    Hamilton Barnes ?

    New York, NY
    3 days ago
  •  ...New York City or Chicago (Hybrid) A technology-driven investment firm is expanding its Platform Engineering organization and is seeking an experienced Senior Site Reliability Engineer to help shape reliability practices across its infrastructure and production... 
    Senior

    Mission Staffing

    New York, NY
    3 days ago
  •  ...Role : Senior Site Reliability Engineer Location : West Lake, CA or Carrolton, TX (ONSITE) FTE ONLY Job Description Must Have Technical/Functional Skills o 5-7 years of professional experience in a Site Reliability, DevOps, or Systems Engineering... 
    Senior
    Permanent employment

    AceStack LLC

    Carrollton, TX
    1 day ago
  •  ...A leading quantitative trading firm is seeking a Senior Site Reliability Engineer to build and evolve the reliability, observability, and automation capabilities powering a highly performance-sensitive trading environment. Working at the intersection of software and infrastructure... 
    Senior

    Acquire Me

    New York, NY
    3 days ago
  • $104.9k - $174.7k

     ...immediately hire a highly skilled and proactive Senior SRE to join our dynamic team. You will...  ...fault‑tolerant systems within agreed reliability objectives, whilst enabling the fast...  ...skills. About team; This diverse team of Engineers in assisting multiple product teams as we... 
    Senior
    Local area
    Immediate start

    RELX

    Ewing, NJ
    3 days ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment... 
    Senior
    Work experience placement

    Samprasoft

    Washington DC
    4 days ago
  •  ...Job Title: Senior Site Reliability Engineer Duration: 18 months (possibility to extend or convert to FTE) Location: Charlotte, NC - Hybrid Role (3 days onsite in a week) Interview process: 2 rounds #1-hour virtual panel #1 hour on site technical panel... 
    Senior
    Shift work
    3 days per week

    Veracity

    Charlotte, NC
    17 hours ago
  •  ...Seeking a full-time Senior Site Reliability Engineer with expertise in C# and .NET to ensure production reliability for customer-facing platforms and weather data services in a remote setting, focusing on high availability, incident response, and operational excellence... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  •  ...Responsibilities Improve the reliability of mission-critical solutions, applications, and platforms Software development for enterprises...  ...Windows and Linux Years of Experience: 5 Years of Software Engineering Seniority level Mid-Senior level Employment type Full-time Job... 
    Senior
    Full time
    Work experience placement

    InterEx Group

    New York, NY
    3 days ago
  •  ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient... 
    Senior

    TechChain Talent

    New York, NY
    3 days ago
  •  ...infrastructure that support Akamai’s Compute products and services. The Senior Engineer creates solutions to improve automation and efficiency for...  ...deployment, monitoring, and resolving incidents. Focus on reliability, scalability, and efficiency through automation and resource... 
    Senior
    Flexible hours

    Akamai

    Poland, NY
    3 days ago
  •  ...founders with PhDs in AI, Math, and Computer Science - is poised to redefine computing. About the Role We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and... 
    Senior

    Hyperbolic Labs

    San Francisco, CA
    3 days ago
  •  ...The Role We're looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of the production systems that power Nectar's platform. We run high-volume data ingestion pipelines and real-time AI agents on top of a fast... 
    Senior
    Remote work

    Nectar Social

    United States
    2 days ago
  •  ...To support the Department of Veterans Affairs' enterprise healthcare platforms, the remote Senior Site Reliability Engineer will enhance reliability engineering, cloud operations, and automation while collaborating with cross-functional teams to improve service delivery... 
    Senior
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  •  ...degree in Computer Science, Information Systems Management, Engineering or related field or equivalent experience. 3+ years of experience...  ...teams can leverage to accelerate innovation in the areas of reliability, scalability and velocity. Design and maintain software... 
    Senior

    Software Technology Inc

    Lancaster, CA
    2 days ago
  •  ...Senior Site Reliability Engineer Partner with software developers, platform engineers, and IT staff to improve system design, operability, deployment safety, and production support readiness. Define and maintain operational standards, runbooks, support procedures... 
    Senior
    Work at office
    Remote work

    ARA Brand

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer We are looking for a Senior Reliability Engineer to join our Platform team. In this position, you will be responsible for maintaining, designing, implementing and upgrading our cloud infrastructure to support our microservices platforms... 
    Senior
    Temporary work
    Flexible hours

    1872 Consulting

    Chicago, IL
    2 days ago
  •  ...Senior Site Reliability Engineer As a Senior Site Reliability Engineer at Blitzy's Cambridge headquarters, you will be the backbone of our platform's reliability, scalability, and operational excellence. You'll work at the intersection of software engineering and infrastructure... 
    Senior

    Blitzy

    Cambridge, MA
    2 days ago
  •  ...FL or Pittsburgh, PA Our client seeks a Senior SRE Developer to lead technical...  ...deployment automation and infrastructure reliability. The role emphasizes building software...  ...improvements. Requirements Proven software engineering background with an SRE mindset and experience... 
    Senior
    Contract work

    Eliassen Group

    Florida, NY
    3 days ago
  • $100 per hour

     ...- join early. As our Senior SRE, you'll be in charge of...  ...create impact Improve reliability of our systems Build & maintain...  ...frameworks and solutions to engineering problems Fast-moving: you...  ...~401k benefits ~ On-site team culture - high collaboration... 
    Senior
    Immediate start
    Weekend work

    DualEntry

    New York, NY
    23 minutes ago
  •  ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...culture of innovation and continuous improvement.The Senior Site Reliability Engineer acts as an advanced senior... 
    Senior
    Work at office
    Shift work

    Bank of America

    Chandler, MN
    3 days ago
  • $100k - $115k

     ...Software And Systems Engineer Why you will love this job: Great opportunity to use software and systems engineering to build a...  ...and design automation strategy across platform Troubleshoot site down issues and respond to emergency outages Work with engineering... 
    Senior
    Remote work
    Work from home

    MRINetwork

    United States
    4 days ago
  •  ...Senior Site Reliability Engineer The primary responsibility of the Senior Site Reliability Engineer (SRE) to lead reliability engineering initiatives across our Azure estate and Command Center operations. This role focuses on scripting, automation, and observability... 
    Senior
    Shift work
    Night shift

    Las Vegas Sands Corp.

    Dallas, TX
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!