Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer -AI Infrastructure Operations

$170k - $265k
Full-time

Nscale

About Nscale
Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work.


The Role
This is a senior SRE role for someone who sets the reliability bar and then pulls the rest of the team up
to it. You'll own the hardest problems on the platform: the automation other engineers build on, the
services that can't go down, and the design decisions that determine whether either holds up at scale.
You'll still carry a pager, but the real job is making sure it fires less, for everyone, over time.


What You'll Do
• Own reliability for critical production services end to end; set the direction, not just respond to what
breaks.
• Grow the team, not just the systems; mentor other SREs through design review, pairing, and
incident debriefs, and hold the bar that pulls everyone up to it.
• Set the standards the rest of the team works to: the SLO framework, the incident process, and the
on-call practices that keep it sustainable.
• Get in early on design reviews and architecture decisions, so reliability is built in rather than bolted
on after the first outage.
• Lead the hardest incidents and the root causes nobody else can crack; turn each one into a change
that keeps it from coming back.
• Build the tooling and automation that removes toil for the whole team, not just your own surface
area.


What You'll Bring
• 6-10 years in SRE, systems engineering, or software engineering, with real ownership of production
at scale in a data center or cloud environment.
• Strong software engineering skills (Python, Go, or similar); you build tools other engineers adopt,
not scripts that run once and rot.
• Deep command of Linux, networking, and distributed systems, plus the judgment to know where
the real failure modes hide.
• Hands-on with Kubernetes and virtualized or bare-metal environments; comfortable close to the
metal, not just the cloud console.
• Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to
get there fast.
• Reliability practices you put in place that outlasted you: SLOs, observability and alerting at scale,
incident process, on-call that people can actually live with.
• A track record as the senior voice in incidents and design reviews, trusted to make the call under
pressure.
• A habit of raising the people around you without being asked to.
Nice to Have
• Familiarity with high-performance networking (InfiniBand, RDMA).


On-Call and Pace
A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
and some weeks are busier than others. As a senior on the team, you help set how that rotation runs
and, more to the point, how we make it lighter over time. We share the load fairly, and we treat every
page as a signal worth acting on rather than just an interruption. The goal is to leave the systems
quieter than you found them, so each rotation asks less of the person carrying it. If that's the kind of
ownership you're drawn to, you'll do well here.


What We Offer
• Competitive base plus equity, reviewed every 12 months.
• Real ownership from the start, and a direct hand in how reliability works across the platform.
• Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
your day.


Salary Range
$170,000 - $265,000 USD. Actual compensation varies with skill set, experience, and location, and the
role may be eligible for bonus and equity.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$170,000—$265,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer -AI Infrastructure Operations in San Francisco, CA vacancy
  •  ...RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform...  ...with reliability, observability, and operational excellence at the core. You’ll...  ...to build, automate, and maintain the infrastructure that powers our core platform—including... 
    Operations
    Senior

    Alembic

    San Francisco, CA
    2 days ago
  • DescriptionWe are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco...  ...capacity planning.• Contribute to infrastructure and delivery workflows across AWS, Terraform... 
    Operations
    Senior

    Robert Half

    San Francisco, CA
    3 days ago
  • $117k - $209.33k

     ...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud...  ...engineering, security, compliance, platform, and infrastructure teams to ensure services are reliable,... 
    Operations
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  • $190k - $240k

     ...You will build the infrastructure that turns physical operation into reproducible training...  ..., replay, and reliable access for training and...  ...evaluation. This is a senior software engineering role responsible for the...  ...This role is based on-site in San Francisco and... 
    Operations
    Senior
    Full time
    Contract work

    GRAM

    San Francisco, CA
    10 hours ago
  • $174k - $252k

     ...impact on hardware, network, or service operations and quality.Minimum qualifications:...  ...years of experience developing large-scale infrastructure, distributed systems or networks, or...  ...accessible technologies. Google's software engineers develop the next-generation... 
    Operations
    Senior

    Google

    San Francisco, CA
    2 days ago
  • Cogent is seeking a Senior Storage Infrastructure Engineer to own how we store, protect, and operate data at scale on our Core Platform. You will build the backup, observability, and operations layer across relational, streaming, and analytical stores, making it agent-... 
    Operations
    Senior

    Doist

    San Francisco, CA
    12 hours ago
  • $150k - $160k

    Join Goodwin’s Global Operations Team as a Senior Data Engineer and make an impact on a global scale. In this role, you'll manage and enhance data platforms, ensuring scalability and data governance while collaborating with various teams to foster a data-driven culture.... 
    Operations
    Senior

    Goodwin Procter

    San Francisco, CA
    2 days ago
  •  ...San Francisco is seeking a skilled Network Engineer to design and maintain high-performance network infrastructure for their programmatic platform. This role involves...  ...closely with engineering teams to ensure reliable operations. Ideal candidates possess extensive... 
    Operations
    Senior

    Rzr Co

    San Francisco, CA
    1 day ago
  • Disney Cruise Line - The Walt Disney Company is seeking a senior big data engineer to design and maintain Identity and Device data...  ...collaborate across squads including PM, design, QA, and operations to deliver reliable data solutions. Responsibilities include on-call... 
    Operations
    Senior

    Disney Cruise Line - The Walt Disney Company

    San Francisco, CA
    12 hours ago
  •  ...Technologies is seeking a Manager, BI + Analytics Engineer in San Francisco to shape data modeling...  ...with Product, Sales, Finance, and Operations to turn data into action. You’ll define...  ..., automate workflows, and elevate data infrastructure, working closely with leaders to... 
    Operations
    Senior

    Doist

    San Francisco, CA
    2 days ago
  • $110.7k - $218.3k

     ...shareholder value, and optimize operational efficiency. As an Oracle Senior Consultant at Deloitte, you will...  ...Information Technology, Software Engineering, or a related field.Ability to travel...  ...platforms such as Oracle Cloud Infrastructure (OCI), Amazon Web Services (AWS),... 
    Operations
    Senior
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    2 days ago
  • $192k - $240k

     ...proliferation of services and multiple operations systems, this data originates from...  ...you will help build the company’s Data Infrastructure and also work across the entire data stack...  ...like Kafka ~ Experience in backend engineering or full‑stack development ~ Strong... 
    Operations
    Senior
    Work at office
    Monday to Friday
    3 days per week

    United States Digital Space LLC

    San Francisco, CA
    16 hours ago
  • Mainz Brady Group seeks an experienced Cloud Data Engineer to support Investment Operations & Fund Treasury. The role emphasizes hands-on design of Snowflake, Azure, SQL and dbt data solutions for scalable, secure platforms. You'll design data warehouses, build modular... 
    Operations
    Senior

    Mainz Brady Group

    San Francisco, CA
    4 days ago
  •  ...video, and text data at scale. As a senior IC, you’ll own architecture and reliability of ingestion, storage, transformation, and data operations that ensure data quality across the...  ...teams to translate needs into durable infrastructure, guiding architecture decisions and... 
    Operations
    Senior

    Walden Robotics

    San Francisco, CA
    2 days ago
  •  ...What the job involves As a Senior Software Engineer on the Platform - Cloud Events team, you will ensure that data captured by Hayden’s devices...  ...run in the Cloud. In this role you will level up our ML operations by enabling a more efficient model improvement lifecycle.... 
    Operations
    Senior
    Work at office
    3 days per week

    Hayden AI

    San Francisco, CA
    16 hours ago
  • $174k - $252k

     ...years of experience developing large-scale infrastructure, distributed systems or networks, or...  ...ABOUT THE JOB: Google's software engineers develop the next-generation technologies...  ...impact on hardware, network, or service operations and quality. Google Cloud accelerates... 
    Operations
    Senior

    Socket.dev

    San Francisco, CA
    16 hours ago
  •  ...into scalable data solutions. You will work closely with finance, supply chain, and operations to deliver reliable datasets and automations while deploying pipelines on cloud infrastructure and maintaining code quality through CI/CD practices. #J-18808-Ljbffr... 
    Operations
    Senior
    Local area

    NextGenEnergyJobs

    San Francisco, CA
    2 days ago
  • $174k - $252k

     ...with low-level system debugging, core operating systems concepts, and network configurations...  ...a technical leadership role, guiding engineering direction or coordinating technical...  ...available technologies for Beam OS infrastructure, startup stability, and networking.... 
    Operations
    Senior
    Temporary work
    Local area
    Remote work

    Google

    San Francisco, CA
    16 hours ago
  • $160k - $200k

     ...anywhere. We design, build, and operate the world's largest...  ...critical supplies quickly and reliably. Today, Zipline operates on...  ...where they are needed. As a Senior Data Engineer on the Data Platform team,...  ...concurrency, storage efficiency, and infrastructure or warehouse cost.... 
    Operations
    Senior
    Full time
    Local area

    Zipline

    San Francisco, CA
    16 days ago
  • $174k - $252k

     ...impact on hardware, network, or service operations and quality.Minimum qualifications:...  ...years of experience developing large-scale infrastructure, distributed systems or networks, or...  ...accessible technologies. Google's software engineers develop the next-generation... 
    Operations
    Senior

    Google

    San Bruno, CA
    1 day ago
  • $85 - $90 per hour

     ...posting.OverviewLABUR is partnering with a client to find a Senior Software Engineer, Platform Engineering to join a team focused on building...  ...software and thrives on moving fast, delivering value, and operating with boldness, humility, and accountability. This person... 
    Operations
    Senior

    Labur

    San Francisco, CA
    3 days ago
  • $200.7k - $250.9k

     ...fundamentally different engineering philosophies and...  ...Rendezvous, Proximity Operations, and Docking (RPOD) subsystems...  ...most consequential infrastructure work at Mercury....  .... As a Senior Software Engineer on...  ...clean boundaries and reliable contracts. Help shape... 
    Operations
    Senior

    Mercury

    San Francisco, CA
    4 days ago
  • $127k - $249k

     ...We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands‑on technically while also mentoring a small team of SREs. The InfraSec team collaborates... 
    Senior
    Local area
    Remote work
    Flexible hours

    Insider, Inc.

    San Francisco, CA
    16 hours ago
  •  ...activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in...  ...through the perception team.We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models.... 
    Operations
    Senior
    Work at office
    Flexible hours

    Voxel

    San Francisco, CA
    3 days ago
  • $172k - $215k

     ...and build real world value.THE WORKAs a Senior Software Engineer on the RippleX Platform Engineering...  ...play a critical role in ensuring the reliability, scalability, and security of XRPL infrastructure. The Platform team operates at the intersection of software engineering... 
    Operations
    Senior
    Full time
    Work at office
    Local area

    Ripple

    San Francisco, CA
    1 day ago
  •  ...Infrastructure Engineering RoleThe Infrastructure Engineering team is crucial to...  ...peak performance, maximum reliability, and cost-efficiency across...  ...modeling best practices in site reliability, proactive system...  ...experience architecting and operating at scale within the Amazon... 
    Senior
    Shift work

    Hayden AI

    San Francisco, CA
    3 days ago
  •  ...A leading IT service provider is seeking a Senior Infrastructure & Security Engineer to support IT operations and security initiatives. This role, based in San Francisco...  ...with the possibility of extension, involving both on-site and remote work options.#J-18808-Ljbffr... 
    Operations
    Senior
    Contract work
    Remote work

    CDW

    San Francisco, CA
    16 hours ago
  •  ...Handshake is seeking a Senior Software Engineer on the Allocation & Onboarding team...  ...to scalable platform infrastructure. You’ll collaborate with Growth, Product, Operations, ML, and Platform to translate...  ...quality technical solutions and reliable systems. In this role you... 
    Operations
    Senior

    Apply

    San Francisco, CA
    16 hours ago
  • $80 per hour

    Infrastructure Site Reliability Engineer (Local only) Direct message the job poster from Maxonic Inc. Job Description...  ...core SRE principles to automate operational tasks, monitor system health, and...  ...execution‑focused, supporting the senior team in ensuring our services are... 
    Operations
    Full time
    Contract work
    For contractors
    Local area
    2 days per week

    Maxonic Inc.

    San Francisco, CA
    3 days ago
  •  ...TeamThe Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the...  ...service, cluster scheduler, and reliability tooling that powers the company's data...  ...end-user tooling.About the RoleAs a Senior Software Engineer on Spark Platform, you will set the technical... 
    Operations
    Senior
    Hourly pay
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours

    Doordash

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer -AI Infrastructure Operations. Be the first to apply!