Senior Site Reliability Engineer -AI Infrastructure Operations
$170k - $265kNscale
About Nscale
Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work.
The Role
This is a senior SRE role for someone who sets the reliability bar and then pulls the rest of the team up
to it. You'll own the hardest problems on the platform: the automation other engineers build on, the
services that can't go down, and the design decisions that determine whether either holds up at scale.
You'll still carry a pager, but the real job is making sure it fires less, for everyone, over time.
What You'll Do
• Own reliability for critical production services end to end; set the direction, not just respond to what
breaks.
• Grow the team, not just the systems; mentor other SREs through design review, pairing, and
incident debriefs, and hold the bar that pulls everyone up to it.
• Set the standards the rest of the team works to: the SLO framework, the incident process, and the
on-call practices that keep it sustainable.
• Get in early on design reviews and architecture decisions, so reliability is built in rather than bolted
on after the first outage.
• Lead the hardest incidents and the root causes nobody else can crack; turn each one into a change
that keeps it from coming back.
• Build the tooling and automation that removes toil for the whole team, not just your own surface
area.
What You'll Bring
• 6-10 years in SRE, systems engineering, or software engineering, with real ownership of production
at scale in a data center or cloud environment.
• Strong software engineering skills (Python, Go, or similar); you build tools other engineers adopt,
not scripts that run once and rot.
• Deep command of Linux, networking, and distributed systems, plus the judgment to know where
the real failure modes hide.
• Hands-on with Kubernetes and virtualized or bare-metal environments; comfortable close to the
metal, not just the cloud console.
• Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to
get there fast.
• Reliability practices you put in place that outlasted you: SLOs, observability and alerting at scale,
incident process, on-call that people can actually live with.
• A track record as the senior voice in incidents and design reviews, trusted to make the call under
pressure.
• A habit of raising the people around you without being asked to.
Nice to Have
• Familiarity with high-performance networking (InfiniBand, RDMA).
On-Call and Pace
A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
and some weeks are busier than others. As a senior on the team, you help set how that rotation runs
and, more to the point, how we make it lighter over time. We share the load fairly, and we treat every
page as a signal worth acting on rather than just an interruption. The goal is to leave the systems
quieter than you found them, so each rotation asks less of the person carrying it. If that's the kind of
ownership you're drawn to, you'll do well here.
What We Offer
• Competitive base plus equity, reviewed every 12 months.
• Real ownership from the start, and a direct hand in how reliability works across the platform.
• Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
your day.
Salary Range
$170,000 - $265,000 USD. Actual compensation varies with skill set, experience, and location, and the
role may be eligible for bonus and equity.
Equal Opportunities Statement
At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there’s anything we can do to accommodate your specific situation, please let us know.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Salary Range
$170,000—$265,000 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
- ...RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform... ...with reliability, observability, and operational excellence at the core. You’ll... ...to build, automate, and maintain the infrastructure that powers our core platform—including...OperationsSenior
- DescriptionWe are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco... ...capacity planning.• Contribute to infrastructure and delivery workflows across AWS, Terraform...OperationsSenior
$117k - $209.33k
...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud... ...engineering, security, compliance, platform, and infrastructure teams to ensure services are reliable,...OperationsSeniorFull timeFor contractors$190k - $240k
...You will build the infrastructure that turns physical operation into reproducible training... ..., replay, and reliable access for training and... ...evaluation. This is a senior software engineering role responsible for the... ...This role is based on-site in San Francisco and...OperationsSeniorFull timeContract work$174k - $252k
...impact on hardware, network, or service operations and quality.Minimum qualifications:... ...years of experience developing large-scale infrastructure, distributed systems or networks, or... ...accessible technologies. Google's software engineers develop the next-generation...OperationsSenior- Cogent is seeking a Senior Storage Infrastructure Engineer to own how we store, protect, and operate data at scale on our Core Platform. You will build the backup, observability, and operations layer across relational, streaming, and analytical stores, making it agent-...OperationsSenior
$150k - $160k
Join Goodwin’s Global Operations Team as a Senior Data Engineer and make an impact on a global scale. In this role, you'll manage and enhance data platforms, ensuring scalability and data governance while collaborating with various teams to foster a data-driven culture....OperationsSenior- ...San Francisco is seeking a skilled Network Engineer to design and maintain high-performance network infrastructure for their programmatic platform. This role involves... ...closely with engineering teams to ensure reliable operations. Ideal candidates possess extensive...OperationsSenior
- Disney Cruise Line - The Walt Disney Company is seeking a senior big data engineer to design and maintain Identity and Device data... ...collaborate across squads including PM, design, QA, and operations to deliver reliable data solutions. Responsibilities include on-call...OperationsSenior
- ...Technologies is seeking a Manager, BI + Analytics Engineer in San Francisco to shape data modeling... ...with Product, Sales, Finance, and Operations to turn data into action. You’ll define... ..., automate workflows, and elevate data infrastructure, working closely with leaders to...OperationsSenior
$110.7k - $218.3k
...shareholder value, and optimize operational efficiency. As an Oracle Senior Consultant at Deloitte, you will... ...Information Technology, Software Engineering, or a related field.Ability to travel... ...platforms such as Oracle Cloud Infrastructure (OCI), Amazon Web Services (AWS),...OperationsSeniorLocal areaVisa sponsorship$192k - $240k
...proliferation of services and multiple operations systems, this data originates from... ...you will help build the company’s Data Infrastructure and also work across the entire data stack... ...like Kafka ~ Experience in backend engineering or full‑stack development ~ Strong...OperationsSeniorWork at officeMonday to Friday3 days per week- Mainz Brady Group seeks an experienced Cloud Data Engineer to support Investment Operations & Fund Treasury. The role emphasizes hands-on design of Snowflake, Azure, SQL and dbt data solutions for scalable, secure platforms. You'll design data warehouses, build modular...OperationsSenior
- ...video, and text data at scale. As a senior IC, you’ll own architecture and reliability of ingestion, storage, transformation, and data operations that ensure data quality across the... ...teams to translate needs into durable infrastructure, guiding architecture decisions and...OperationsSenior
- ...What the job involves As a Senior Software Engineer on the Platform - Cloud Events team, you will ensure that data captured by Hayden’s devices... ...run in the Cloud. In this role you will level up our ML operations by enabling a more efficient model improvement lifecycle....OperationsSeniorWork at office3 days per week
$174k - $252k
...years of experience developing large-scale infrastructure, distributed systems or networks, or... ...ABOUT THE JOB: Google's software engineers develop the next-generation technologies... ...impact on hardware, network, or service operations and quality. Google Cloud accelerates...OperationsSenior- ...into scalable data solutions. You will work closely with finance, supply chain, and operations to deliver reliable datasets and automations while deploying pipelines on cloud infrastructure and maintaining code quality through CI/CD practices. #J-18808-Ljbffr...OperationsSeniorLocal area
$174k - $252k
...with low-level system debugging, core operating systems concepts, and network configurations... ...a technical leadership role, guiding engineering direction or coordinating technical... ...available technologies for Beam OS infrastructure, startup stability, and networking....OperationsSeniorTemporary workLocal areaRemote work$160k - $200k
...anywhere. We design, build, and operate the world's largest... ...critical supplies quickly and reliably. Today, Zipline operates on... ...where they are needed. As a Senior Data Engineer on the Data Platform team,... ...concurrency, storage efficiency, and infrastructure or warehouse cost....OperationsSeniorFull timeLocal area$174k - $252k
...impact on hardware, network, or service operations and quality.Minimum qualifications:... ...years of experience developing large-scale infrastructure, distributed systems or networks, or... ...accessible technologies. Google's software engineers develop the next-generation...OperationsSenior$85 - $90 per hour
...posting.OverviewLABUR is partnering with a client to find a Senior Software Engineer, Platform Engineering to join a team focused on building... ...software and thrives on moving fast, delivering value, and operating with boldness, humility, and accountability. This person...OperationsSenior$200.7k - $250.9k
...fundamentally different engineering philosophies and... ...Rendezvous, Proximity Operations, and Docking (RPOD) subsystems... ...most consequential infrastructure work at Mercury.... .... As a Senior Software Engineer on... ...clean boundaries and reliable contracts. Help shape...OperationsSenior$127k - $249k
...We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands‑on technically while also mentoring a small team of SREs. The InfraSec team collaborates...SeniorLocal areaRemote workFlexible hours- ...activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in... ...through the perception team.We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models....OperationsSeniorWork at officeFlexible hours
$172k - $215k
...and build real world value.THE WORKAs a Senior Software Engineer on the RippleX Platform Engineering... ...play a critical role in ensuring the reliability, scalability, and security of XRPL infrastructure. The Platform team operates at the intersection of software engineering...OperationsSeniorFull timeWork at officeLocal area- ...Infrastructure Engineering RoleThe Infrastructure Engineering team is crucial to... ...peak performance, maximum reliability, and cost-efficiency across... ...modeling best practices in site reliability, proactive system... ...experience architecting and operating at scale within the Amazon...SeniorShift work
- ...A leading IT service provider is seeking a Senior Infrastructure & Security Engineer to support IT operations and security initiatives. This role, based in San Francisco... ...with the possibility of extension, involving both on-site and remote work options.#J-18808-Ljbffr...OperationsSeniorContract workRemote work
- ...Handshake is seeking a Senior Software Engineer on the Allocation & Onboarding team... ...to scalable platform infrastructure. You’ll collaborate with Growth, Product, Operations, ML, and Platform to translate... ...quality technical solutions and reliable systems. In this role you...OperationsSenior
$80 per hour
Infrastructure Site Reliability Engineer (Local only) Direct message the job poster from Maxonic Inc. Job Description... ...core SRE principles to automate operational tasks, monitor system health, and... ...execution‑focused, supporting the senior team in ensuring our services are...OperationsFull timeContract workFor contractorsLocal area2 days per week- ...TeamThe Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the... ...service, cluster scheduler, and reliability tooling that powers the company's data... ...end-user tooling.About the RoleAs a Senior Software Engineer on Spark Platform, you will set the technical...OperationsSeniorHourly payWork at officeLocal areaRemote workRelocationFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer -AI Infrastructure Operations. Be the first to apply!
- site reliability engineer San Francisco, CA
- site reliability engineer remote San Francisco, CA
- site reliability engineer sre San Francisco, CA
- infrastructure engineering manager San Francisco, CA
- data infrastructure engineer San Francisco, CA
- security infrastructure engineer San Francisco, CA
- senior infrastructure engineer San Francisco, CA
- remote infrastructure engineer San Francisco, CA
- lead infrastructure engineer San Francisco, CA
- entry level infrastructure engineer San Francisco, CA



