Senior Infrastructure/Site Reliability Engineer On-Call
Jobleads-US
ABOUT THE ROLE
This is a senior on-call SRE role at an early-stage AI infrastructure company, where you will be the technical expert enterprise clients depend on when critical systems fail. You will own incident response across Kubernetes clusters, Ceph storage, and bare metal servers, keeping high-value AI workloads running at all times. Your calm judgment and deep distributed systems expertise will have a direct and immediate impact on client operations.
WHAT YOU'LL DO
- Respond to and resolve production incidents across client infrastructure spanning Kubernetes, Ceph, and bare metal environments.
- Troubleshoot complex distributed systems problems including pod scheduling failures, CNI networking issues, storage performance degradation, and hardware faults.
- Handle escalations requiring deep expertise in etcd clusters, Ceph RGW authentication, Cilium networking, and bare metal load balancers.
- Communicate directly with enterprise clients during incidents, providing clear status updates and resolution timelines.
- Participate in a follow-the-sun on-call rotation with engineers across multiple time zones.
- Document incidents thoroughly and improve runbooks based on recurring patterns.
- Collaborate with the infrastructure team on long-term reliability improvements and automation to reduce incident frequency.
WHAT WE'RE LOOKING FOR
- 5 or more years of production experience with Kubernetes in enterprise environments, including cluster operations, bare metal troubleshooting, admission controllers, and control plane architecture.
- Deep production experience with distributed storage systems, particularly Ceph, or equivalent platforms such as Weka or VAST.
- Production experience with at least one CNI plugin, preferably Cilium or Calico.
- Strong modern Linux systems administration skills and comfort with bare metal infrastructure, IPMI, hardware troubleshooting, and networking.
- Demonstrated ability to systematically diagnose and resolve complex distributed systems issues under pressure during live outages.
- Production experience with etcd cluster management, including backup and restore procedures.
- Experience with GPU infrastructure for AI/ML workloads, including the NVIDIA Kubernetes operator.
- Familiarity with infrastructure-as-code tools such as Ansible, Kubespray, or similar orchestration frameworks.
- Clear, composed communication with both technical and non-technical audiences during incidents.
- Background in AI inference or training infrastructure is a strong plus.
LOCATION
On-site in San Francisco, CA. This role involves a follow-the-sun on-call rotation and requires timezone flexibility. Visa sponsorship is not available.
#J-18808-Ljbffr Jobleads-US$152.5k - $205k
...applications, and programmable blockchain infrastructure. Circle’s platform includes the... ...What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll... ...reliability by participating in on-call, leading incident response, performing...SeniorFlexible hours- ...in 2015 to build the infrastructure global commerce runs... ...curiosity, and make calls from first principles... ...full.About the teamThe Engineering team at Airwallex is... ...together to build scalable, reliable, and secure products... ....What you’ll doAs a Senior Site Reliability Engineer,...SeniorTemporary workLocal area
$210k - $240k
...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided... ...build, automate, and maintain the infrastructure that powers our core platform—including... ...processes (SLOs, runbooks, on‑call rotations) Collaborate across engineering...SeniorFull time- ...RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with... ...to build, automate, and maintain the infrastructure that powers our core platform—... ...response processes (SLOs, runbooks, on-call rotations)Collaborate across engineering...Senior
$250k
...Join a rapidly scaling AI cloud infrastructure provider building a next-generation... .... The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC... ...Participate in an on-call rotation supporting mission-critical...SeniorFull timeRemote work$117k - $209.33k
...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build... ...security, compliance, platform, and infrastructure teams to ensure services are... ...applicableParticipate in a 24x7 on-call rotation for production servicesFunction...SeniorFull timeFor contractors- ...Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates... ...engineering and applies them to infrastructure and operations problems. The main goals... ...performance; Participate in on-call rotation to provide 24/7 support for...SeniorImmediate startRemote workWorldwide
$148.5k - $223.9k
...Details Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in SanFrancisco.... ...with counterparts in the Infrastructure and R&D organizations, this organization... ...-the-sun model with weekend on‑call, the Site Reliability team keeps...SeniorWorldwideWeekend work$153k - $191.3k
...data processing, and software engineering, our office is a truly... ...Planet's Direct Access Service Infrastructure team, directly contributing... ...environments, to guarantee the reliability, scalability, and... ...tests Participate in on-call rotations to ensure operational...SeniorFull timeTemporary workFor contractorsWork at officeLocal areaRemote workHome office3 days per week$220k - $235k
...strategic, high‑output Staff/Senior Staff SRE to define... ...and champion engineering excellence across Ironclad... ...direction for the Site Reliability Engineering team and... ...systems Be on an on‑call rotation to respond... ...Ability to build resilient infrastructure ~ Modern GitOps –...SeniorFull timeWork at office- Senior Software Engineer At Commure, we're building the AI Operating System... ...Software Engineer on the Infrastructure team to own the foundational... ...'ll make architectural calls, write the code that... ...infrastructure, platform, or site reliability engineering roles. Experience...SeniorLocal areaImmediate start
- Software Engineer Voxel's perception system is the technical core of everything we ship.... ...strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision... ..., write code, make architecture calls, and partner closely with applied CV, ML...SeniorWork at officeFlexible hours
- Platform Engineer HUD is building infrastructure to create RL training data and evals for frontier AI agents,... ...Platform Engineer who can own the reliability, scale, performance, and developer... ...logs, traces, SLOs, runbooks, and on-call workflows so failures are detected,...SeniorFull timeWork at officeRemote workRelocationVisa sponsorship
$200k - $260k
...team of the world's best engineers and operators. If you are... ...Role We are looking for a Senior Software Engineer to join our Infrastructure Team , focused on building secure, reliable, and scalable... ...response processes, including on-call practices, runbooks, postmortems...SeniorLocal areaFlexible hours$232k - $319k
...secures AI by building the trusted, neutral infrastructure that enables organizations to safely... ...the service with great people and reliable, cost-effective, and efficient infrastructure... ...the velocity of SRE and product engineering by developing robust platforms, powerful...SeniorPermanent employmentLocal areaWorldwideFlexible hours- ...US Corp. is seeking a Lead Site Reliability Engineer to spearhead our mission of delivering highly available and performant systems. With an... ...be responsible for designing and implementing automated infrastructure using Terraform, managing containerized workloads within...Senior
$160k - $195k
...vertically integrated AI infrastructure company built from... ...We’re seeking a Senior Cloud Infrastructure Engineer to own the design, implementation... ...IT, Security, and site-specific operations to keep systems reliable, secure, and ready... ...; participate in on-call rotation as the...SeniorTemporary workWork at office$180k - $240k
Senior Cloud Infrastructure Engineer Loft Orbital is revolutionizing access to space by building reliable, shareable satellites that drastically reduce the... ...spacecraft control. Yes, we call it SatDevOps, and we'll... ...with financial support Off-sites and many social events...SeniorTemporary workWork at officeRelocation packageFlexible hours$175k - $250k
...000.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance... ...scalability, performance, and reliability across environments. What You’ll Do Design...SeniorFull timeRemote workRelocationRelocation package$164k - $205k
...maintain production systems Build and operate cloud infrastructure on AWS, using Terraform to codify and version-control... ...alerting and observability systems Collaborate with engineering teams to embed reliability into the development lifecycle, shifting left on...SeniorWork experience placementSummer holidayLive outWork at officeLocal areaFlexible hoursShift work2 days per week- Role SummaryWe have an opening for a Senior Software Infrastructure Engineer in our Infrastructure team focused on driving technical strategy, alignment... ...team focused on improving the scalability and reliability of Temporal’s core infrastructure. In this role, you will...Senior
$189k - $283.6k
..., you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics-driven, systems-oriented, and focused... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience...SeniorFull timeRelocation packageFlexible hoursShift work$181k - $263k
## Senior Staff Site Reliability EngineerApplylocations: San Franciscotime type: Full timeposted on: Posted... ...for a Senior Staff Site Reliability Engineer who will set the technical direction... ...across LiveRamp's global infrastructure. This is a senior individual contributor...SeniorWork from homeFlexible hoursNight shift- ...Required skills ~ Engineering ~322 Infra About Anyscale: At Anyscale... ...: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide... ...architecture discussions Provide on-call support, working closely with...Remote jobFull time
$131.75k - $178.25k
...practices and entrepreneurial spirit allow exceptional opportunities for professional achievement and career growth. The Senior Infrastructure DevOps Engineer, under the direction of the Senior Manager of Infrastructure, will collaborate with our internal IT teams to design,...SeniorFull timeWork experience placementRemote workWorldwide- ...AfterQuery is seeking a Senior Software Engineer - Infrastructure in San Francisco to design and build core infrastructure for data generation, evaluation... ...-scale experiments and ensure they are scalable and reliable. You will collaborate with the founding team to set...Senior
- ...Grow Therapy in San Francisco is seeking a Senior AI Enablement Engineer to define how AI transforms operations across the organization. You will design and build foundational AI infrastructure that enhances efficiency. Responsibilities include implementing AI systems...SeniorFlexible hours3 days per week
- ...Engg, a San Francisco–based AI infrastructure company, seeks a senior on-call SRE to own incident response for Kubernetes, Ceph storage, and bare metal... ...restores, and help reduce incident frequency through reliability improvements and IaC tooling. #J-18808-Ljbffr...Senior
- ...Crusoe is seeking a Senior Software Engineer to architect, design, and develop Cloud Infrastructure management systems for Crusoe Cloud. You will deliver end-to-end workflows... ...You'll collaborate across teams to build reliable, secure, and cost-efficient cloud platforms,...Senior
$183k - $220k
...all of us. ???? The Role As a senior member of our small but mighty engineering team, you'll work closely with... ...evolution of our backend tech stack and infrastructure. Whether iterating on core... ...our customer base and perform reliably for our users. Product development...SeniorFull timeFor contractorsWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Infrastructure/Site Reliability Engineer On-Call. Be the first to apply!
- lead infrastructure engineer San Francisco, CA
- security infrastructure engineer San Francisco, CA
- principal infrastructure engineer San Francisco, CA
- entry level infrastructure engineer San Francisco, CA
- data infrastructure engineer San Francisco, CA
- infrastructure engineer San Francisco, CA
- remote infrastructure engineer San Francisco, CA
- senior infrastructure engineer San Francisco, CA
- infrastructure developer San Francisco, CA
- infrastructure engineering manager San Francisco, CA




