Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Infrastructure/Site Reliability Engineer On-Call

Jobleads-US

ABOUT THE ROLE

This is a senior on-call SRE role at an early-stage AI infrastructure company, where you will be the technical expert enterprise clients depend on when critical systems fail. You will own incident response across Kubernetes clusters, Ceph storage, and bare metal servers, keeping high-value AI workloads running at all times. Your calm judgment and deep distributed systems expertise will have a direct and immediate impact on client operations.

WHAT YOU'LL DO

  • Respond to and resolve production incidents across client infrastructure spanning Kubernetes, Ceph, and bare metal environments.
  • Troubleshoot complex distributed systems problems including pod scheduling failures, CNI networking issues, storage performance degradation, and hardware faults.
  • Handle escalations requiring deep expertise in etcd clusters, Ceph RGW authentication, Cilium networking, and bare metal load balancers.
  • Communicate directly with enterprise clients during incidents, providing clear status updates and resolution timelines.
  • Participate in a follow-the-sun on-call rotation with engineers across multiple time zones.
  • Document incidents thoroughly and improve runbooks based on recurring patterns.
  • Collaborate with the infrastructure team on long-term reliability improvements and automation to reduce incident frequency.

WHAT WE'RE LOOKING FOR

  • 5 or more years of production experience with Kubernetes in enterprise environments, including cluster operations, bare metal troubleshooting, admission controllers, and control plane architecture.
  • Deep production experience with distributed storage systems, particularly Ceph, or equivalent platforms such as Weka or VAST.
  • Production experience with at least one CNI plugin, preferably Cilium or Calico.
  • Strong modern Linux systems administration skills and comfort with bare metal infrastructure, IPMI, hardware troubleshooting, and networking.
  • Demonstrated ability to systematically diagnose and resolve complex distributed systems issues under pressure during live outages.
  • Production experience with etcd cluster management, including backup and restore procedures.
  • Experience with GPU infrastructure for AI/ML workloads, including the NVIDIA Kubernetes operator.
  • Familiarity with infrastructure-as-code tools such as Ansible, Kubespray, or similar orchestration frameworks.
  • Clear, composed communication with both technical and non-technical audiences during incidents.
  • Background in AI inference or training infrastructure is a strong plus.

LOCATION

On-site in San Francisco, CA. This role involves a follow-the-sun on-call rotation and requires timezone flexibility. Visa sponsorship is not available.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Infrastructure/Site Reliability Engineer On-Call in San Francisco, CA vacancy
  • $152.5k - $205k

     ...applications, and programmable blockchain infrastructure. Circle’s platform includes the...  ...What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll...  ...reliability by participating in on-call, leading incident response, performing... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  •  ...in 2015 to build the infrastructure global commerce runs...  ...curiosity, and make calls from first principles...  ...full.About the teamThe Engineering team at Airwallex is...  ...together to build scalable, reliable, and secure products...  ....What you’ll doAs a Senior Site Reliability Engineer,... 
    Senior
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    2 days ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided...  ...build, automate, and maintain the infrastructure that powers our core platform—including...  ...processes (SLOs, runbooks, on‑call rotations) Collaborate across engineering... 
    Senior
    Full time

    Alembic Technologies

    San Francisco, CA
    3 days ago
  •  ...RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with...  ...to build, automate, and maintain the infrastructure that powers our core platform—...  ...response processes (SLOs, runbooks, on-call rotations)Collaborate across engineering... 
    Senior

    Alembic Limited

    San Francisco, CA
    1 day ago
  • $250k

     ...Join a rapidly scaling AI cloud infrastructure provider building a next-generation...  .... The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC...  ...Participate in an on-call rotation supporting mission-critical... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $117k - $209.33k

     ...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build...  ...security, compliance, platform, and infrastructure teams to ensure services are...  ...applicableParticipate in a 24x7 on-call rotation for production servicesFunction... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  •  ...Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates...  ...engineering and applies them to infrastructure and operations problems. The main goals...  ...performance; Participate in on-call rotation to provide 24/7 support for... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    18 hours ago
  • $148.5k - $223.9k

     ...Details Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in SanFrancisco....  ...with counterparts in the Infrastructure and R&D organizations, this organization...  ...-the-sun model with weekend on‑call, the Site Reliability team keeps... 
    Senior
    Worldwide
    Weekend work

    Jobleads-US

    San Francisco, CA
    18 hours ago
  • $153k - $191.3k

     ...data processing, and software engineering, our office is a truly...  ...Planet's Direct Access Service Infrastructure team, directly contributing...  ...environments, to guarantee the reliability, scalability, and...  ...tests Participate in on-call rotations to ensure operational... 
    Senior
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office
    3 days per week

    Planet Labs PBC

    San Francisco, CA
    6 hours ago
  • $220k - $235k

     ...strategic, high‑output Staff/Senior Staff SRE to define...  ...and champion engineering excellence across Ironclad...  ...direction for the Site Reliability Engineering team and...  ...systems Be on an on‑call rotation to respond...  ...Ability to build resilient infrastructure ~ Modern GitOps –... 
    Senior
    Full time
    Work at office

    Ironclad Inc

    San Francisco, CA
    18 hours ago
  • Senior Software Engineer At Commure, we're building the AI Operating System...  ...Software Engineer on the Infrastructure team to own the foundational...  ...'ll make architectural calls, write the code that...  ...infrastructure, platform, or site reliability engineering roles. Experience... 
    Senior
    Local area
    Immediate start

    commure

    San Francisco, CA
    3 days ago
  • Software Engineer Voxel's perception system is the technical core of everything we ship....  ...strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision...  ..., write code, make architecture calls, and partner closely with applied CV, ML... 
    Senior
    Work at office
    Flexible hours

    Voxel

    San Francisco, CA
    3 days ago
  • Platform Engineer HUD is building infrastructure to create RL training data and evals for frontier AI agents,...  ...Platform Engineer who can own the reliability, scale, performance, and developer...  ...logs, traces, SLOs, runbooks, and on-call workflows so failures are detected,... 
    Senior
    Full time
    Work at office
    Remote work
    Relocation
    Visa sponsorship

    Hud (yc W25)

    San Francisco, CA
    3 days ago
  • $200k - $260k

     ...team of the world's best engineers and operators. If you are...  ...Role We are looking for a Senior Software Engineer to join our Infrastructure Team , focused on building secure, reliable, and scalable...  ...response processes, including on-call practices, runbooks, postmortems... 
    Senior
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    23 days ago
  • $232k - $319k

     ...secures AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...the service with great people and reliable, cost-effective, and efficient infrastructure...  ...the velocity of SRE and product engineering by developing robust platforms, powerful... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    5 days ago
  •  ...US Corp. is seeking a Lead Site Reliability Engineer to spearhead our mission of delivering highly available and performant systems. With an...  ...be responsible for designing and implementing automated infrastructure using Terraform, managing containerized workloads within... 
    Senior

    Axiom Pursuits

    San Francisco, CA
    18 hours ago
  • $160k - $195k

     ...vertically integrated AI infrastructure company built from...  ...We’re seeking a Senior Cloud Infrastructure Engineer to own the design, implementation...  ...IT, Security, and site-specific operations to keep systems reliable, secure, and ready...  ...; participate in on-call rotation as the... 
    Senior
    Temporary work
    Work at office

    Crusoe

    San Francisco, CA
    a month ago
  • $180k - $240k

    Senior Cloud Infrastructure Engineer Loft Orbital is revolutionizing access to space by building reliable, shareable satellites that drastically reduce the...  ...spacecraft control. Yes, we call it SatDevOps, and we'll...  ...with financial support Off-sites and many social events... 
    Senior
    Temporary work
    Work at office
    Relocation package
    Flexible hours

    Loft Orbital

    San Francisco, CA
    5 days ago
  • $175k - $250k

     ...000.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do Design... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    3 days ago
  • $164k - $205k

     ...maintain production systems Build and operate cloud infrastructure on AWS, using Terraform to codify and version-control...  ...alerting and observability systems Collaborate with engineering teams to embed reliability into the development lifecycle, shifting left on... 
    Senior
    Work experience placement
    Summer holiday
    Live out
    Work at office
    Local area
    Flexible hours
    Shift work
    2 days per week

    SupportFinity

    San Francisco, CA
    18 hours ago
  • Role SummaryWe have an opening for a Senior Software Infrastructure Engineer in our Infrastructure team focused on driving technical strategy, alignment...  ...team focused on improving the scalability and reliability of Temporal’s core infrastructure. In this role, you will... 
    Senior

    Temporal Technologies

    San Francisco, CA
    16 hours ago
  • $189k - $283.6k

     ..., you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics-driven, systems-oriented, and focused...  ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience... 
    Senior
    Full time
    Relocation package
    Flexible hours
    Shift work

    Block Inc

    San Francisco, CA
    18 hours ago
  • $181k - $263k

    ## Senior Staff Site Reliability EngineerApplylocations: San Franciscotime type: Full timeposted on: Posted...  ...for a Senior Staff Site Reliability Engineer who will set the technical direction...  ...across LiveRamp's global infrastructure. This is a senior individual contributor... 
    Senior
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    18 hours ago
  •  ...Required skills ~ Engineering ~322 Infra About Anyscale: At Anyscale...  ...: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide...  ...architecture discussions Provide on-call support, working closely with... 
    Remote job
    Full time
    San Francisco, CA
    a month ago
  • $131.75k - $178.25k

     ...practices and entrepreneurial spirit allow exceptional opportunities for professional achievement and career growth. The Senior Infrastructure DevOps Engineer, under the direction of the Senior Manager of Infrastructure, will collaborate with our internal IT teams to design,... 
    Senior
    Full time
    Work experience placement
    Remote work
    Worldwide

    Wilson Sonsini Goodrich & Rosati

    San Francisco, CA
    17 hours ago
  •  ...AfterQuery is seeking a Senior Software Engineer - Infrastructure in San Francisco to design and build core infrastructure for data generation, evaluation...  ...-scale experiments and ensure they are scalable and reliable. You will collaborate with the founding team to set... 
    Senior

    AfterQuery

    San Francisco, CA
    3 days ago
  •  ...Grow Therapy in San Francisco is seeking a Senior AI Enablement Engineer to define how AI transforms operations across the organization. You will design and build foundational AI infrastructure that enhances efficiency. Responsibilities include implementing AI systems... 
    Senior
    Flexible hours
    3 days per week

    Grow Therapy

    San Francisco, CA
    18 hours ago
  •  ...Engg, a San Francisco–based AI infrastructure company, seeks a senior on-call SRE to own incident response for Kubernetes, Ceph storage, and bare metal...  ...restores, and help reduce incident frequency through reliability improvements and IaC tooling. #J-18808-Ljbffr... 
    Senior

    Jobleads-US

    San Francisco, CA
    1 day ago
  •  ...Crusoe is seeking a Senior Software Engineer to architect, design, and develop Cloud Infrastructure management systems for Crusoe Cloud. You will deliver end-to-end workflows...  ...You'll collaborate across teams to build reliable, secure, and cost-efficient cloud platforms,... 
    Senior

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $183k - $220k

     ...all of us. ???? The Role As a senior member of our small but mighty engineering team, you'll work closely with...  ...evolution of our backend tech stack and infrastructure. Whether iterating on core...  ...our customer base and perform reliably for our users. Product development... 
    Senior
    Full time
    For contractors
    Work at office
    Remote work

    Siteline

    San Francisco, CA
    18 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Infrastructure/Site Reliability Engineer On-Call. Be the first to apply!