Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Runloop AI, Inc

About Runloop

Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.

The Role

We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset.

Responsibilities
  • Design and maintain our production infrastructure on cloud platforms like AWS, GCP, Azure, and emergent Neo-Clouds
  • Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users
  • Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind
  • Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment
  • Participate in an on-call rotation to support our production systems
  • Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing
  • Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems
  • Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement
  • Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience
  • Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes
Qualifications
  • Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience
  • 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations
  • Strong programming skills in languages like Python or Go
  • Deep expertise in containerization technologies such as Docker and Kubernetes
  • Experience with cloud infrastructure and tools like Terraform and/or Pulumi
  • Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog
  • A solid understanding of networking, security, and Linux systems administration
  • Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)
  • Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity
  • Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems
  • Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance
Bonus Points
  • Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery
Benefits
  • Competitive salary and equity
  • Comprehensive health, dental, and vision insurance for employee and dependents
  • Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering
  • Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks
Location :
  • Onsite 4 days a week in San Francisco; Optional 1 day a week remote

Join Us! If you're excited about shaping the future of AI-driven software engineering and empowering developers to build the next generation of AI powered coding tools, we want to hear from you. Join the Runloop team and be at the forefront of the AI revolution in software development.

Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in San Francisco, CA vacancy
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Suggested

    Alembic

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,... 
    Suggested
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Suggested
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    2 days ago
  • $139.76k - $287.75k

     ...their business.We are seeking a Senior Site ReliabilityEngineer to help operate, scale...  ...will be instrumental in advancing the reliability, scalability, automation, observability,...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Suggested
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    1 day ago
  • $167.7k - $245.2k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

     ...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical... 
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  •  ...let’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ..., working together to build scalable, reliable, and secure products that empower businesses...  ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    2 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Temporary work

    TextNow

    San Francisco, CA
    23 hours ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure... 
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  • $150k

     ...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our... 

    VantageScore

    San Francisco, CA
    1 day ago
  •  ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built... 
    Relocation package

    1872 Consulting

    San Francisco, CA
    1 day ago
  • $150k - $250k

     ...Site Reliability Engineer role USC or GC only are considered at this time. San Francisco - Local to Bay area only but role is remote and occasion meeting required Latest update, 03/31/2026: The Site Reliability Engineer role is critical for... 
    Work experience placement
    Casual work
    Local area
    Immediate start
    Remote work

    3B Staffing LLC

    San Francisco, CA
    1 day ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description... 
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    San Francisco, CA
    3 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    San Francisco, CA
    1 day ago
  •  ...founders with PhDs in AI, Math, and Computer Science - is poised to redefine computing. About the Role We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and... 

    Hyperbolic Labs

    San Francisco, CA
    1 day ago
  • $86k - $105k

     ...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best...  ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    San Francisco, CA
    3 days ago
  • $98.58k - $138.02k

     ...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering... 
    Work at office

    Restaurant365

    San Francisco, CA
    22 hours ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    1 day ago
  • $200k - $300k

     ...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round — Healthcare, AI Office Type: Onsite Salary: $200K–$300K Company Description We're representing a dynamic company... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    1 day ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    1 day ago
  •  ...access to life-saving treatment. What We Look for in a Great Engineer You have the intensity and technical mastery to own...  ...support high-velocity feature release while maintaining the highest reliability. DevX Support: Support Developer Experience (DevX) work to... 
    Work at office

    Latent

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human... 
    Remote work
    1 day per week

    Runloop AI

    San Francisco, CA
    2 days ago
  • $163.71k - $306k

     ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 

    Retool

    San Francisco, CA
    2 days ago
  • $155k - $222.6k

     ...global cloud platform. As a team of six engineers distributed across the US, Canada, and the...  ...with a strong focus on automation, reliability, and operational excellence. We are one...  ...Qualifications ~2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    San Francisco, CA
    2 days ago
  •  ...healthcare, we'd love to meet you. Apply now to join our growing team. About the Role Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    1 day ago
  •  ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring through a focused Engineering Hiring Sprint...  ...: Platform Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers AI Agents Engineers... 
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    2 days ago
  • $160k - $250k

     ...public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able... 

    Hive

    San Francisco, CA
    1 day ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!