Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Runloop AI, Inc

About Runloop

Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.

The Role

We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset.

Responsibilities
  • Design and maintain our production infrastructure on cloud platforms like AWS, GCP, Azure, and emergent Neo-Clouds
  • Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users
  • Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind
  • Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment
  • Participate in an on-call rotation to support our production systems
  • Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing
  • Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems
  • Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement
  • Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience
  • Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes
Qualifications
  • Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience
  • 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations
  • Strong programming skills in languages like Python or Go
  • Deep expertise in containerization technologies such as Docker and Kubernetes
  • Experience with cloud infrastructure and tools like Terraform and/or Pulumi
  • Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog
  • A solid understanding of networking, security, and Linux systems administration
  • Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)
  • Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity
  • Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems
  • Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance
Bonus Points
  • Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery
Benefits
  • Competitive salary and equity
  • Comprehensive health, dental, and vision insurance for employee and dependents
  • Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering
  • Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks
Location :
  • Onsite 4 days a week in San Francisco; Optional 1 day a week remote

Join Us! If you're excited about shaping the future of AI-driven software engineering and empowering developers to build the next generation of AI powered coding tools, we want to hear from you. Join the Runloop team and be at the forefront of the AI revolution in software development.

Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in San Francisco, CA vacancy
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering...  ...You will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes... 
    Suggested

    Socket

    San Francisco, CA
    2 days ago
  •  ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage... 
    Suggested

    Lambda

    San Francisco, CA
    1 day ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business... 
    Suggested

    J.P. Morgan

    San Francisco, CA
    4 days ago
  • $194k - $237k

     ...employer, at the date of hire. This position is ineligible for employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and... 
    Suggested
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    San Francisco, CA
    1 day ago
  • $200.7k - $250.9k

     ...washed away in a flood in 1942, the Royal Engineers rebuilt it. Then it washed away again in...  ...opportunities for improvements in reliability/observability/performance/preparedness and...  ...candidate for the role: Has past Site Reliability Engineering or DevOps experience... 
    Suggested

    Embedded Shishya

    San Francisco, CA
    1 day ago
  • $200k - $240k

     ...Senior Site Reliability Engineer (SRE)Location: San Francisco, CAWork Model: OnsiteIndustry: Renewable EnergyComp: $200,000 - $240,000 We’re partnering with a fast-growing energy technology company looking for a Senior Site Reliability Engineer to take ownership of... 

    Lawrence Harvey Search & Selection

    San Francisco, CA
    1 day ago
  •  ...troubleshooting, providing rubric-based written feedback. This role requires hands-on Kubernetes expertise in EKS/GKE/AKS or self-managed clusters, with strong scripting in Go, Python, or TypeScript, and ability to document findings clearly for engineering #J-18808-Ljbffr

    Obsidian

    San Francisco, CA
    3 days ago
  • $350k

     ...with leading AI companies and infrastructure providers to build reliable, high-performance platforms supporting next-generation AI workloads. This opportunity is for a Staff Site Reliability Engineer to lead the reliability of large-scale GPU infrastructure, covering... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    1 day ago
  • $200k - $240k

     ...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to...  ...and Helm ~ Understand the importance of performant and reliable systems ~ Education - Ideally looking for a B.A. / B.S. degree... 
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    21 hours ago
  •  ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built... 
    Relocation package

    1872 Consulting

    San Francisco, CA
    4 days ago
  • $170k - $220k

     ...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes. Our innovative approach blends technology with deep legal expertise, making us a leader in our field. We go beyond surface-level... 
    Work at office
    Remote work
    Flexible hours

    Supio

    San Francisco, CA
    21 hours ago
  •  ...enterprise that runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging... 
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    1 day ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round — Healthcare, AI Office Type: Onsite Salary: $200K–$300K Company Description We're representing a dynamic company... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    4 days ago
  •  ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here...  ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with... 
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    21 hours ago
  • $98.58k - $138.02k

     ...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering... 
    Work at office

    Restaurant365

    San Francisco, CA
    2 days ago
  •  ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical... 
    Remote work

    Specter Services LLC

    San Francisco, CA
    1 day ago
  • $130k - $200k

     ...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership... 
    Shift work

    Nscale

    San Francisco, CA
    4 days ago
  •  ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online...  ...foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at... 
    Permanent employment
    Shift work

    Arena AI

    San Francisco, CA
    3 days ago
  • $150k

     ...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security... 

    VantageScore®

    San Francisco, CA
    21 hours ago
  •  ...JOB DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will help Shipt understand where we can improve stability and reliability. There will be a focus on the intersection of systems... 

    BayOne Solutions

    San Francisco, CA
    1 day ago
  •  ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You'll be building and operating the core systems that power agentic AI at scale. Your mission: keep... 

    Blaxel, Inc

    San Francisco, CA
    2 days ago
  •  ...had design and make. Operate is the phase that tells you what actually happened, and it is ours. We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability, observability, and developer autonomy as we scale our platform. In this role,... 
    Remote work

    MaintainX

    San Francisco, CA
    3 days ago
  • $260k - $300k

     ...software agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding...  ...faster than anyone expects. You will own both the production reliability of our user-facing products and the platform engineering that... 

    Cognition AI

    San Francisco, CA
    4 days ago
  •  ...access to life-saving treatment. What We Look for in a Great Engineer You have the intensity and technical mastery to own...  ...support high-velocity feature release while maintaining the highest reliability. DevX Support: Support Developer Experience (DevX) work to... 
    Work at office

    LATENT

    San Francisco, CA
    9 hours ago
  •  ...globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You... 
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...0/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance of...  ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design,... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    1 day ago
  •  ...guarantees and certifications. We're hiring staff-level SREs to help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions for how we deploy, observe, and secure the platform, and help... 
    Remote work
    Flexible hours

    Akka

    San Francisco, CA
    17 days ago
  • $165k - $227k

     ...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    10 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!