Site Reliability Engineer
Runloop AI, Inc
About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale. The Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset. Responsibilities
- Design and maintain our production infrastructure on cloud platforms like AWS, GCP, Azure, and emergent Neo-Clouds
- Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users
- Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind
- Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment
- Participate in an on-call rotation to support our production systems
- Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing
- Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems
- Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement
- Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience
- Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes
- Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience
- 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations
- Strong programming skills in languages like Python or Go
- Deep expertise in containerization technologies such as Docker and Kubernetes
- Experience with cloud infrastructure and tools like Terraform and/or Pulumi
- Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog
- A solid understanding of networking, security, and Linux systems administration
- Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)
- Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity
- Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems
- Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance
- Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery
- Competitive salary and equity
- Comprehensive health, dental, and vision insurance for employee and dependents
- Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering
- Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks
- Onsite 4 days a week in San Francisco; Optional 1 day a week remote
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in San Francisco, CA vacancy
- ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built...SuggestedRelocation package
$150k - $250k
...Site Reliability Engineer role USC or GC only are considered at this time. San Francisco - Local to Bay area only but role is remote and occasion meeting required Latest update, 03/31/2026: The Site Reliability Engineer role is critical for...SuggestedWork experience placementCasual workLocal areaImmediate startRemote work$150k
...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security...Suggested$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering...SuggestedWork at office$100k - $170k
...Site Reliability Engineer Houston; San Francisco; Seattle About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services...SuggestedFlexible hoursShift work$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description...Work at officeVisa sponsorshipFlexible hours- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient...
$113.4k - $162k
...Site Reliability Engineer San Francisco, CA We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet...Temporary work$160k - $250k
...public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able...$163.71k - $306k
...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system... ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially...$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,...Work at officeLocal area1 day per week$86k - $105k
...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best... ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours- ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge...Work at officeWeekend work
$200k - $300k
...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round - Healthcare, AI Office Type: Onsite Salary: $200K-$300K Company Description We're representing a dynamic...Work at office$210k - $240k
Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $2...Full time- ...company valued at $10 billion. We work in‑person five days a week in our new San Francisco headquarters. About the Role As a Site Reliability Engineer (SRE) at Mercor, you’ll own production reliability across our most critical systems, partnering directly with...
$175k - $250k
...0/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance of... ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design,...Full timeRemote workRelocationRelocation package- ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it...Temporary workWork experience placement
- ...Staff Site Reliability Engineer (SRE) Location: San Francisco, CA Job Responsibilities As our Staff SRE, you'll be the primary expert responsible for our entire compute ecosystem. Your key responsibilities will include: As a Staff SRE, you'll operate at the...
$205k - $305k
...Director Of Site Reliability Engineering Interested in working on cutting-edge blockchain technology and creating equitable access to the global financial system? Since 2014, the mission-driven team at the Stellar Development Foundation (SDF) has helped fuel the tremendous...Temporary workWork at officeLocal areaWorldwideFlexible hours$204k - $281k
.... This is an opportunity to do career-defining work. We’re all in on this mission. If you are too, let’s talk. Manager, Site Reliability Engineering San Francisco, California Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on...Permanent employmentWorldwideFlexible hours$210.8k - $272.8k
About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user...Local area- ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...platform capabilities in partnership with architects and product engineering Build a world-class observability platform and monitoring...
$174.92k - $209.91k
...: to make access to data as simple and reliable as electricity. With Fivetran, customer... ...canonical and ready to query, with no engineering or maintenance required. We’re proud that... ...integrate our teams, systems, and career sites. About the Role Fivetran is building...Full timeWork at officeRemote work- ...ambitious goals and attract incredibly creative scientists and engineers from leading academic institutions and from frontier AI labs... ...human brain. Position Summary We are looking for a Site Reliability Engineer to own the digital infrastructure that powers our...Visa sponsorship
- ...design of information and operational support systems. Required Skills/Qualifications: BS/MS degree in Computer Science, Engineering, or a related subject. Equivalent experience accepted. Proven working experience in installing, configuring, and troubleshooting...Permanent employmentWork experience placementStart working todayRemote workFlexible hours
$160k - $230k
...As a Site Reliability Engineer (SRE) at Together, you are responsible for keeping all user-facing services and production systems running smoothly. You are a blend of a pragmatic operator and a software engineer that applies sound engineering principles, operational discipline...Remote jobFull timeWork experience placement$250k
...Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves working...Permanent employmentRemote work- ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will...
$300k
...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring...Permanent employment
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer sre San Francisco, CA
- site reliability engineer San Francisco, CA
- site safety San Francisco, CA
- historic site San Francisco, CA
- IT site lead San Francisco, CA
- site leader San Francisco, CA
- site recruiter San Francisco, CA
- website coordinator San Francisco, CA
- website content developer San Francisco, CA
- site services specialist San Francisco, CA

