Software Engineer, Site Reliability (SRE)
Sierra
About us At Sierra, we're creating a platform to help businesses build better, more human customer experiences with AI. We are primarily an in-person company based in San Francisco, with growing offices in Atlanta, New York, London, Paris, Madrid, Munich, Singapore, Tokyo, and Sydney. We are guided by a set of values that are at the core of our actions and define our culture: Trust, Customer Obsession, Craftsmanship, Intensity, and Family. These values are the foundation of our work, and we are committed to upholding them in everything we do. Our co-founders are Bret Taylor and Clay Bavor. Bret currently serves as Board Chair of OpenAI. Previously, he was co-CEO of Salesforce (which had acquired the company he founded, Quip) and CTO of Facebook. Bret was also one of Google's earliest product managers and co-creator of Google Maps. Before founding Sierra, Clay spent 18 years at Google, where he most recently led Google Labs. Earlier, he started and led Google's AR/VR effort, Project Starline, and Google Lens. Before that, Clay led the product and design teams for Google Workspace.
What you'll do As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability, observability, and scalability across Sierra's AI-driven infrastructure. You'll partner closely with our core engineering and product teams to ensure our systems are highly available, efficient, and built for growth.
These benefits are further detailed in Sierra's policies, may vary by region, and are subject to change at any time, consistent with the terms of any applicable compensation or benefits plans. Eligible full-time employees can participate in Sierra's equity plans subject to the terms of the applicable plans and policies. Be you, with us We're working to bring the transformative power of AI to every organization in the world. To do so, it is important to us that the diversity of our employees represents the diversity of our customers. We believe that our work and culture are better when we encourage, support, and respect different skills and experiences represented within our team. We encourage you to apply even if your experience doesn't precisely match the job description. We strive to evaluate all applicants consistently without regard to race, color, religion, gender, national origin, age, disability, veteran status, pregnancy, gender expression or identity, sexual orientation, citizenship, or any other legally protected class.
What you'll do As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability, observability, and scalability across Sierra's AI-driven infrastructure. You'll partner closely with our core engineering and product teams to ensure our systems are highly available, efficient, and built for growth.
- Own Sierra's observability stack-monitoring, alerting, logging, and tracing-to give engineers clear visibility into system health and performance.
- Partner with product and platform engineers to design systems that are reliable and scalable from day one-not as an afterthought.
- Design and implement scalable, reliable, and secure cloud infrastructure (AWS) using Terraform and modern DevOps tooling.
- Improve the reliability and scalability of our LLM deployments, ensuring robust, performant, and cost-effective operation.
- Lead improvements to deployment pipelines, CI/CD tooling, and incident management processes to reduce downtime and response time.
- Define the foundation of SRE practices at Sierra, influencing culture, tooling, and best practices across the engineering org.
- 5+ years of hands-on experience in Site Reliability or Infrastructure engineering roles for complex SaaS or cloud-based systems.
- Experience designing for availability, scalability, and reliability at both infrastructure and application layers.
- Deep experience with Terraform, AWS services, container orchestration, and cloud networking (including IAM and VPC architecture).
- Strong background in observability systems (e.g., Prometheus, Grafana, Datadog, or similar).
- Experience working with enterprise customers and familiarity with their compliance and networking needs along with integration patterns.
- Comfortable working in fast-moving environments and collaborating across product, ML, and core engineering teams.
- Degree in Computer Science or a related field, or equivalent professional experience.
- Experience with LLM infrastructure - optimizing inference performance, managing fine-tuned models, or large-scale model deployment.
- Past experience in an early-stage startup environment, especially defining SRE culture and tooling from scratch.
- Familiarity with incident management automation or self-healing infrastructure patterns.
- Trust: We build trust with our customers with our accountability, empathy, quality, and responsiveness. We build trust in AI by making it more accessible, safe, and useful. We build trust with each other by showing up for each other professionally and personally, creating an environment that enables all of us to do our best work.
- Customer Obsession: We deeply understand our customers' business goals and relentlessly focus on driving outcomes, not just technical milestones. Everyone at the company knows and spends time with our customers. When our customer is having an issue, we drop everything and fix it.
- Craftsmanship: We get the details right, from the words on the page to the system architecture. We have good taste. When we notice something isn't right, we take the time to fix it. We are proud of the products we produce. We continuously self-reflect to continuously self-improve.
- Intensity: We know we don't have the luxury of patience. We play to win. We care about our product being the best, and when it isn't, we fix it. When we fail, we talk about it openly and without blame so we succeed the next time.
- Family: We know that balance and intensity are compatible, and we model it in our actions and processes. We are the best technology company for parents. We support and respect each other and celebrate each other's personal and professional achievements.
- Flexible (unlimited) paid time off
- Medical, dental, and vision benefits for you and your family
- Life insurance and disability benefits
- Retirement plan dependent on country of employment
- Parental leave
- Fertility and family building benefits through Carrot
- Lunch, as well as delicious snacks and coffee to keep you energized
- Discretionary benefit stipend giving people the ability to spend where it matters most
- Free alphorn lessons
These benefits are further detailed in Sierra's policies, may vary by region, and are subject to change at any time, consistent with the terms of any applicable compensation or benefits plans. Eligible full-time employees can participate in Sierra's equity plans subject to the terms of the applicable plans and policies. Be you, with us We're working to bring the transformative power of AI to every organization in the world. To do so, it is important to us that the diversity of our employees represents the diversity of our customers. We believe that our work and culture are better when we encourage, support, and respect different skills and experiences represented within our team. We encourage you to apply even if your experience doesn't precisely match the job description. We strive to evaluate all applicants consistently without regard to race, color, religion, gender, national origin, age, disability, veteran status, pregnancy, gender expression or identity, sexual orientation, citizenship, or any other legally protected class.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer, Site Reliability (SRE) in San Francisco, CA vacancy
- ...design teams for Google Workspace. What you'll do As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and... ...downtime and response time. Define the foundation of SRE practices at Sierra, influencing culture, tooling, and...WebsiteFull timeFlexible hours
- ...Software Engineer As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability,... ...and response time. Defining the foundation of SRE practices at Sierra, influencing culture, tooling...WebsiteFull timeFlexible hours
$350k
...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower humanity through advancing collaborative general... ...or site reliability engineering. Proficiency writing software to solve reliability problems, including building tooling...WebsiteLocal areaVisa sponsorshipWork visaRelocation package- ...management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the... ...of our systems. You will work at the intersection of software development and operations, implementing automation,...WebsiteTemporary workWorldwide
- ...We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities... ...startups or added to our vetted SRE network for future projects. This role... ...scale and improving how teams ship software, you’ll fit right in. Key Responsibilities...WebsiteLocal area
$180k - $250k
You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing... ...production issues, and improve software development speed,... ...automation, runbooks, and chaos engineering Requirements 5+ years experience...WebsiteCurrently hiringRemote workRelocation package$160k - $300k
...and market leadership. The Role We are looking for a Site Reliability Engineer who thinks like a software engineer first. You will own critical production... ...or Rust Proven experience as a Production Engineer, SRE, or software engineer with a deep infrastructure focus...Website$180k - $200k
About The Role As a Senior Site Reliability Engineer on our rapidly growing team, you’ll measure how well our software is working, and from that baseline propel us towards high impact... ...Think) You’ll Need To Do It 5+ years of SRE, DevOps, or Platform engineering experience...WebsiteWork at office3 days per week- Gravity Engineering Services Pvt Ltd. is seeking a Site Reliability Engineer to ensure system reliability and efficiency. The role includes defining SLOs, implementing observability stacks, and automating processes alongside incident retrospectives. Candidates should possess...Website
- Gravity Engineering Services Pvt Ltd. is seeking a Site Reliability Engineer in San Francisco, California. The role involves defining and tracking SLOs, designing observability stacks, and automating system reliability. The ideal candidate should have 3+ years of experience...Website
$325k
Anthropic is seeking a Reliability Engineer to enhance the reliability of AI services in San Francisco. The role involves developing Service Level... ...candidates have strong backgrounds in distributed systems or site reliability engineering, excellent communication skills, and...Website- Plaid is seeking a Staff Site Reliability Engineer on Release Engineering to scale reliability across product engineering. You will architect... ...service tooling, lead incident responses, and shape deployment health for faster, safer software delivery. #J-18808-Ljbffr PlaidWebsite
$200k - $260k
Senior Software Engineer, Site Reliability Engineer (SRE) Why Harvey At Harvey, we’re transforming how legal and professional services operate — not incrementally, but end‑to‑end. By combining frontier agentic AI, an enterprise‑grade platform, and deep domain expertise...WebsiteRelocation package$300k
...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and... .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting...WebsitePermanent employment- ...help build the future of identity with a team that holds a high bar for itself — keep reading. We are hiring exceptional Site Reliability Engineers who take pride in building and operating mission-critical, production-grade systems. This role is for engineers who own...Website
$180k - $200k
...days for team or company events. _ Software Engineer, Platform Infrastructure sits under the... ...systems to provide our customers with reliable, secure, and scalable software. Roles... ...Responsibilities: Be part of the Cloud Platform SRE Team, focused on building our Cloud...WebsiteContract workWork at office- ...Open Source LLM Gateway Engineer LiteLLM is an open-source LLM Gateway with 34K+ stars... ...seeking our 6th Engineer focused on owning reliability, performance, and infrastructure... ...on) or JWT (JSON Web Tokens). As the SRE, you'll own the reliability and performance...Website
$325k
About The Role AIRE (AI Reliability Engineering) partners with teams across Anthropic... ...serving - critical for both site reliability and Anthropic’s... ...for reliability‑minded software engineers and SREs Curious... ...Qualifications Experience as an SRE, Production Engineer, or in...WebsiteVisa sponsorship- A profitable AI development firm in San Francisco is seeking an experienced Site Reliability Engineer (SRE) to own production reliability across critical systems. You will build the SRE function from the ground up and define priorities for reliability and production safety...Website
- ...DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will... ...Skill Requirements: - Engineering Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production...Website
$260k - $300k
...Applied AI Lab Building Software Agents We are an applied AI lab... ...Devin, the first AI software engineer, and Windsurf, the AI-native... ...will own both the production reliability of our user-facing products and... ...engineering fundamentals; SRE at Cognition means writing real...Website- ...About the Role We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate, and maintain...Website
- ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure... ...your decisions. Required skills ~3+ years in SRE, DevOps, or infrastructure engineering roles ~ Strong...Website
- ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract... ...owners, scrum masters, and architects. The SRE ensures that both our internally critical... ...system design review, developing software platforms and frameworks, capacity planning...WebsiteContract workLocal area
- ...Job: Staff Site Reliability Engineer (SRE) Location: San Francisco, CA Job Responsibilities As our Staff SRE, you'll be the primary expert responsible for our entire compute ecosystem. Your key responsibilities will include: As a Staff SRE, you...Website
- ...Senior SRE Unify is building the first AI-native outbound... ...you'll tackle the scaling and reliability challenges that come with adding... ..., and alerting that give engineers clear visibility into system... ...Who You Are ~5+ years of software engineering experience with a...Website
- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... .... You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting...WebsiteFlexible hours
- ...billion. We work in-person five days a week in our San Francisco, NYC, or London offices. About the Role As a Site Reliability Engineer (SRE) at Mercor, you'll own production reliability across our most critical systems, partnering directly with infrastructure...WebsiteWork at officeRelocation package
- ...To achieve our ambitious goals, we're looking for an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth....WebsiteWorldwideHome officeFlexible hours
$227.2k - $324.5k
A leading streaming service located in San Francisco is seeking a Senior SRE Manager to spearhead the Site Reliability Engineering team. This role will focus on driving the technical strategy for observability and automation while fostering a culture of innovation and excellence...Website
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Site Reliability (SRE). Be the first to apply!
Related searches
- ngo software engineer San Francisco, CA
- software developer San Francisco, CA
- software developer internship no experience San Francisco, CA
- junior software developer San Francisco, CA
- part time software developer remote San Francisco, CA
- financial software developer San Francisco, CA
- senior software engineer ruby on rails San Francisco, CA
- software engineer amazon San Francisco, CA
- senior software design engineer San Francisco, CA
- software engineer part time San Francisco, CA


