Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$195k - $240k

Y.O.U.

Senior Site Reliability Engineer

San Francisco (Hybrid)

At You.com, we are building the AI Search Infrastructure that powers modern AI systems. Our goal is to create the trusted knowledge layer that agents, applications, and enterprises rely on to retrieve real-time, accurate, and citation-backed information.

Our platform combines proprietary vertical indexes with LLM-optimized retrieval systems to power AI agents, applications, and enterprise workflows. We are solving hard problems across search, large language models, and large-scale infrastructure to make AI systems more reliable, transparent, and useful.

Our team includes engineers, researchers, product builders, and operators who care about solving meaningful problems and delivering real-world impact. Whether you are improving core infrastructure, shaping product experiences, or helping bring new AI capabilities to market, your work will help define how modern AI finds and uses knowledge.

About the Role

As a Site Reliability Engineer, you will own parts of the reliability, observability, and incident response posture for You.com's production services. Your work will ensure that every user query, every API call, and every data pipeline runs with measurable, defensible uptime, and when something breaks, the tools and dashboards you developed will help the team identify the issue, respond, and learn from it. Additionally, you will partner with teams to help them implement best practices, establish reliability objectives, and ensure the engineering team can build reliable services with minimal friction.

Responsibilities
  • Instrument services end-to-end using OpenTelemetry metrics and structured logging to ensure every critical path is measurable.
  • Develop and maintain SRE standards and patterns (instrumentation guidelines, incident playbooks, service templates) that engineering teams adopt by default in new and existing services. Build internal tooling and automation in Python, Bash and Terraform to improve deployment safety, reliability, and operational efficiency.
  • Design and maintain actionable dashboards that surface real user impact, not vanity metrics, for service owners and leadership.
  • Tune alerting rules continuously to maximize signal-to-noise ratio; tie alerts to SLO-based error-budget burn rates rather than arbitrary thresholds.
  • Own reliability incident response end-to-end: detection, triage, communication, escalation, resolution, and stakeholder updates.
  • Track and run blameless postmortems that focus on systemic contributing factors, not individual fault, producing actionable remediation items with owners and deadlines.
  • Track remediation follow-through as a first-class metric. Ensure postmortem action items are completed, not just documented.
  • Continuously improve MTTD and MTTR by feeding incident learnings back into monitoring, runbooks, and automation.
  • Collaborate with Customer Success and ensure we by feed incident learnings back into monitoring, runbooks, and automation.
  • Define meaningful SLOs for all production services grounded in critical user journeys, historical performance data, and business requirements.
  • Eliminate alert fatigue by auditing, categorizing, and deprecating noisy or non-actionable alerts on a regular cadence.
  • Help manage incident management processes and playbooks.
Qualifications
  • 2+ years of full-time experience in an SRE or similar role
  • 3+ years of experience working in AWS with EKS and Github (GHA) & CI/CD
  • Strong hands-on experience with Git, Python, and Bash. Comfortable building production-grade automation and tooling.
  • Experience establishing SRE practices across multiple teams (SLO definitions, alert hygiene, postmortem culture).
  • Built or maintained Prometheus-based monitoring with dashboards they have in Grafana.
  • Demonstrated experience scoping and delivering infrastructure projects from proposal through production deployment
  • Demonstrated experience managing incidents and response to service outage
  • Hands-on experience integrating AI with SRE efforts to improve reliability, development and velocity
  • Demonstrated track record of collaborating with teams to define SLOs, instrument services against measurable SLIs, and operationalize error-budget burn-rate alerting that teams use independently to balance risk and delivery speed.

Our salary bands are structured based on a combination of geographic tiers and internal leveling. Compensation is determined by multiple factors assessed during the interview process, with the final offer reflecting these considerations.

Salary Band

$195,000 - $240,000 USD

Company Perks:
  • Hubs in San Francisco and New York City offering regular in-person gatherings and co-working sessions
  • Flexible PTO with U.S. holidays observed and a week shutdown in December to rest and recharge*
  • A competitive health insurance plan covers 100% of the policyholder and 75% for dependents*
  • 12 weeks of paid parental leave in the US*
  • 401k program, 3% match - vested immediately!*
  • $500 work-from-home stipend to be used up to a year of your start date*
  • $600 technology stipend to support a portion of our hybrid/remote team's cell phone and internet expenses*
  • $1,200 per year Health & Wellness Allowance to support your personal goals*
  • The chance to collaborate with a team at the forefront of AI research

*Certain perks and benefits are limited to full-time employees only

You.com participates in E-Verify. We will provide the Social Security Administration (SSA) and, if necessary, the Department of Homeland Security (DHS) with information from each new employee's Form I-9 to confirm work authorization. (English/Spanish: E-Verify Participation / Right to Work ) We are also an inclusive, equitable, and accessible workplace. Please let us know if you require accommodation for any portion of the recruitment and hiring process.

Beware of recruiting scams: You.com will only contact you through official @ You.com email addresses and will never ask for payment or sensitive personal information during the hiring process.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
  •  ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate... 
    Senior

    Alembic Technologies

    San Francisco, CA
    4 days ago
  •  ...Senior SRE Unify is building the first AI-native outbound platform where agents and...  ...Senior SRE, you'll tackle the scaling and reliability challenges that come with adding...  ...tracing, metrics, and alerting that give engineers clear visibility into system behavior and... 
    Senior

    Unify

    San Francisco, CA
    1 day ago
  • $181.69k - $213.75k

     ...Senior Site Reliability Engineer San Francisco, California; Santa Clara, California; Seattle, WA The Company You'll Join Carta connects founders, investors, and limited partners through world-class software, purpose-built for everyone in venture capital, private... 
    Senior
    Full time
    Work at office

    Carta

    San Francisco, CA
    4 days ago
  • $159.2k - $301.6k

     ...running Graphs on the cloud. In this reliability-focused role, you will own the availability...  .... You'll partner with the backend engineers building these APIs to make sure the system...  ...Science. ~5-10 years of experience in site reliability engineering, infrastructure,... 
    Senior
    Temporary work
    Local area
    Worldwide

    Adobe

    San Francisco, CA
    4 days ago
  • $117k - $209.33k

     ...Job Requisition ID # 26WD99273 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products. As part of a... 
    Senior
    For contractors

    Autodesk

    San Francisco, CA
    5 days ago
  • $148.5k - $223.9k

     ...duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce...  ...future of Salesforce. Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely... 
    Senior
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  •  ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    1 day ago
  • $127k - $249k

    THE TEAM Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational...  ..., alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    5 days ago
  • $166.9k - $225.9k

     ...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team...  ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building... 
    Senior
    Work at office
    Immediate start
    Worldwide
    Monday to Friday
    Flexible hours

    Drata Inc

    San Francisco, CA
    1 day ago
  • $287k

     ...Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was...  ...expect us to hit our SLAs. What? We're looking for a Senior or Staff Site level Reliability Engineer as part of Infrastructure team to: Own... 
    Senior
    Contract work
    Work at office
    Remote work

    IVO Inc

    San Francisco, CA
    4 days ago
  • $220k - $235k

     ...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move faster, insights surface instantly, and agents push work forward, all with you in control. Whether you're buying or selling... 
    Senior
    Full time
    Contract work
    Work at office

    Ironclad Inc

    San Francisco, CA
    1 day ago
  • $160k - $300k

    Hebbia is looking for a Site Reliability Engineer to manage and improve critical production systems. This role requires writing production-quality code and collaborating with product engineering teams to enhance system reliability. Applicants should have 5+ years in software... 
    Senior

    Hebbia

    San Francisco, CA
    5 days ago
  • $170k - $190k

    Medrio is seeking a Senior Site Reliability Engineer in San Francisco, California. The role involves maintaining and supporting the Medrio Platform, troubleshooting issues, and implementing solutions. Candidates should have experience in Kubernetes, cloud services, and... 
    Senior
    Flexible hours

    Medrio

    San Francisco, CA
    3 days ago
  • $210.8k - $272.8k

    Thumbtack is hiring for a Site Reliability Engineer to enhance the reliability and scalability of our services in San Francisco, CA. You'll design and support resilient systems, ensuring a smooth user experience. The ideal candidate will have extensive experience with AWS... 
    Senior

    Thumbtack

    San Francisco, CA
    5 days ago
  • Early Warning is seeking a Staff Site Reliability Engineer to enhance application performance and resiliency while guiding development teams. This role involves designing automation and monitoring systems, improving scalability and availability, and participating in a 2... 
    Senior

    Early Warning

    San Francisco, CA
    5 days ago
  • A tech company specializing in AI is seeking a Site Reliability Engineer to ensure the reliability and observability of their production services. You will instrument services, develop SRE standards, and manage incident response. The ideal candidate has experience in AWS... 
    Senior
    Remote work
    Flexible hours

    You.com

    San Francisco, CA
    5 days ago
  • $180k - $200k

    Parabola is seeking a Senior Site Reliability Engineer to join their team in San Francisco. In this role, you will monitor and improve software performance, maintain infrastructure, and collaborate with engineering teams. Candidates should have over 5 years of experience... 
    Senior

    Parabola

    San Francisco, CA
    5 days ago
  • $181k - $263k

     ...and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a... 
    Senior
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    1 day ago
  •  ...multimodality is critical for intelligence. This requires a massive, reliable, and performant GPU infrastructure that pushes the boundaries...  ...You Come In We are looking for a hands‑on, first‑principles engineer who is fluent in Linux, comfortable operating close to the... 
    Senior
    Work experience placement

    Luma AI

    San Francisco, CA
    2 days ago
  • Location San Francisco, CA Employment Type Full time Department Engineering Who We Are Hyperbolic Labs is on a mission to democratize AI...  ...to redefine computing. About the Role We\'re seeking a Site Reliability Engineer to ensure Hyperbolic\'s GPU marketplace and AI infrastructure... 
    Senior
    Full time

    Hyperbolic

    San Francisco, CA
    5 days ago
  • CloudDevs works with fast-moving, venture-backed startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will either be placed directly into one of our partner startups or added to our... 
    Senior
    Local area

    Breakout Tools

    San Francisco, CA
    1 day ago
  • $170k - $190k

    Position Medrio Senior Site Reliability Engineer Responsibilities Build, maintain, and support all environments which host the Medrio Platform Monitor environments for issues, configuring and building alerting/self‑healing of issues Update and maintain documentation... 
    Senior
    Temporary work
    Flexible hours

    Medrio

    San Francisco, CA
    3 days ago
  • $175k - $250k

     ...00.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    1 day ago
  • $210.8k - $272.8k

    About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user... 
    Senior
    Local area

    Thumbtack

    San Francisco, CA
    5 days ago
  •  ...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics...  ...of accountability A strong desire to perform and grow as an engineer 5+ years of software development experience Technologies We... 
    Senior
    Flexible hours

    Block, Inc.

    San Francisco, CA
    5 days ago
  • $200k - $260k

    Senior Software Engineer, Site Reliability Engineer (SRE) Why Harvey At Harvey, we’re transforming how legal and professional services operate — not incrementally, but end‑to‑end. By combining frontier agentic AI, an enterprise‑grade platform, and deep domain expertise... 
    Senior
    Relocation package

    Harvey

    San Francisco, CA
    5 days ago
  •  ...systems. It's designed so Stellar's ecosystem can make a real-world, lasting impact. About the Role SDF is looking for a Senior Site Reliability Engineer to help build and operate the foundation that powers our engineering teams. You'll ensure the reliability and... 
    Senior

    TechChain Talent

    San Francisco, CA
    5 days ago
  • $160k - $250k

     ...time Location Type Hybrid Department Engineering Spend Compensation 160,000 - 250,000;...  ...ownership, working together to build scalable, reliable, and secure products that empower...  ...Global services. What you’ll do As a Senior Site Reliability Engineer, you’ll work... 
    Senior
    Full time
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    5 days ago
  • $213k - $263k

     ...driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Software Reliability Engineers (SREs) are responsible for the stable operation of Waymo's fully autonomous systems and supporting infrastructure. As an SRE... 
    Senior
    Full time
    Remote work

    Waymo

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!