Staff Site Reliability Engineer
$200k - $260kSight Machine
Team Culture Great things happen when people can bring their authentic selves to work. We empower all of our employees to share their perspectives, passions and experiences because collectively we make a better, stronger team. Our team members collaborate closely with peers & cross functional stakeholders throughout the business, our clients on the forefront of digital transformation, and the cutting edge of digital manufacturing thought leadership.
We take pride in our self-starter culture where employees are enabled and encouraged to achieve their professional goals through leadership guidance, learning and development. Our philosophy is that careers are continuous journeys, and we dedicate time and offer resources so that employees can reach their full potential. Benefits + Perks We value you at and outside of work and know your loved ones are important. Our benefits are designed to support you and your family's health through life's expected and unexpected events. Our Benefits Include:
Sight Machine has offices in San Francisco, CA and Ann Arbor, Mi. We do have a remote-friendly culture with people based all around the US and the rest of the world. For this role in particular, the ideal candidate is located near either of our offices and willing to work in a hybrid capacity. We would still consider 100% remote for exceptional candidates if they aren't located near an office. About the role Join the Cloud Infrastructure Team as a technical leader driving reliability, automation, and scalability across the systems running Sight Machine's platform. You'll operate at the intersection of classic SRE discipline which include IaC, CI/CD, observability, incident response and the emerging demands of running agentic AI systems in production: LLM gateways, agent orchestration, and the operational patterns that come with non-deterministic workloads. This is a senior level IC role. You'll help set and drive technical direction for infrastructure and reliability practices across teams, mentor senior engineers, and be a primary escalation point for the org's hardest systems problems while still being hands-on with code, infrastructure, and incidents. Success requires deep technical range, sound judgment on risk vs. customer impact, and the ability to influence architecture decisions across Development Engineering without formal authority. What You'll Actually Work On
We take pride in our self-starter culture where employees are enabled and encouraged to achieve their professional goals through leadership guidance, learning and development. Our philosophy is that careers are continuous journeys, and we dedicate time and offer resources so that employees can reach their full potential. Benefits + Perks We value you at and outside of work and know your loved ones are important. Our benefits are designed to support you and your family's health through life's expected and unexpected events. Our Benefits Include:
- Competitive Salary + Stock Options
- Health Care Coverage + Life Insurance + Health Savings Account + Flexible Spending
- Account (includes spouse + children)
- Flexible Vacation Policy
- Adaptable Working Schedule and Environment
- Our Perks Include:
- Casual Dress Attire
- Hybrid work flexibility
- Catered Lunches, Snacks and Beverages
- Commuter Savings Program
- Company Outings
- Designated Volunteering Hours + Group Volunteer Events
Sight Machine has offices in San Francisco, CA and Ann Arbor, Mi. We do have a remote-friendly culture with people based all around the US and the rest of the world. For this role in particular, the ideal candidate is located near either of our offices and willing to work in a hybrid capacity. We would still consider 100% remote for exceptional candidates if they aren't located near an office. About the role Join the Cloud Infrastructure Team as a technical leader driving reliability, automation, and scalability across the systems running Sight Machine's platform. You'll operate at the intersection of classic SRE discipline which include IaC, CI/CD, observability, incident response and the emerging demands of running agentic AI systems in production: LLM gateways, agent orchestration, and the operational patterns that come with non-deterministic workloads. This is a senior level IC role. You'll help set and drive technical direction for infrastructure and reliability practices across teams, mentor senior engineers, and be a primary escalation point for the org's hardest systems problems while still being hands-on with code, infrastructure, and incidents. Success requires deep technical range, sound judgment on risk vs. customer impact, and the ability to influence architecture decisions across Development Engineering without formal authority. What You'll Actually Work On
- Champion an agentic-AI-first engineering mindset: identify where AI-driven automation and agent-based tooling can replace manual toil, and hold that work to the same quality, testing, and reliability bar as any other production system
- Evolve reliability practices for meeting reliability SLO's, error budgets, drive incident postmortems to systemic (not just symptomatic) fixes, and lead reliability reviews for new services before they hit production
- Troubleshoot and resolve the org's most complex, cross-layer systems problems CI/CD, container orchestration, networking, OS, cloud resources, databases, and increasingly, agentic AI/LLM orchestration layers
- Design, build, and operate the infrastructure supporting agentic AI workloads, LLM gateway routing, agent orchestration frameworks, monitoring of non-deterministic/AI-driven services, and the operational tooling needed to run them reliably at scale
- Architect and instrument monitoring, alerting, and observability infrastructure for critical services, with an eye toward what "critical" means for AI-driven systems specifically
- Author and continuously improve operational runbooks and automation, increasingly incorporating agentic/AI-assisted tooling (e.g., automated triage, AI-assisted incident response) where it measurably reduces toil
- Design and build internal platforms and developer tooling that other engineers build on top of
- Participate in on-call coverage and help evolve the program as we scale including escalation paths and reducing avoidable pages through better automation
- Bring a startup mindset of daily engagement: staying close to what's breaking, what customers are hitting, and where the team needs help, even outside a formal ticket or rotation
- Mentor senior and mid-level engineers; act as a technical sounding board across teams
- Proactively identify and drive cross-team initiatives that improve stability, reliability, and availability, this is expected to be self-directed, not assigned
- Demonstrated experience designing, building, or operating agentic AI/LLM-based systems in production, held to the same quality-first, test-driven rigor as traditional infrastructure code, not just prototype-grade work
- Embody a quality-first and security-first culture in all that you do
- 10+ years of experience with Kubernetes/Docker in at least one top-tier cloud provider (Azure, GCP, AWS), including production-scale multi-tenant or multi-cluster environments
- 10+ years coding experience (Python, Go, Java, or similar) with a track record of building tools/platforms used by other engineers, not just scripts
- 10+ years with IaC and CI/CD tooling (Terraform/OpenTofu, FluxCD or similar GitOps tooling, Jenkins/GitHub Actions)
- Strong, provable Linux and networking fundamentals (TCP/IP and application-layer)
- Practical experience integrating or operating LLM/agentic AI systems in a production context this can be API-based orchestration, LLM gateways, or agent frameworks
- A track record of authoring technical documentation (design docs, ADRs, runbooks) that other engineers actually use
- Demonstrated mentorship of other engineers, without needing formal management authority to do it
- Strong bias for action over endless planning, hands-on, has made mistakes, learned from them, and can weigh risk vs. customer impact under pressure
- Clear, empathetic communicator, comfortable pushing back on architecture decisions across teams
- Operational experience with monitoring/alerting systems (Prometheus, Grafana, Loki, Sentry, Signoz or equivalents)
- Deep understanding of cloud performance, able to diagnose and resolve bottlenecks others can't
- Experience with elements of our current tech stack are a plus: Kubernetes, FluxCD, Terraform, Helm Charts, Prometheus, Elasticsearch, Python, Java, Kafka, Postgres, and Jenkins
- Previous experience or a keen interest in industrial IoT, analytics, or manufacturing a plus
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in San Francisco, CA vacancy
- ...our Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to-... ...expect us to hit our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team to: Own uptime,...SuggestedContract workWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours
- ...and actionable to everyone, everywhere. That everyone now includes AI agents. The Role: You'll be the infrastructure and reliability engineer on the Data Replication team - a full-stack product team running over 3 million sync jobs a week powering thousands of data...SuggestedFull timeWork at officeLocal areaFlexible hours
$204k - $306k
...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Manager, Site Reliability Engineering San Francisco, California Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of...SuggestedPermanent employmentFull timeWork at officeLocal areaWorldwideFlexible hours2 days per week$210k - $240k
...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $24...SuggestedFull time$175k - $250k
...00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance of... ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build...SuggestedFull timeRemote workRelocationRelocation package- ...company valued at $10 billion. We work in‑person five days a week in our new SanFrancisco headquarters. About the Role As a Site Reliability Engineer (SRE) at Mercor, you’ll own production reliability across our most critical systems, partnering directly with...
$155k - $222.6k
...global cloud platform. As a team of six engineers distributed across the US, Canada, and the... ...with a strong focus on automation, reliability, and operational excellence. We are one... ...Qualifications ~2+ years of experience in Site Reliability Engineering, DevOps,...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours- ...JOB DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will help Shipt understand where we can improve stability and reliability. There will be a focus on the intersection of systems...
$148.5k - $223.9k
...you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1... ...is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts...WorldwideWeekend work$260k - $300k
...agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding... ...than anyone expects. You will own both the production reliability of our user-facing products and the platform engineering that...- ...globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You...Temporary workWorldwide
- ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves...
- A leading technology firm is looking for a Manager to expand their Cloud Site Reliability team. The ideal candidate will have extensive Linux administration experience, a passion for automation, and be comfortable in a remote, diverse workplace. This position emphasizes...Remote work
- ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online... ...foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at...Permanent employmentShift work
- ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production...WorldwideHome officeFlexible hours
- ...Open Source LLM Gateway Engineer LiteLLM is an open-source LLM Gateway with 34K+ stars on GitHub and trusted by companies like NASA... ...expanding and seeking our 6th Engineer focused on owning reliability, performance, and infrastructure stability for the LiteLLM proxy...
- ...enterprise that runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging...WorldwideShift work
- ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You'll be building and operating the core systems that power agentic AI at scale. Your mission: keep...
$120k - $168.49k
...Site Reliability Engineer, Cloud Infrastructure About Quizlet At Quizlet, our mission is to help every learner achieve their outcomes in the most effective and delightful way. Our $1B+ learning platform serves tens of millions of students every month, including two-thirds...InternshipWork at office3 days per week- ...Senior Site Reliability Engineer Location: Global Remote / San Francisco • Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only...Full timeRemote work
$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...Temporary workImmediate startFlexible hoursShift work- ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical...Remote work
$166.9k - $225.9k
...SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close... ...with product engineering leads and staff engineers to define SLOs and SLIs... ...bring: ~6+ years of experience in Site Reliability Engineering, Cloud...Work at officeImmediate startWorldwideMonday to FridayFlexible hours- ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally...Contract workLocal area
$140k - $205k
...Senior Technology Site Reliability Engineer Cooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operationsteam. Position summary: The Senior Technology Site Reliability Engineer("SRE") is responsible for ensuring the reliability...Full timeTemporary workWork at officeFlexible hoursWeekend work$230k - $310k
...millions of daily users while enabling our engineering teams to ship fast. You'll own the... ...building automation and tooling that improves reliability and partnering with engineering to... ...What You'll Bring ~5+ years in site reliability engineering, DevOps, or systems...Full timeWork at officeWork from home$181k - $263k
...and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a...Work from homeFlexible hoursNight shift- ...Lead Site Reliability Engineer Stuut is transforming accounts receivable for B2B companies—making collections smarter and faster for companies that have historically relied on manual processes that are labor intensive and costly. Our platform is gaining traction with...Full timeFlexible hours
$220k - $235k
...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move faster, insights surface instantly, and agents push work forward, all with you in control. Whether you're buying or selling...Full timeContract workWork at office$210.8k - $272.8k
About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
Related searches
- staff security engineer San Francisco, CA
- project engineer assistant project manager San Francisco, CA
- assistant chief engineer San Francisco, CA
- staff data engineer San Francisco, CA
- senior staff engineer San Francisco, CA
- engineering aide San Francisco, CA
- software engineer staff San Francisco, CA
- assistant engineer San Francisco, CA
- assistant engineering manager San Francisco, CA
- technology administrator San Francisco, CA



