Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Staff Site Reliability Engineer

$232k - $263k

Obsidian

Sr. Staff Site Reliability Engineer As a Sr. Staff SRE at Obsidian, you will define and drive the company‑wide reliability vision for a complex, multi‑tenant SaaS platform serving enterprise and financial customers. You will operate as a strategic partner to DevOps and Platform Engineering leadership, shaping a unified reliability strategy that scales across the organization. Your core mandate: ensure Obsidian detects, diagnoses, and communicates system issues before customers are impacted—consistently and predictably. This is a hands‑on technical role that involves architecting and leading the implementation of systems that handle real‑world complexity, including upstream SaaS dependencies, sparse and noisy signals, and mission‑critical enterprise workloads. Key Responsibilities Reliability Strategy & Architecture – Define and lead long‑term reliability strategy across services. Establish end‑to‑end system visibility frameworks and guide architecture for observability, detection, and resilience. Cross‑Org Leadership – Partner across teams to embed reliability, standardize SLI/SLOs, and serve as a technical escalation expert. Detection & Observability – Build intelligent detection systems (anomaly detection, connector health models) and enable self‑service observability. Incident Management – Define and evolve a tiered incident communication strategy , improve response practices, and lead post‑mortems to strengthen reliability and customer trust. Execution – Contribute hands‑on to system design, monitoring, and debugging across distributed systems and data pipelines. Required Qualifications 5+ years in SRE, Production Engineering, or related roles. 3+ years operating at a senior or technical leadership level (Staff or equivalent scope). Deep expertise in AWS and/or GCP. Kubernetes and Helm. Observability stacks (Prometheus, Grafana, or equivalent). Proven experience designing and scaling reliability systems for multi‑tenant SaaS platforms. Strong debugging and systems thinking across distributed microservices and legacy systems. Demonstrated ability to lead initiatives that improve incident detection, response, and system resilience. Hands‑on engineering approach with a track record of building—not just configuring—reliability systems. Preferred Qualifications Experience in B2B SaaS serving enterprise or financial customers. Familiarity with third‑party SaaS connector architectures and ingestion patterns. Experience building anomaly detection or intelligent alerting systems. Experience designing customer‑facing status pages and incident communication frameworks. Why This Role Drive org‑wide reliability strategy. Own and build new detection & observability systems. Tackle complex distributed systems challenges. Safeguard critical infrastructure for financial customers. What Success Looks Like Issues caught and resolved before customer impact. Reliability is measurable and continuously improving. Teams self‑serve observability with scalable tools. Clear, proactive incident communication builds trust. Reliability becomes a competitive advantage. Employee Benefits Competitive compensation with equity and 401k. Comprehensive healthcare with dental and vision coverage. Flexible paid time off and paid holiday time off. 12 weeks of new parent or family leave. Personal and professional development resources. Pay Transparency Please note that the base pay range is a guideline and for candidates who receive an offer, the base pay will vary based on factors such as work location, as well as the knowledge, skills and experience of the candidate. In addition to a competitive base salary, this position is eligible for equity awards and may be eligible for sales commission or incentive compensation based on the role or function within the company. Equal Employment Opportunity Statement At Obsidian, we are proud to be an equal‑opportunity employer. We value diversity and hire for talent, passion, and compassion. In compliance with federal law, all persons hired will be required to submit satisfactory proof of identity and legal authorization. If you have a need that requires accommodation, please contact View email address on click.appcast.io. Information collected and processed as part of any job applications you choose to submit is subject to Obsidian’s Applicant Privacy Policy. Base Salary Range

$232,000 - $263,000 USD

#J-18808-Ljbffr Obsidian

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Sr. Staff Site Reliability Engineer in Palo Alto, CA vacancy
  • $90k - $135k

     ...containerization infrastructure. Troubleshoot operating system and engineering issues within our Linux environment. Collaborate on...  .... The Impact You Will Have: Improving the reliability and performance of our engineering environment. Focusing on... 
    Senior
    Local area
    Remote work

    Synopsys

    Mountain View, CA
    2 days ago
  •  ...professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of...  ...professionals and position yourself among the top echelon in site reliability.  As a Senior Lead Site Reliability Engineer at JPMorgan Chase... 
    Senior

    J.P. Morgan

    Palo Alto, CA
    3 days ago
  •  ...The Role We're looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of the production systems that power Nectar's platform. We run high-volume data ingestion pipelines and real-time AI agents on top of a fast... 
    Senior
    Remote work

    Nectar Social

    Palo Alto, CA
    3 days ago
  •  ...that keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity...  ...are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in... 
    Senior
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    1 day ago
  • $167.2k - $316.6k

     ...architecture, workflows, and technical specifications. Qualifications Bachelor’s or Master’s degree in Computer Science, Software Engineering, or equivalent combination of relevant education and experience. 8+ years of software development experience, particularly in... 
    Senior
    Immediate start
    Visa sponsorship
    Flexible hours

    Ford Motor Company

    Palo Alto, CA
    5 days ago
  • $137.77k - $194.59k

     ...distributed team of roughly 80 scientists and engineers building and operating Rubin's petascale...  ...Your role: \n You will own the reliability and robustness of Rubin Observatory's...  ...nature of this position, SLAC is open to on-site, hybrid, and remote work options. \n \... 
    Senior
    Remote work
    Flexible hours
    Night shift

    Stanford University

    Menlo Park, CA
    4 days ago
  • $224k - $257k

     ...center infrastructure, enabling the next giant leaps in human progress. The company invented the world’s first 3D-stacked photonics engine, Passage™, capable of connecting thousands to millions of processors at the speed of light in extreme-scale data centers for the... 
    Senior
    Full time
    Temporary work
    Flexible hours

    Lightmatter

    Mountain View, CA
    5 days ago
  •  ...join our small team focused on growth and productivity. The role involves scaling our platform and infrastructure while enhancing reliability and the overall developer experience. Ideal candidates will have strong expertise in distributed systems, cloud-native... 
    Senior
    Remote job

    BuildBuddy

    Palo Alto, CA
    2 days ago
  • Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident... 
    Senior

    Kody

    Palo Alto, CA
    5 days ago
  • $120k - $200k

    Sr Site Reliability Engineer (Prisma Access) 2 days ago Be among the first 25 applicants Job Description This role requires US Citizenship. Your Career Palo Alto Networks runs a large infrastructure and is one of the biggest GCP customers. As a Principal SRE, you'll... 
    Senior
    Rotating shift

    Palo Alto Networks

    Santa Clara, CA
    5 days ago
  • $232.75k - $325k

     ...in transformative projects. Together, let's push boundaries and achieve unparalleled success. As a Senior Director of Site Reliability Engineering at JPMorgan Chase within the I nfrastructure Platforms and Foundational Services (IPFS) team, you are deemed as a force... 
    Senior

    JPMorgan Chase

    Palo Alto, CA
    5 days ago
  • $179.2k - $268.8k

     ...sensors and compute systems, test operations, systems and safety engineering - all dedicated to redefining the relationship between people and their vehicles for millions of customers. As a Site Reliability Engineer on the team, you will be responsible for helping to... 
    Senior
    Permanent employment
    Full time
    Work at office
    Immediate start
    Visa sponsorship

    Latitude AI

    Palo Alto, CA
    5 days ago
  • $213k - $266.3k

     ...as reference platforms for future bring‑up, system integration, or system‑level validation. Collaborate with system performance engineers, hardware design and software teams to create comprehensive validation plans that surface key system‑level performance metrics, locate... 
    Senior
    Full time
    Local area

    Rivian and Volkswagen Group Technologies

    Palo Alto, CA
    5 days ago
  • $165.5k - $289.6k

    Sr Staff Site Reliability Engineer - Veza Full-time Employee Type: Regular Region: AMS - North America and Canada Work Persona: Flexible or Remote Veza is the pioneer in identity security, purpose-built to answer the fundamental question enterprises face: who can... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    5 days ago
  • $262k - $365k

    Senior Staff Site Reliability Engineer, AViD, YouTube Ads Mountain View, CA, USA; London, UK Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Mountain View, CA, USA; London, UK . Advanced Experience... 
    Senior

    Google

    Mountain View, CA
    5 days ago
  • $115.8k - $160k

    Tencent is seeking a skilled professional to manage and optimize PaaS products in North America. This role involves monitoring product stability, resolving technical issues, and applying tools like CI/CD to enhance operational efficiency. Candidates should have a Bachelor...
    Senior

    Tencent

    Palo Alto, CA
    5 days ago
  • $151.6k - $245.3k

     ...Site Reliability Engineer Palo Alto Networks runs a large hybrid infrastructure and is one of the largest GCP customers. As a Site Reliability Engineer, you will be part of a team supporting the services running on this infrastructure. This includes automation, architecture... 

    Palo Alto Networks

    Palo Alto, CA
    4 days ago
  • $205.5k - $278k

     ...Lead the ideation, technical development, and launch of innovative agentic AI features and experiences, partnering deeply with AI/ML engineers, data scientists, designers, and tax experts to deliver high-accuracy, trustworthy, and confidence-building solutions. Drive... 
    Senior
    Work at office
    3 days per week

    Intuit

    Mountain View, CA
    5 days ago
  • $189k - $232k

     ...Site Reliability Engineer Mountain View, US About EarnIn As one of the first pioneers of earned wage access, our passion at EarnIn is building products that deliver real-time financial flexibility for those with the unique needs of living paycheck to paycheck.... 
    Full time
    Work at office
    Local area
    2 days per week

    Earnin

    Mountain View, CA
    1 day ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability... 

    Chase

    Palo Alto, CA
    4 days ago
  • $170.7k - $300.2k

     ...business value. We do all this with an outstanding group of software engineers, data scientists, SRE/MLOps engineers and managers. We are...  ...and providing insight for the Infrastructure service reliability and availability through extensible services & platforms. Design... 
    Senior
    Relocation

    Apple Inc.

    Cupertino, CA
    5 days ago
  • $200k - $260k

     ...Site Reliability Engineering Lead Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. Your role is pivotal in ensuring our services meet stringent... 
    Work at office
    Home office

    Colorwave Inc

    Mountain View, CA
    5 days ago
  •  ...Lead Site Reliability Engineer Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within... 

    Chase

    Palo Alto, CA
    5 days ago
  • $217.57k - $260k

     ...description explicitly states otherwise, all roles are on-site five days per week at one of our offices in McLean, VA; Mountain...  ...which can be found here. Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high-scale infrastructure... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours
    Shift work

    ID.me

    Mountain View, CA
    4 days ago
  • $170k - $220k

     ...center infrastructure, enabling the next giant leaps in human progress. The company invented the world’s first 3D-stacked photonics engine, Passage™, capable of connecting thousands to millions of processors at the speed of light in extreme‑scale data centers for the... 
    Senior
    Full time
    Temporary work
    Flexible hours

    Lightmatter

    Mountain View, CA
    5 days ago
  • $205.5k - $278k

    Intuit is seeking a Senior Staff PM to develop the Small Business Health product, focused on providing insights to small business customers. You'll own the vision and strategy, set key engagement metrics, and ensure alignment across multiple teams. With a competitive compensation... 
    Senior

    Intuit

    Mountain View, CA
    5 days ago
  • Coupang in Mountain View, CA is seeking a Senior Staff Technical Program Manager to lead company-wide, technically complex programs...  ...drive strategy, architecture, and data-driven decisions across engineering and business teams to deliver critical platform capabilities. The... 
    Senior

    Coupang

    Mountain View, CA
    5 days ago
  • $165k - $190k

     ...DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high‑...  ...security platform Address complex challenges around scalability, reliability, observability, and cost efficiency Collaborate with Engineering... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    5 days ago
  • $100k - $200k

    OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    1 day ago
  • $115.8k - $160k

    Responsibilities Monitor and maintain Tencent Cloud's PaaS products in the North American region to ensure stability and reliability, resolving technical issues and mitigating risks to keep services operating smoothly in various technical scenarios. Utilize tools or platforms... 
    Relocation package

    Tencent

    Palo Alto, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Staff Site Reliability Engineer. Be the first to apply!