Staff Platform & Reliability Engineer
Full-time
Interface AI
Responsibilities
- Define customer-facing SLIs and SLOs, operate error budgets, and govern release decisions through reliability targets.
- Own regional failover, disaster recovery, RTO/RPO definitions, third-party dependency resilience, and tested recovery procedures.
- Build and operate GitOps-based progressive delivery with automated analysis, one-click rollback, drift detection, and change auditing.
- Own AWS infrastructure-as-code, Kubernetes, service mesh, capacity, cost, and scale across the platform.
- Lead paging, severity policies, incident-command rotations, status-page automation, post-mortems, and customer SLA reporting.
- Build metrics, logging, distributed tracing, dashboards, alerts, and burn-rate alerting as code.
- Develop AI-native operational automation, safe runbooks, self-healing workflows, and golden-path service templates.
- Own cloud and cluster security, including IAM least privilege, secrets management, admission and network policy, image signing, SBOMs, scanning, and threat-detection feeds.
- Produce infrastructure compliance evidence for SOC 2 Type II, NCUA/FFIEC examinations, and GLBA reviews.
- Protect AI systems against prompt injection, tool-based data exfiltration, tenant-isolation failures, endpoint abuse, and unsafe AI-assisted engineering workflows.
Requirements
- Senior-most hands-on individual contributor experience running production systems with real uptime commitments and pager participation.
- Deep production Kubernetes on AWS, including service mesh and GitOps delivery across many services.
- Experience managing infrastructure-as-code at multi-account scale and reshaping large existing estates.
- Experience building SLO and error-budget practices, burn-rate alerting, and multi-region or disaster-recovery capabilities with tested failover.
- Daily-practice security experience with least-privilege IAM, secrets management, admission and network policies, software supply-chain controls, and auditor-facing evidence.
- Daily fluency with frontier AI tools such as Claude Code and Cursor, including judgment about agent operating boundaries.
- Strong programming ability in TypeScript and/or Python plus Bash, with clear technical and regulatory-facing writing.
- BS or BA in Computer Science required; an MS or PhD is preferred.
- Preferred experience includes real-time voice or telephony systems, streaming or analytical data platforms, durable workflow engines, chaos engineering, regulated industries, LLM threat modeling, cloud cost engineering, and LLM spend attribution.
- Must be San Francisco-based, committed to working onsite, and able to participate in on-call coverage.
Benefits
- 100% paid health, dental, and vision care.
- 401(k), financial wellness perks, daily meals, commuter benefit, monthly wellness stipend, and mental health, wellness, and family benefits.
- Claude Enterprise and frontier AI tools for every employee.
- Discretionary PTO and paid parental leave.
- Onsite work at the brand-new 21st-floor San Francisco office at 44 Montgomery.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Platform & Reliability Engineer in San Francisco, CA vacancy
$140k - $230k
...logistics system, delivering critical supplies quickly and reliably. Today, Zipline operates on four continents, makes a... ...lives with minimal maintenance. As a Senior or Staff Electrical Hardware Reliability Engineer focused on power electronics, you will serve as the reliability...SuggestedLocal area- ...The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports...SuggestedFull timeWork at officeRemote workWorldwide
- ...Amsterdam.Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems... ...shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and...SuggestedPermanent employmentWork experience placementWork at officeLocal area
- ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage...Suggested
$113.4k - $162k
...down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is aboutimpactatscale.You...Suggested$200.7k - $250.9k
...washed away in a flood in 1942, the Royal Engineers rebuilt it. Then it washed away again... ...team at Mercury has primarily built platform-level constructs that product teams adopt... ...identify opportunities for improvements in reliability/observability/performance/preparedness...$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$120k - $168.49k
...Site Reliability Engineer, Cloud Infrastructure About Quizlet At Quizlet, our mission is to help every learner achieve their outcomes in the most effective and delightful way. Our $1B+ learning platform serves tens of millions of students every month, including two-thirds...InternshipWork at office3 days per week- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering... ...You will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes...
- ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves...
$167.7k - $245.2k
...Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital... ...Observability portfolios. Your Impact As a Senior Site Reliability Engineer (SRE), you will lead the design and management of large-scale...Full timeTemporary workWork experience placementWork at officeLocal areaFlexible hours$127k - $249k
...The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering... ...deployment pipelines and global observability platforms.Within Platform Engineering, the Fabric... ...a pivotal role in engineering the reliable, globally connected, multi-cloud network...Local areaRemote workWorldwideFlexible hours- ...innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business problems with simple and...
$180.1k - $278.7k
...Staff Infrastructure Reliability EngineerThe Staff Infrastructure Reliability Engineer is responsible for the technical leadership of Redfin's production database and storage systems. They will work with the database team manager and other database and storage engineers...Minimum wageImmediate start$194k - $237k
...employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering... ...Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering,...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours- ...Amperesand in the San Francisco Bay Area seeks a Sr. Manager of Reliability Engineering to lead test operations across California and Nevada. This role designs new reliability labs, drives failure analysis, and ensures products meet stringent grid and hyperscaler requirements...
- ...troubleshooting, providing rubric-based written feedback. This role requires hands-on Kubernetes expertise in EKS/GKE/AKS or self-managed clusters, with strong scripting in Go, Python, or TypeScript, and ability to document findings clearly for engineering #J-18808-Ljbffr
$200k - $240k
...high performing SRE to join our growing Platform team. In this role, you will focus on performance... .... You will collaborate closely with engineering leadership, product managers, and cross-... ...the importance of performant and reliable systems ~ Education - Ideally looking...Work at officeImmediate start3 days per week$350k
...with leading AI companies and infrastructure providers to build reliable, high-performance platforms supporting next-generation AI workloads. This opportunity is for a Staff Site Reliability Engineer to lead the reliability of large-scale GPU infrastructure, covering...$204k - $306k
...you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure... ...Office.The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and... ...tooling.As the Manager of Infrastructure Platform and Shared Services, you will oversee multiple...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week- ...the leading AI contracting platform that transforms agreements into... ...a strategic, high-output Staff SRE to define the future of... ...cloud platform and champion engineering excellence across Ironclad.... ...strategic direction for the Site Reliability Engineering team and our...Full timeContract workWork at office
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the...
- ...Job Description Description Akka is the platform for building and running AI agents and... ...guarantees and certifications. We're hiring staff-level SREs to help run and evolve that... ..., working alongside the senior engineers already on the team. You'll contribute to...Remote workFlexible hours
$114.3k - $235.32k
About Pinterest:Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and... ...can trust to grow their business.We are seeking a Site Reliability Engineer to help operate, scale, and continuously improve a cloud-native...Work at officeLocal areaRelocationRelocation package- ...the only unified payments and financial platform for global businesses. Powered by our unique... ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ..., working together to build scalable, reliable, and secure products that empower businesses...Temporary workLocal areaWorldwide
$140k - $180k
...delivering critical supplies quickly and reliably. Today, Zipline operates on four... ...Senior Mechanical Hardware Reliability Engineer, you will own the reliability of critical... ...systems across Zipline's autonomous delivery platform. Your primary scope may include systems...Local area$117k - $209.33k
...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure... ...closely with product engineering, security, compliance, platform, and infrastructure teams to ensure services are reliable, scalable...Full timeFor contractors$250k
...cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference... ...United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud...Full timeRemote work- ...design of information and operational support systems. Required Skills/Qualifications: BS/MS degree in Computer Science, Engineering, or a related subject. Equivalent experience accepted. Proven working experience in installing, configuring, and troubleshooting...Full timeWork experience placementRemote workFlexible hours
$150k
...Job Description Job Description About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture,...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Platform & Reliability Engineer. Be the first to apply!
Related searches
- engineering aide San Francisco, CA
- technology administrator San Francisco, CA
- staff security engineer San Francisco, CA
- assistant engineering manager San Francisco, CA
- project engineer assistant project manager San Francisco, CA
- assistant chief engineer San Francisco, CA
- staff data engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- senior staff engineer San Francisco, CA
- staff engineer San Francisco, CA


