Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Akka

Job Description

Job Description

Description

Akka is the platform for building and running AI agents and distributed systems at scale. Our actor-model runtime powers agent and workflow systems for companies like Manulife, John Deere, Capital One, and CERN. We also run Akka Automated Operations, our managed platform for customer workloads on dedicated, BYOC, and BYOK8s deployments, backed by extensive availability guarantees and certifications.

We're hiring staff-level SREs to help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions for how we deploy, observe, and secure the platform, and help shape the standards the team builds against. It's not a team that watches someone else's dashboards, and it's not a role with a long runway: we expect you to be contributing meaningfully within your first few weeks .

The core of the platform is a set of Kubernetes operators written in Go. Every customer service, route, datastore, secret, and region flows through a reconciler, and when something goes wrong in production the fix usually starts with an operator's logs and a custom resource's status. Customer clusters on AWS, Azure, and GCP are provisioned with Crossplane compositions and delivered with Flux; with Terraform at the core. Each region runs managed Postgres that both our control plane and customer workloads depend on. Everything we run is defined and changed through code.

What we're looking for
  • 7+ years in infrastructure or platform engineering, with production experience across AWS and Azure. GCP is a plus.
  • Experience operating Kubernetes controllers built on controller-runtime in production: CRDs, admission webhooks, finalizers, status conditions, and debugging a reconcile loop that isn't converging. You can read Go well enough to trace a reconciler to a root cause.
  • Infrastructure as code with real production work, Terraform, and Crossplane at the level of authoring Compositions and XRDs rather than only applying claims, delivered through Flux and Kustomize.
  • Operating managed Postgres (RDS, Cloud SQL, or Azure Flexible Server) in production, point-in-time restore, major-version upgrades, and moving a live database between instances inside a bounded outage window.
  • Observability at scale with the Prometheus operator, scrape and relabel configuration, cardinality and cost control, alert rules as code, and distributed tracing with OpenTelemetry. Running a long-term metrics store (Cortex, Mimir, Thanos) is a plus, not a requirement.
  • Securing Kubernetes clusters, service mesh (Linkerd or similar), OIDC/workload identity, mTLS, cert-manager for PKI including trust-anchor rotation, and secrets management via cloud KMS.
  • Production on-call experience, you've carried a pager and written up what happened afterward.
  • Skilled use of LLMs as a tool to sharpen your work, not to run on autopilot.
  • Strong written communication, we weigh this heavily in our process.

Nice to have
  • Sizing JVM services in containers, heap versus container limits, direct memory, GC behavior, and reading a heap dump. Our platform computes JVM flags per service, and getting it wrong shows up as OOMKilled.
  • Operating event-sourced systems, projection lag, offsets, replay semantics, and what a journal replay does to a read model. Our own control plane is event-sourced, and so are our customers' workloads.
  • Messaging or streaming systems (Kafka, Pub/Sub, or similar) at production scale.
  • Teleport or a similar access plane, managed as code.
  • Distributed, stateful, or actor-based systems (Akka, Erlang/OTP).
Benefits
  • Competitive salary with performance-based incentives.
  • Comprehensive health and wellness benefits.
  • Opportunities for professional development and continuous learning.
  • Flexible remote working environment.
  • Collaborative, inclusive, and innovative company culture.
  • A transparent, distributed work environment with a strong focus on work-life balance .
  • Challenging work that interacts with innovative applications used by millions.
  • A collaborative culture that attracts the "brightest minds" in the technology community.

Vacancy posted 11 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in San Francisco, CA vacancy
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering, this role will help... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    4 days ago
  •  ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here...  ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with... 
    Suggested
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling... 
    Suggested
    Flexible hours

    Baseten

    San Francisco, CA
    3 days ago
  •  ...the globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You... 
    Suggested
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    21 hours ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Suggested
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    2 days ago
  •  ...About the Role We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate, and maintain... 

    Alembic Limited

    San Francisco, CA
    2 days ago
  • $260k - $300k

     ...software agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding...  ...faster than anyone expects. You will own both the production reliability of our user-facing products and the platform engineering that... 

    Cognition AI

    San Francisco, CA
    1 day ago
  •  ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You'll be building and operating the core systems that power agentic AI at scale. Your mission: keep... 

    Blaxel, Inc

    San Francisco, CA
    1 day ago
  •  ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online...  ...foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at... 
    Permanent employment
    Shift work

    Arena AI

    San Francisco, CA
    2 days ago
  •  ...JOB DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will help Shipt understand where we can improve stability and reliability. There will be a focus on the intersection of systems... 

    BayOne Solutions

    San Francisco, CA
    21 hours ago
  •  ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical... 
    Remote work

    Specter Services LLC

    San Francisco, CA
    15 hours ago
  •  ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally... 
    Contract work
    Local area

    InterSources

    San Francisco, CA
    3 days ago
  • $230k - $310k

     ...millions of daily users while enabling our engineering teams to ship fast. You'll own the...  ...building automation and tooling that improves reliability and partnering with engineering to...  ...What you'll bring ~5+ years in site reliability engineering, DevOps, or systems... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    2 days ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    4 days ago
  •  ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    13 hours ago
  • $170k - $220k

     ...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes. Our innovative approach blends technology with deep legal expertise, making us a leader in our field. We go beyond surface-level... 
    Work at office
    Remote work
    Flexible hours

    Supio

    San Francisco, CA
    4 days ago
  •  ...enterprise that runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging... 
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    21 hours ago
  • $117k - $209.33k

     ...Full time 26WD99273 Job Requisition ID # 26WD99273 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    4 days ago
  •  ...design of information and operational support systems.  Required Skills/Qualifications: BS/MS degree in Computer Science, Engineering, or a related subject. Equivalent experience accepted.   Proven working experience in installing, configuring, and troubleshooting... 
    Full time
    Work experience placement
    Remote work
    Flexible hours
    San Francisco, CA
    more than 2 months ago
  • $250k

     ...Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves working... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $194k - $267k

     ...something more than once, automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    13 hours ago
  • $195k - $257.5k

     ...Staff Site Reliability Engineer Circle (NYSE: CRCL) is one of the world's leading internet financial platform companies, building the foundation of a more open, global economy through digital assets, payment applications, and programmable blockchain infrastructure.... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  •  ...Reserve and our economy, and we’re building a dynamic and diverse team for our future. We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role... 
    Full time
    Temporary work
    Part time
    Shift work

    Federal Reserve Bank

    San Francisco, CA
    3 days ago
  • $150k - $220k

     ...Manager, Site Reliability Engineer San Francisco, California, United States At Forge, we know our team is our greatest asset. As technology innovators in the private market, our vision is to deliver a richer future for everyone. We live that vision through our values... 
    Local area

    FORGE

    San Francisco, CA
    4 days ago
  • $181k - $263k

     ...supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a senior... 
    Full time
    Work from home
    Worldwide
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    21 hours ago
  • $200k - $260k

     ...Join the Cloud Infrastructure Team as a technical leader driving reliability, automation, and scalability across the systems running Sight...  ...and reliability practices across teams, mentor senior engineers, and be a primary escalation point for the org's hardest systems... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    Sight Machine

    San Francisco, CA
    2 days ago
  • $120.6k - $150.9k

     ...Staff Site Reliability Engineer (SRE) We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative impact across WEX's platform reliability and operational excellence. This... 
    Flexible hours

    WEX

    San Francisco, CA
    1 day ago
  •  ...Job: Staff Site Reliability Engineer (SRE) Location: San Francisco, CA Job Responsibilities As our Staff SRE, you'll be the primary expert responsible for our entire compute ecosystem. Your key responsibilities will include: As a Staff SRE, you... 

    United IT Solutions

    San Francisco, CA
    2 days ago
  • $61k - $101k

     ...Salary: $61,000 - 101,000 per year Requirements: We expect formal training or certification in site reliability engineering, plus 3+ years of hands-on experience. We want strong familiarity with SRE culture and the practical application of reliability principles... 
    Full time

    J.P. Morgan

    San Francisco, CA
    3 days ago
  • $221.2k - $300k

     ...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering ~ link Copy link corporate_fare Google place San Francisco, CA, USA Advanced Experience owning outcomes and decision making, solving ambiguous... 
    Full time
    Work at office

    Google Inc.

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!