Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Reliability Engineer: Scale Systems, Observe & Automate

Jobleads-US

A leading AI research company based in San Francisco is seeking experienced reliability engineers to scale their infrastructure and ensure system performance and reliability. This role involves collaborating with diverse teams to develop resilient systems and enhance operations. Candidates should have strong cloud proficiency, experience in containerization technologies, and a bachelor's degree in a related field. #J-18808-Ljbffr Jobleads-US

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Reliability Engineer: Scale Systems, Observe & Automate in San Francisco, CA vacancy
  • TwelveLabs is seeking a GTM Systems Engineer to design, build, and scale the revenue operations stack. You will own HubSpot-based workflows, Clay automation, and AWS co-sell pipelines, translating GTM strategy into reliable, automated systems that drive growth across sales... 
    Suggested

    Twelve-Labs

    San Francisco, CA
    1 day ago
  • $310k

     ...leading AI research organization is looking for a Software Engineer for their Platform Systems team in San Francisco. You will design and build systems for large-scale AI training workloads, focusing on reliability and performance. Ideal candidates should have a deep... 
    Suggested

    OpenAI

    San Francisco, CA
    1 day ago
  • $107k - $150k

     ...build sensors and tools for engineers, roboticists, and researchers...  ...digital device powered by one chip-scale laser array and one CMOS...  ..., optimize, and maintain our automated production test infrastructure...  ...hardware to our data systems.We treat manufacturing test infrastructure... 
    Suggested
    Work experience placement
    Local area
    Flexible hours

    Ouster

    San Francisco, CA
    1 day ago
  • $174.5k - $240k

     ...About this roleGTM Engineering builds and operates the intelligent systems, integrations, and automations that power Faire's...  ...accountable for the reliability, adoption, and continuous...  ...we ship is observable, reversible, and measured...  ....Equipped to scale: We invest in what... 
    Suggested
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San...  ...solutions that blend automation, observability, and AI-powered...  ...but proactively design systems that prevent them, applying...  ...reliability at scale. By leveraging cutting... 
    Suggested
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  •  ...looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence...  ...data scientists to build, automate, and maintain the...  ...workloads, and real-time analytics systems.This is a hands-on, high-... 

    Alembic

    San Francisco, CA
    8 hours ago
  •  ...designing, building, and running large-scale, distributed, fault-tolerant systems that power most of Varo's...  ...oriented mindset.We are an automation and observability focused team and we strive to automate...  ...platform that enables our engineers to accomplish their own goals instead... 
    For contractors

    Varo Money

    San Francisco, CA
    2 days ago
  •  ...embedded finance at a global scale.Proudly founded in...  ...next.About the teamThe Engineering team at Airwallex is a...  ...together to build scalable, reliable, and secure products...  ...incident response, observability, and automation across critical systems.Own team-level SLOs, runbooks... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    8 hours ago
  • $165k - $225.6k

     ...) team delivers the systems, tools, and services...  ...functions to drive scale, reliability, and innovation through...  ...Site Reliability Engineer OpportunityReporting...  ...With a strong focus on automation, testing, and...  ....Containerization & Observability: Strong hands-on experience... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    8 hours ago
  • $167.7k - $245.2k

     ...customers deploy at scale while also...  ...Collaboration, and Observability portfolios Your ImpactThe...  ...for talented engineers with a software...  ...available distributed systems in the cloud. You...  ...to ensure the reliability, performance and...  ....Drive and build automation enabling our... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within...  ...service mesh), and observability and alerting systems.The Fleet...  ...that ensure cluster reliability and security (e.g.,...  ...our infrastructure scales to support new use...  ...to resolutionPrefer automation over manual processes... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    2 days ago
  • $152.5k - $205k

     ...power trusted, internet-scale financial innovation....  ...forThe Site Reliability Engineer builds and maintains...  ...and operate production systems.You will also help teams...  ...appropriate reliability, observability, access controls,...  ...and software delivery automation across engineering teams... 
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  • $117k - $209.33k

     ...world? As a Senior Site Reliability Engineer at Autodesk, you can help...  ..., reliability practices, automation, and engineering standards...  ...experience operating production systems at scale, an automation-first...  ..., incident management, observability, resilience testing, and... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    1 day ago
  • $113.4k - $162k

     ...looking for motivated Site Reliability Engineer to own infrastructure...  ...is about impact at scale. You’ll shape how...  ...builds and operates its systems in an AI-first...  ...infrastructure and services.Automation & Infrastructure as...  ...and improve observability tools, logging, and monitoring... 
    Temporary work

    TextNow

    San Francisco, CA
    3 days ago
  • WorkOS is seeking a Systems Engineer to architect and automate our internal IT systems. You will design and implement...  ..., and build automation that scales with a fast-growing company. You will...  ...functional teams to reduce toil and improve reliability. #J-18808-Ljbffr WorkOS
    Remote job

    WorkOS

    San Francisco, CA
    5 days ago
  • $152.5k - $205k

     ...power trusted, internet-scale financial innovation....  ...for:As a Senior Site Reliability Engineer on Circle’s platform...  ...services and automation, developing reliable...  ...solving hard distributed-systems problems, taking ownership...  ....Define and evolve observability practices across metrics... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $249.5k - $273.5k

     ...advantage.Unlike legacy systems built to route and...  ...busywork through automation while seamlessly...  ..., Design, and Engineering leadership, you will...  ...from concept to reliable production systems...  ...and evaluation or observability capabilities.Build and Scale: Stay close to the... 
    Work at office

    Dialpad

    San Francisco, CA
    4 days ago
  • $195k - $257.5k

     ...power trusted, internet-scale financial innovation....  ...responsible for:As a Staff Site Reliability Engineer on Circle’s Platform...  ..., performance, and automation of distributed blockchain systems across multiple public...  ...operators, controllers, and observability tooling.Hands-on... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...StaffObservabilitySite Reliability Engineer with a specialty in...  ...comprehensive, scalable Observability Platform that enables...  ..., Python, or Ruby—to automate the deployment of...  ...complex distributed systems.Key ResponsibilitiesAutomated...  ...the deployment and scaling of observability... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  •  ...or Staff level Site Reliability Engineer to strengthen the reliability...  ...health, refining observability, and partnering with...  ...teams to build systems that perform consistently...  ...experience, a strong automation mindset, and a practical...  ...systems at scale.• Hands-on experience... 

    Robert Half

    San Francisco, CA
    1 day ago
  • $194k - $267k

     ...passion for solving large-scale automation, testing, and tuning problems...  ...Overview:The Site Reliability Engineer (SRE) will play a key role...  ...communication, security, and observability within the Kubernetes clusters...  ...troubleshoot, and resolve system issues related to... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $190k - $270k

     ...AI Infrastructure Engineer at Together, you...  ...services and production systems running smoothly....  ..., and mature automation to our operating...  ...availability, reliability and scalability,...  ...Kubernetes to enable scaling to a massive...  ...in monitoring and observability practicesKnowledge... 
    Full time

    Together AI

    San Francisco, CA
    1 day ago
  • $157k - $239k

     ...the adventure?As a Site Reliability Engineer with strong networking skills...  ...You keep them reliable, observable, and secure as we scale. You coordinate with IT...  ...), instrumenting the systems that prove it, managing...  ...the network as code, and automating away toil so on-call... 
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    3 days ago
  • $165k - $241.4k

     ...supporting customers in scaling deployments while...  ..., Collaboration, and Observability portfolios.Your ImpactWe...  ...skilled Senior Site Reliability Engineer (SRE) in Production...  ...available distributed systems in the cloud, collaborating...  ...platform reliability.Automate production operations... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    2 days ago
  •  ...to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation...  ...mission-critical enterprise systems. You'll work across networking, systems, automation, observability, and reliability engineering... 

    Alembic

    San Francisco, CA
    2 days ago
  • $160k - $194k

     ...built by pioneers in grid-scale batteries, energy...  ...The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work...  ...enhance system performance, reliability, and security to meet...  ...to develop deployment automation, observability, monitoring, and... 
    Full time
    Remote work
    Flexible hours

    Ad Verse

    San Francisco, CA
    8 hours ago
  •  ...About the Team The Scaling team is responsible for...  ...the architectural and engineering backbone of OpenAI’s...  ...and deliver advanced systems that support the deployment...  ..., scheduled, and observable. You’ll sit at the intersection...  ...-up, provisioning automation, fleet/cluster... 
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...intersection of product, systems, and cloud...  ...as the company scales. You’ll partner closely with engineers and leadership to...  ...developer workflows and observability, you’ll turn...  ...Strengthen the reliability of our production...  ..., templates, and automation that help teams ship... 
    Full time

    Midstream Health

    San Francisco, CA
    8 hours ago
  • $200k - $260k

     ...investor support, we're scaling fast and defining a...  ...a Software Engineer on the Site Reliability team at Harvey, you...  ...product, owning the systems that keep our platform...  ...50+ regions to automating mission-critical operations...  ...familiarity with observability tools (Datadog,... 
    Relocation package

    Harvey

    San Francisco, CA
    1 day ago
  •  ...organizations of all sizes to easily build, scale, and run modern applications by...  ..., providing a highly scalable yet observable system for customers and engineers. The Atlas Search product is...  ...-strength backend software and automation in complex codebases ~ Experience... 
    Full time
    Work at office
    Local area
    Worldwide

    Mongodb

    San Francisco, CA
    8 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Reliability Engineer: Scale Systems, Observe & Automate. Be the first to apply!