Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Salesforce

Senior Engineering Role at Salesforce

Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn't a buzzword — it's a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.

Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service availability and ready to swiftly repair any service-impacting issues. Five days a week, 24 hours a day, in a follow-the-sun model with weekend oncall, the Site Reliability team keeps the Salesforce cloud and our customers protected.

As an SRE, you will be a technical leader of the team driving Salesforce's operational resilience by engineering solutions that blend automation, observability, and AI-powered platforms. You will not only respond to incidents but proactively design systems that prevent them, applying software engineering principles to operations to reduce toil and improve reliability at scale. By leveraging cutting-edge software engineering practices within SRE function and AI-driven insights, you will help transform how services are built, monitored, and operated — ensuring that Salesforce delivers always-on, high-performance experiences to customers worldwide.

Build and run reliable, scalable, and efficient systems by applying software engineering principles to operations. Our mission is to ensure services are highly available, performant, and resilient — while continuously improving the balance between operational work and engineering innovation.

  • Reliability as the Priority: Ensure that systems meet defined Service Level Indicators (SLIs) and Service Level Objectives (SLOs), using error budgets to guide engineering and release decisions.
  • Engineering for Operations: Apply software engineering practices — automation, monitoring, self-healing systems — to eliminate toil and improve operational efficiency.
  • Incident Management: Lead the coordinated response to incidents as an Incident Commander, drive fast recovery (low TTR), and ensure lasting improvements through blameless postmortems.
  • Continuous Improvement: Identify and remove sources of toil, enhance observability, and optimize systems to reduce Time to Detect (TTD) and Time to Restore (TTR).
  • Collaboration with Development: Partner with product and engineering teams early in the lifecycle to design, build, and operate systems that are reliable by default.
  • Long-Term Focus: Leverage AI-driven automation to eliminate manual workflows, enabling the team to focus on complex problem-solving and strategic innovation while reducing operational overhead to less than 20% of capacity.

What You'll Actually Be Doing:

  • Lead incident detection, response, and resolution—driving root cause analysis, postmortems, and proactive measures to ensure high uptime, rapid recovery, and prevention of future issues.
  • Lead post-incident reviews, drive systemic fixes through corrective actions, and ensure customer-facing services maintain peak performance and reliability.
  • Understanding of AI/ML concepts applied to operations (e.g., anomaly detection, predictive analysis).
  • Independently drive the design and implementation of complex automation platforms, self-healing systems, and AI-powered operational tooling using durable workflow engines (Temporal, Airflow, Argo Workflows).
  • Architect and build production-grade observability solutions — monitoring, logging, alerting, and tracing systems — that enable proactive detection and autonomous remediation.
  • Design and implement AI/ML-powered operations tools including anomaly detection systems, predictive analysis pipelines, intelligent runbook automation, and prompt-engineered operational agents (MCP-based).
  • Drive optimization of system performance, reliability, and cost-effectiveness through proactive monitoring and tuning.
  • Ensuring that work carried out by the Site Reliability team is executed in such a way as to comply with the company's internal compliance policy and directives.
  • Identifying opportunities and driving the creation of comprehensive technical epics that include well-defined problem statements, detailed project and implementation documentation, and clearly measurable business outcomes aligned with team objectives.
  • Provide technical coaching to junior team members through pair programming, design reviews, and code reviews — helping grow their skills and knowledge.
  • Collaborate with engineering and product teams to define and uphold SLAs/SLOs, driving improvements in service reliability and customer experience.
  • Build and ship high-quality, production-grade software using modern engineering practices, with AI as a core part of your development workflow by pushing the boundaries of AI development tools to deliver secure, optimized, and high-quality code.
  • Design and orchestrate complex systems where AI agents integrate seamlessly into human workflows, driving efficiency and innovation at scale.
  • Critically evaluate code (Human or AI-generated) for correctness, quality, security, and performance
  • Contribute to building and maintaining the shared system context, an explicit repository of system designs, constraints, and standards that enables AI to operate accurately and reliably.

You're Our Person If You Have:

  • 5+ years of experience in systems engineering and software engineering for large-scale, internet-facing services.
  • Hands-on expertise with containerized architectures (Docker, Kubernetes) and orchestration platforms.
  • Strong knowledge of distributed systems and Linux/Unix internals, with experience tuning performance and troubleshooting at scale.
  • Familiarity with large-scale internet service architectures (DNS, Load Balancing, caching, etc.).
  • Proven proficiency in Python and Go (GoLang) with strong software engineering practices (testing, code review, CI/CD).
  • Production experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog, or similar)
  • Solid background in incident management, including on-call participation, root cause analysis, and postmortem practices.
  • Strong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, and capacity planning.
  • Hands-on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows, or similar) for building durable automation pipelines.
  • Experience applying AI/ML to operations — including anomaly detection, predictive analysis, LLM-based automation, and prompt engineering to build intelligent operational agents and workflows.
  • Excellent communication skills with demonstrated ability to lead during high-pressure incidents, present technical designs to leadership, and mentor junior engineers.
  • Track record of mentoring and technically coaching other engineers.
  • Ability to work in a 24/7 global operations model, managing multiple priorities under time-sensitive conditions.
  • Growth mindset with curiosity to explore new technologies and drive continuous improvement.
  • A demonstrated, genuine AI-first approach to engineering. Using AI to move faster, build fluency across the stack, and contribute well beyond your core specialty.
  • Experience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflows
  • Advanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production-ready.
  • A related technical degree required.

Even Better If You Have:

  • Experience with AI agent frameworks, MCP (Model Context Protocol), or building LLM-powered operational tools.
  • Contributions to open-source reliability/observability tooling.
  • AWS/GCP professional-level certifications.
  • Prior experience in SRE organizations supporting multi-cloud or hyperscale environments.
  • Python and Go proficiency for systems-level tooling.
  • Experience with chaos engineering and game day exercises.

Unleash Your Potential

When you join Salesforce, you'll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we'll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future — but to redefine what's possible — for yourself, for AI, and the world.

Accommodations

If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form.

Please

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    5 days ago
  • We are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering... 
    Senior

    Robert Half

    San Francisco, CA
    5 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    6 days ago
  • $190.8k - $267.1k

     ...while helping Reddit grow its business. The reliability of our Ads systems directly impacts...  ...Reliability team partners closely with Ads Engineering teams to improve reliability,...  ...advertising ecosystem.We're looking for a Staff Site Reliability Engineer who will define and... 
    Senior
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    Reddit

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  • $165k - $227k

     ...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    5 days ago
  •  ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ...ownership, working together to build scalable, reliable, and secure products that empower...  ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Senior
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    5 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    5 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    5 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    7 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations... 
    Senior
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  • $167.7k - $245.2k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    4 days ago
  •  ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'... 
    Senior

    Alembic

    San Francisco, CA
    7 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    5 days ago
  • $170k - $220k

     ...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes. Our innovative approach blends technology with deep legal expertise, making us a leader in our field. We go beyond surface-level... 
    Senior
    Work at office
    Remote work
    Flexible hours

    Supio

    San Francisco, CA
    5 days ago
  •  ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    17 hours ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    5 days ago
  • $167.7k - $245.2k

     ...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    7 days ago
  • $181k - $263k

     ...and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a... 
    Senior
    Full time
    Work from home
    Worldwide
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    1 day ago
  • $250k

     ...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $232k - $319k

     ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...  ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $175k - $250k

     ...00.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    6 days ago
  • $15k

     ...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    3 days ago
  • $139.76k - $287.75k

     ...to grow their business.We are seeking a Senior Site ReliabilityEngineer to help operate,...  ...will be instrumental in advancing the reliability, scalability, automation, observability...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Senior
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    4 days ago
  • $174.92k - $209.91k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We’re proud...  ...integrate our teams, systems, and career sites.About the RoleFivetran is building data... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    7 days ago
  • $174.92k - $209.91k

     ...High-Performance Engineer For Site Reliability Engineering Team Fivetran is building data pipelines to power the modern data stack for thousands of companies. Fivetran is looking for a high-performance, experienced engineer to be a part of a team of Site Reliability... 
    Senior
    Full time
    Work at office
    Remote work

    dbt Labs

    Oakland, CA
    2 days ago
  • $262k - $364k

    Lead a team of software/systems engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...Engineering.Experience with machine learning infrastructure.Site Reliability Engineering (SRE) combines software and systems engineering... 
    Senior

    Google

    San Bruno, CA
    4 days ago
  • $300k

     ...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    4 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Temporary work

    TextNow

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!