Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$148.5k - $223.9k

100 Salesforce, Inc.

About Salesforce

Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.

Job Details

Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in SanFrancisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service availability and ready to swiftly repair any service-impacting issues. Five days a week, 24 hours a day, in a follow-the-sun model with weekend on‑call, the Site Reliability team keeps the Salesforce cloud and our customers protected.

The Experience

As an SRE, you will be a technical leader of the team driving Salesforce’s operational resilience by engineering solutions that blend automation, observability, and AI‑powered platforms. You will not only respond to incidents but proactively design systems that prevent them, applying software engineering principles to operations to reduce toil and improve reliability at scale. By leveraging cutting‑edge software engineering practices within SRE function and AI‑driven insights, you will help transform how services are built, monitored, and operated — ensuring that Salesforce delivers always‑on, high‑performance experiences to customers worldwide.

Build and run reliable, scalable, and efficient systems by applying software engineering principles to operations. Our mission is to ensure services are highly available, performant, and resilient — while continuously improving the balance between operational work and engineering innovation.

Reliability as the Priority

Ensure that systems meet defined Service Level Indicators (SLIs) and Service Level Objectives (SLOs), using error budgets to guide engineering and release decisions.

Engineering for Operations

Apply software engineering practices — automation, monitoring, self‑healing systems — to eliminate toil and improve operational efficiency.

Incident Management

Lead the coordinated response to incidents as an Incident Commander, drive fast recovery (low TTR), and ensure lasting improvements through blameless postmortems.

Continuous Improvement

Identify and remove sources of toil, enhance observability, and optimize systems to reduce Time to Detect (TTD) and Time to Restore (TTR).

Collaboration with Development

Partner with product and engineering teams early in the lifecycle to design, build, and operate systems that are reliable by default.

Long‑Term Focus

Leverage AI‑driven automation to eliminate manual workflows, enabling the team to focus on complex problem‑solving and strategic innovation while reducing operational overhead to less than 20% of capacity.

What You’ll Actually Be Doing
  • Lead incident detection, response, and resolution—driving root cause analysis, postmortems, and proactive measures to ensure high uptime, rapid recovery, and prevention of future issues.
  • Lead post‑incident reviews, drive systemic fixes through corrective actions, and ensure customer‑facing services maintain peak performance and reliability.
  • Understanding of AI/ML concepts applied to operations (e.g., anomaly detection, predictive analysis).
  • Independently drive the design and implementation of complex automation platforms, self‑healing systems, and AI‑powered operational tooling using durable workflow engines (Temporal, Airflow, Argo Workflows).
  • Architect and build production‑grade observability solutions — monitoring, logging, alerting, and tracing systems — that enable proactive detection and autonomous remediation.
  • Design and implement AI/ML‑powered operations tools including anomaly detection systems, predictive analysis pipelines, intelligent runbook automation, and prompt‑engineered operational agents (MCP‑based).
  • Drive optimization of system performance, reliability, and cost‑effectiveness through proactive monitoring and tuning.
  • Ensuring that work carried out by the Site Reliability team is executed in such a way as to comply with the company’s internal compliance policy and directives.
  • Identifying opportunities and driving the creation of comprehensive technical epics that include well‑defined problem statements, detailed project and implementation documentation, and clearly measurable business outcomes aligned with team objectives.
  • Provide technical coaching to junior team members through pair programming, design reviews, and code reviews — helping grow their skills and knowledge.
  • Collaborate with engineering and product teams to define and uphold SLAs/SLOs, driving improvements in service reliability and customer experience.
  • Build and ship high‑quality, production‑grade software using modern engineering practices, with AI as a core part of your development workflow by pushing the boundaries of AI development tools to deliver secure, optimized, and high‑quality code.
  • Design and orchestrate complex systems where AI agents integrate seamlessly into human workflows, driving efficiency and innovation at scale.
  • Critically evaluate code (Human or AI‑generated) for correctness, quality, security, and performance.
  • Contribute to building and maintaining the shared system context, an explicit repository of system designs, constraints, and standards that enables AI to operate accurately and reliably.
You’re Our Person If You Have
  • 5+ years of experience in systems engineering and software engineering for large‑scale, internet‑facing services.
  • Hands‑on expertise with containerized architectures (Docker, Kubernetes) and orchestration platforms.
  • Strong knowledge of distributed systems and Linux/Unix internals, with experience tuning performance and troubleshooting at scale.
  • Familiarity with large‑scale internet service architectures (DNS, Load Balancing, caching, etc.).
  • Proven proficiency in Python and Go (GoLang) with strong software engineering practices (testing, code review, CI/CD).
  • Production experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog, or similar).
  • Solid background in incident management, including on‑call participation, root cause analysis, and postmortem practices.
  • Strong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, and capacity planning.
  • Hands‑on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows, or similar) for building durable automation pipelines.
  • Experience applying AI/ML to operations — including anomaly detection, predictive analysis, LLM‑based automation, and prompt engineering to build intelligent operational agents and workflows.
  • Excellent communication skills with demonstrated ability to lead during high‑pressure incidents, present technical designs to leadership, and mentor junior engineers.
  • Track record of mentoring and technically coaching other engineers.
  • Ability to work in a 24/7 global operations model, managing multiple priorities under time‑sensitive conditions.
  • Growth mindset with curiosity to explore new technologies and drive continuous improvement.
  • A demonstrated, genuine AI‑first approach to engineering.
  • Using AI to move faster, build fluency across the stack, and contribute well beyond your core specialty.
  • Experience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflows.
  • Advanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production‑ready.
  • A related technical degree required.
Even Better If You Have
  • Experience with AI agent frameworks, MCP (Model Context Protocol), or building LLM‑powered operational tools.
  • Contributions to open‑source reliability/observability tooling.
  • AWS/GCP professional‑level certifications.
  • Prior experience in SRE organizations supporting multi‑cloud or hyperscale environments.
  • Python and Go proficiency for systems‑level tooling.
  • Experience with chaos engineering and game day exercises.
Unleash Your Potential

When you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love.

Accommodations

If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form. Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.

Posting Statement

Salesforce is an equal opportunity employer and maintains a policy of non‑discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that’s inclusive, and free from discrimination.

Equality Statement

Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications – without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law.

This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between.

Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education.

In the United States, compensation offered will be determined by factors such as location, job level, job‑related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity or benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link:

Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records. At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions.

Salary

The typical base salary range for this position is $148,500 - $223,900 annually. In select cities within the SanFrancisco and NewYork City metropolitan area, the base salary range for this role is $178,900 - $246,000 annually. The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable.

Company Culture

We’re Salesforce, the Customer Company, inspiring the future of business with AI + Data + CRM. Leading with our core values, we help companies across every industry blaze new trails and connect with customers in a whole new way. And, we empower you to be a Trailblazer, too — driving your performance and career growth, charting new paths, and improving the state of the world.

If you believe in business as the greatest platform for change and in companies doing well and doing good – you've come to the right place.

#J-18808-Ljbffr
Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
  •  ...US Corp. is seeking a Lead Site Reliability Engineer to spearhead our mission of delivering highly available and performant systems. With an average of over 12 years of industry experience, the successful candidate will bridge the gap between software development and systems... 
    Senior

    Axiom Pursuits

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  •  ...principles to see it in full.About the teamThe Engineering team at Airwallex is a diverse group of...  ..., working together to build scalable, reliable, and secure products that empower...  ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Senior
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    2 days ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $2... 
    Senior
    Full time

    Alembic Technologies

    San Francisco, CA
    3 days ago
  • $250k

     ...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 
    Senior

    Alembic Limited

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...000.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do Design... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    3 days ago
  • $164k - $205k

     ...ensuring high availability and performance Design intelligent alerting and observability systems Collaborate with engineering teams to embed reliability into the development lifecycle, shifting left on operational concerns Automate incident response workflows and... 
    Senior
    Work experience placement
    Summer holiday
    Live out
    Work at office
    Local area
    Flexible hours
    Shift work
    2 days per week

    SupportFinity

    San Francisco, CA
    4 days ago
  • $117k - $209.33k

     ...Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  •  ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and operations... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    4 days ago
  • $189k - $283.6k

     ...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics...  ...accountability ~ A strong desire to perform and grow as an engineer ~5+ years of software development experience... 
    Senior
    Full time
    Relocation package
    Flexible hours
    Shift work

    Block Inc

    San Francisco, CA
    4 days ago
  • $153k - $191.3k

     ...hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a...  ...deployments across operating environments, to guarantee the reliability, scalability, and availability of our services. To do this, you... 
    Senior
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office
    3 days per week

    Planet Labs PBC

    San Francisco, CA
    5 days ago
  • $220k - $235k

     ...are seeking a strategic, high‑output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role,...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... 
    Senior
    Full time
    Work at office

    Ironclad Inc

    San Francisco, CA
    4 days ago
  • $181k - $263k

    ## Senior Staff Site Reliability EngineerApplylocations: San Franciscotime type: Full timeposted on: Posted Yesterdayjob requisition id: JR01220...  ...support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability... 
    Senior
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    4 days ago
  •  ...management. We have become a multibillion‑dollar asset manager, and we have ambitious goals for the future. As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage... 
    Senior
    Local area

    The Voleon Group

    Berkeley, CA
    1 day ago
  • $300k

     ...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    1 day ago
  • $210.38k - $243.21k

    Manager, Site Reliability Engineer (Hybrid in South San Francisco)About the RoleWe are seeking an experienced and hands-on Site Reliability Engineering (SRE) Manager to lead our Site Operations and infrastructure initiatives. This role is responsible for ensuring the reliability... 

    Twist Bioscience

    San Francisco, CA
    3 hours ago
  • $173k - $230k

     ...employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering...  .... Establishes enterprise technical direction, develops senior technical leaders, and demonstrates impact well beyond systems... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    San Francisco, CA
    2 days ago
  •  ...Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply...  ...constraints they personally own. Engineering · Mid to Senior · Full-time · On-site · San Francisco, CA... 
    Full time

    Zof AI

    San Francisco, CA
    4 days ago
  •  ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built... 
    Relocation package

    Talentify.io

    San Francisco, CA
    2 days ago
  •  ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site) Responsibilities Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers. Perform regular capacity... 

    Methodic

    San Francisco, CA
    1 day ago
  • Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed...

    BEAM inc.

    San Francisco, CA
    1 day ago
  •  ...access to life-saving treatment. What We Look for in a Great Engineer Tool Proficiency: You are highly proficient with your tools...  ...high-velocity feature release while maintaining the highest reliability. DevX Support: Support Developer Experience (DevX) work to... 
    Work at office

    Latent

    San Francisco, CA
    2 days ago
  • $110k - $160k

     ...more. Base pay range $110,000.00/yr - $160,000.00/yr Site Reliability Engineer Fractal Analytics is a strategic AI partner to...  ...characteristic protected by federal, state or local laws. Seniority level ~ Mid-Senior level Employment type... 
    Hourly pay
    Full time
    Local area
    Relocation package
    Monday to Friday

    Fractal, Inc.

    San Francisco, CA
    3 days ago
  • $98.58k - $138.02k

     ...This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH. Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365’s cloud infrastructure and applications.... 
    Work at office

    Restaurant365

    San Francisco, CA
    1 day ago
  •  ...Sigma, Flow Traders, Tower Research, PDT Partners, SIG, and more. We're looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability of our platform, and help set the... 
    Full time
    Local area
    Remote work

    Databento

    San Francisco, CA
    9 hours ago
  •  ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 

    re-tool®

    San Francisco, CA
    1 day ago
  • $150k

     ...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our cloud infrastructure... 

    VantageScore®

    San Francisco, CA
    2 days ago
  •  ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.   As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad... 

    JPMorgan Chase & Co.

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!