Senior Site Reliability Engineer
$148.5k - $223.9k100 Salesforce, Inc.
About Salesforce
Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.
Job Details
Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in SanFrancisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service availability and ready to swiftly repair any service-impacting issues. Five days a week, 24 hours a day, in a follow-the-sun model with weekend on‑call, the Site Reliability team keeps the Salesforce cloud and our customers protected.
The Experience
As an SRE, you will be a technical leader of the team driving Salesforce’s operational resilience by engineering solutions that blend automation, observability, and AI‑powered platforms. You will not only respond to incidents but proactively design systems that prevent them, applying software engineering principles to operations to reduce toil and improve reliability at scale. By leveraging cutting‑edge software engineering practices within SRE function and AI‑driven insights, you will help transform how services are built, monitored, and operated — ensuring that Salesforce delivers always‑on, high‑performance experiences to customers worldwide.
Build and run reliable, scalable, and efficient systems by applying software engineering principles to operations. Our mission is to ensure services are highly available, performant, and resilient — while continuously improving the balance between operational work and engineering innovation.
Reliability as the Priority
Ensure that systems meet defined Service Level Indicators (SLIs) and Service Level Objectives (SLOs), using error budgets to guide engineering and release decisions.
Engineering for Operations
Apply software engineering practices — automation, monitoring, self‑healing systems — to eliminate toil and improve operational efficiency.
Incident Management
Lead the coordinated response to incidents as an Incident Commander, drive fast recovery (low TTR), and ensure lasting improvements through blameless postmortems.
Continuous Improvement
Identify and remove sources of toil, enhance observability, and optimize systems to reduce Time to Detect (TTD) and Time to Restore (TTR).
Collaboration with Development
Partner with product and engineering teams early in the lifecycle to design, build, and operate systems that are reliable by default.
Long‑Term Focus
Leverage AI‑driven automation to eliminate manual workflows, enabling the team to focus on complex problem‑solving and strategic innovation while reducing operational overhead to less than 20% of capacity.
What You’ll Actually Be Doing
- Lead incident detection, response, and resolution—driving root cause analysis, postmortems, and proactive measures to ensure high uptime, rapid recovery, and prevention of future issues.
- Lead post‑incident reviews, drive systemic fixes through corrective actions, and ensure customer‑facing services maintain peak performance and reliability.
- Understanding of AI/ML concepts applied to operations (e.g., anomaly detection, predictive analysis).
- Independently drive the design and implementation of complex automation platforms, self‑healing systems, and AI‑powered operational tooling using durable workflow engines (Temporal, Airflow, Argo Workflows).
- Architect and build production‑grade observability solutions — monitoring, logging, alerting, and tracing systems — that enable proactive detection and autonomous remediation.
- Design and implement AI/ML‑powered operations tools including anomaly detection systems, predictive analysis pipelines, intelligent runbook automation, and prompt‑engineered operational agents (MCP‑based).
- Drive optimization of system performance, reliability, and cost‑effectiveness through proactive monitoring and tuning.
- Ensuring that work carried out by the Site Reliability team is executed in such a way as to comply with the company’s internal compliance policy and directives.
- Identifying opportunities and driving the creation of comprehensive technical epics that include well‑defined problem statements, detailed project and implementation documentation, and clearly measurable business outcomes aligned with team objectives.
- Provide technical coaching to junior team members through pair programming, design reviews, and code reviews — helping grow their skills and knowledge.
- Collaborate with engineering and product teams to define and uphold SLAs/SLOs, driving improvements in service reliability and customer experience.
- Build and ship high‑quality, production‑grade software using modern engineering practices, with AI as a core part of your development workflow by pushing the boundaries of AI development tools to deliver secure, optimized, and high‑quality code.
- Design and orchestrate complex systems where AI agents integrate seamlessly into human workflows, driving efficiency and innovation at scale.
- Critically evaluate code (Human or AI‑generated) for correctness, quality, security, and performance.
- Contribute to building and maintaining the shared system context, an explicit repository of system designs, constraints, and standards that enables AI to operate accurately and reliably.
You’re Our Person If You Have
- 5+ years of experience in systems engineering and software engineering for large‑scale, internet‑facing services.
- Hands‑on expertise with containerized architectures (Docker, Kubernetes) and orchestration platforms.
- Strong knowledge of distributed systems and Linux/Unix internals, with experience tuning performance and troubleshooting at scale.
- Familiarity with large‑scale internet service architectures (DNS, Load Balancing, caching, etc.).
- Proven proficiency in Python and Go (GoLang) with strong software engineering practices (testing, code review, CI/CD).
- Production experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog, or similar).
- Solid background in incident management, including on‑call participation, root cause analysis, and postmortem practices.
- Strong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, and capacity planning.
- Hands‑on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows, or similar) for building durable automation pipelines.
- Experience applying AI/ML to operations — including anomaly detection, predictive analysis, LLM‑based automation, and prompt engineering to build intelligent operational agents and workflows.
- Excellent communication skills with demonstrated ability to lead during high‑pressure incidents, present technical designs to leadership, and mentor junior engineers.
- Track record of mentoring and technically coaching other engineers.
- Ability to work in a 24/7 global operations model, managing multiple priorities under time‑sensitive conditions.
- Growth mindset with curiosity to explore new technologies and drive continuous improvement.
- A demonstrated, genuine AI‑first approach to engineering.
- Using AI to move faster, build fluency across the stack, and contribute well beyond your core specialty.
- Experience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflows.
- Advanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production‑ready.
- A related technical degree required.
Even Better If You Have
- Experience with AI agent frameworks, MCP (Model Context Protocol), or building LLM‑powered operational tools.
- Contributions to open‑source reliability/observability tooling.
- AWS/GCP professional‑level certifications.
- Prior experience in SRE organizations supporting multi‑cloud or hyperscale environments.
- Python and Go proficiency for systems‑level tooling.
- Experience with chaos engineering and game day exercises.
Unleash Your Potential
When you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love.
Accommodations
If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form. Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.
Posting Statement
Salesforce is an equal opportunity employer and maintains a policy of non‑discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that’s inclusive, and free from discrimination.
Equality Statement
Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications – without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law.
This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between.
Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education.
In the United States, compensation offered will be determined by factors such as location, job level, job‑related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity or benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link:
Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records. At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions.
Salary
The typical base salary range for this position is $148,500 - $223,900 annually. In select cities within the SanFrancisco and NewYork City metropolitan area, the base salary range for this role is $178,900 - $246,000 annually. The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable.
Company Culture
We’re Salesforce, the Customer Company, inspiring the future of business with AI + Data + CRM. Leading with our core values, we help companies across every industry blaze new trails and connect with customers in a whole new way. And, we empower you to be a Trailblazer, too — driving your performance and career growth, charting new paths, and improving the state of the world.
If you believe in business as the greatest platform for change and in companies doing well and doing good – you've come to the right place.
#J-18808-Ljbffr- ...US Corp. is seeking a Lead Site Reliability Engineer to spearhead our mission of delivering highly available and performant systems. With an average of over 12 years of industry experience, the successful candidate will bridge the gap between software development and systems...Senior
$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind...SeniorFlexible hours- ...principles to see it in full.About the teamThe Engineering team at Airwallex is a diverse group of... ..., working together to build scalable, reliable, and secure products that empower... ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work...SeniorTemporary workLocal area
$210k - $240k
...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $2...SeniorFull time$250k
...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves...SeniorFull timeRemote work- ...About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the...Senior
$175k - $250k
...000.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance... ...scalability, performance, and reliability across environments. What You’ll Do Design...SeniorFull timeRemote workRelocationRelocation package$164k - $205k
...ensuring high availability and performance Design intelligent alerting and observability systems Collaborate with engineering teams to embed reliability into the development lifecycle, shifting left on operational concerns Automate incident response workflows and...SeniorWork experience placementSummer holidayLive outWork at officeLocal areaFlexible hoursShift work2 days per week$117k - $209.33k
...Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team...SeniorFull timeFor contractors- ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and operations...SeniorImmediate startRemote workWorldwide
$189k - $283.6k
...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...accountability ~ A strong desire to perform and grow as an engineer ~5+ years of software development experience...SeniorFull timeRelocation packageFlexible hoursShift work$153k - $191.3k
...hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a... ...deployments across operating environments, to guarantee the reliability, scalability, and availability of our services. To do this, you...SeniorFull timeTemporary workFor contractorsWork at officeLocal areaRemote workHome office3 days per week$220k - $235k
...are seeking a strategic, high‑output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role,... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...SeniorFull timeWork at office$181k - $263k
## Senior Staff Site Reliability EngineerApplylocations: San Franciscotime type: Full timeposted on: Posted Yesterdayjob requisition id: JR01220... ...support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability...SeniorWork from homeFlexible hoursNight shift- ...management. We have become a multibillion‑dollar asset manager, and we have ambitious goals for the future. As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage...SeniorLocal area
$300k
...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring...SeniorPermanent employment- ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely... ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale...Permanent employmentWork experience placementWork at officeLocal area
$210.38k - $243.21k
Manager, Site Reliability Engineer (Hybrid in South San Francisco)About the RoleWe are seeking an experienced and hands-on Site Reliability Engineering (SRE) Manager to lead our Site Operations and infrastructure initiatives. This role is responsible for ensuring the reliability...$173k - $230k
...employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering... .... Establishes enterprise technical direction, develops senior technical leaders, and demonstrates impact well beyond systems...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours- ...Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply... ...constraints they personally own. Engineering · Mid to Senior · Full-time · On-site · San Francisco, CA...Full time
- ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built...Relocation package
- ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site) Responsibilities Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers. Perform regular capacity...
- Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed...
- ...access to life-saving treatment. What We Look for in a Great Engineer Tool Proficiency: You are highly proficient with your tools... ...high-velocity feature release while maintaining the highest reliability. DevX Support: Support Developer Experience (DevX) work to...Work at office
$110k - $160k
...more. Base pay range $110,000.00/yr - $160,000.00/yr Site Reliability Engineer Fractal Analytics is a strategic AI partner to... ...characteristic protected by federal, state or local laws. Seniority level ~ Mid-Senior level Employment type...Hourly payFull timeLocal areaRelocation packageMonday to Friday$98.58k - $138.02k
...This role requires a hybrid work schedule based out of one of our office locations: Austin, TX; Irvine, CA; or Akron, OH. Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining Restaurant365’s cloud infrastructure and applications....Work at office- ...Sigma, Flow Traders, Tower Research, PDT Partners, SIG, and more. We're looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability of our platform, and help set the...Full timeLocal areaRemote work
- ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system... ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially...
$150k
...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our cloud infrastructure...- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer San Francisco, CA
- site reliability engineer remote San Francisco, CA
- site reliability engineer sre San Francisco, CA
- senior living director San Francisco, CA
- senior php developer remote San Francisco, CA
- senior manager customer operations San Francisco, CA
- senior support engineer San Francisco, CA
- senior product manager mobile San Francisco, CA
- senior software engineer ruby on rails San Francisco, CA
- sr finance manager San Francisco, CA



