Senior Site Reliability Engineer
$148.5k - $223.9kSalesforce.Com Inc
Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service availability and ready to swiftly repair any service-impacting issues. Five days a week, 24 hours a day, in a follow-the-sun model with weekend oncall, the Site Reliability team keeps the Salesforce cloud and our customers protected.
The Experience
As an SRE, you will be a technical leader of the team driving Salesforce’s operational resilience by engineering solutions that blend automation, observability, and AI-powered platforms. You will not only respond to incidents but proactively design systems that prevent them, applying software engineering principles to operations to reduce toil and improve reliability at scale. By leveraging cutting-edge software engineering practices within SRE function and AI-driven insights, you will help transform how services are built, monitored, and operated — ensuring that Salesforce delivers always-on, high-performance experiences to customers worldwide.
Build and run reliable, scalable, and efficient systems by applying software engineering principles to operations. Our mission is to ensure services are highly available, performant, and resilient — while continuously improving the balance between operational work and engineering innovation.
Reliability as the Priority : Ensure that systems meet defined Service Level Indicators (SLIs) and Service Level Objectives (SLOs), using error budgets to guide engineering and release decisions.
Engineering for Operations : Apply software engineering practices — automation, monitoring, self-healing systems — to eliminate toil and improve operational efficiency.
Incident Management : Lead the coordinated response to incidents as an Incident Commander, drive fast recovery (low TTR), and ensure lasting improvements through blameless postmortems.
Continuous Improvement : Identify and remove sources of toil, enhance observability, and optimize systems to reduce Time to Detect (TTD) and Time to Restore (TTR).
Collaboration with Development : Partner with product and engineering teams early in the lifecycle to design, build, and operate systems that are reliable by default.
Long-Term Focus : Leverage AI-driven automation to eliminate manual workflows, enabling the team to focus on complex problem-solving and strategic innovation while reducing operational overhead to less than 20% of capacity.
What You’ll Actually Be Doing:
Lead incident detection, response, and resolution—driving root cause analysis, postmortems, and proactive measures to ensure high uptime, rapid recovery, and prevention of future issues.
Lead post-incident reviews, drive systemic fixes through corrective actions, and ensure customer-facing services maintain peak performance and reliability.
Understanding of AI/ML concepts applied to operations (e.g., anomaly detection, predictive analysis).
Independently drive the design and implementation of complex automation platforms, self-healing systems, and AI-powered operational tooling using durable workflow engines (Temporal, Airflow, Argo Workflows).
Architect and build production-grade observability solutions — monitoring, logging, alerting, and tracing systems — that enable proactive detection and autonomous remediation.
Design and implement AI/ML-powered operations tools including anomaly detection systems, predictive analysis pipelines, intelligent runbook automation, and prompt-engineered operational agents (MCP-based).
Drive optimization of system performance, reliability, and cost-effectiveness through proactive monitoring and tuning.
Ensuring that work carried out by the Site Reliability team is executed in such a way as to comply with the company’s internal compliance policy and directives.
Identifying opportunities and driving the creation of comprehensive technical epics that include well-defined problem statements, detailed project and implementation documentation, and clearly measurable business outcomes aligned with team objectives.
Provide technical coaching to junior team members through pair programming, design reviews, and code reviews — helping grow their skills and knowledge.
Collaborate with engineering and product teams to define and uphold SLAs/SLOs, driving improvements in service reliability and customer experience.
Build and ship high-quality, production-grade software using modern engineering practices, with AI as a core part of your development workflow by pushing the boundaries of AI development tools to deliver secure, optimized, and high-quality code.
Design and orchestrate complex systems where AI agents integrate seamlessly into human workflows, driving efficiency and innovation at scale.
Critically evaluate code (Human or AI-generated) for correctness, quality, security, and performance
Contribute to building and maintaining the shared system context, an explicit repository of system designs, constraints, and standards that enables AI to operate accurately and reliably.
You’re Our Person If You Have:
- 5+ years of experience in systems engineering and software engineering for large-scale, internet-facing services.
- Hands-on expertise with containerized architectures (Docker, Kubernetes) and orchestration platforms.
- Strong knowledge of distributed systems and Linux/Unix internals, with experience tuning performance and troubleshooting at scale.
- Familiarity with large-scale internet service architectures (DNS, Load Balancing, caching, etc.).
- Proven proficiency in Python and Go (GoLang) with strong software engineering practices (testing, code review, CI/CD).
- Production experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog, or similar)
- Solid background in incident management, including on-call participation, root cause analysis, and postmortem practices.
- Strong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, and capacity planning.
- Hands-on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows, or similar) for building durable automation pipelines.
- Experience applying AI/ML to operations — including anomaly detection, predictive analysis, LLM-based automation, and prompt engineering to build intelligent operational agents and workflows.
- Excellent communication skills with demonstrated ability to lead during high-pressure incidents, present technical designs to leadership, and mentor junior engineers.
- Track record of mentoring and technically coaching other engineers.
- Ability to work in a 24/7 global operations model, managing multiple priorities under time-sensitive conditions.
- Growth mindset with curiosity to explore new technologies and drive continuous improvement.
- A demonstrated, genuine AI-first approach to engineering. Using AI to move faster, build fluency across the stack, and contribute well beyond your core specialty.
- Advanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production-ready.
- A related technical degree required.
Even Better If You Have:
- Experience with AI agent frameworks, MCP (Model Context Protocol), or building LLM-powered operational tools.
- Contributions to open-source reliability/observability tooling.
- AWS/GCP professional-level certifications.
- Prior experience in SRE organizations supporting multi-cloud or hyperscale environments.
- Python and Go proficiency for systems-level tooling.
- Experience with chaos engineering and game day exercises.
Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records.
In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link:
Unleash Your Potential
When you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love.
At Salesforce, we strive to create an accessible and inclusive experience for all candidates.
If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodation Request Form.
Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.
Equal Opportunity Statement.
Salesforce is an equal opportunity employer and maintains a policy of non-discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that’s inclusive, and free from discrimination. Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications - without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law. This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between. Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education.
At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions.
The typical base salary range for this position is $148,500 - $223,900 annually. In select cities within the San Francisco and New York City metropolitan area, the base salary range for this role is $178,900 - $246,000 annually.
The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable.
We can’t wait to meet you!
Join our Talent Community and be the first to know about open roles, career tips, events happening near you, and much more.
#J-18808-Ljbffr- ...Discover exciting DevOps job opportunities and connect with 28,396 DevOps professionals. The Senior Site Reliability Engineer role at Jobicy is designed for experienced professionals who are passionate about enhancing system reliability and operational efficiency....SeniorRemote workFlexible hours
$182.8k - $247.3k
...mission to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement- ...developer-tooling company whose product is used by engineering teams at thousands of software companies for... ...commitments and the SRE team is a senior, well-resourced group of nine. As Senior SRE you will lead reliability initiatives across the platform — from defining...Senior
- ...We need a Senior SRE to ensure Xident's verification platform runs at 99.99% uptime.... ...SLOs, incident response, and production reliability for a system that processes millions of... ...structured logging Implement chaos engineering practices to proactively identify failure...SeniorRemote work
- ...As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture...SeniorWork experience placement
$180k - $230k
...Power Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the...SeniorWork at officeLocal areaImmediate startRemote work3 days per week- ...SQS, Cloudflare, GitHub Actions, PostgreSQL, Redis/BullMQ, Node.js/NestJS, Datadog, TypeScript, React, SQL Position: Senior Site Reliability Engineer Engagement period: Ongoing Interview timeline: ASAP Interview process: 1) CV review 2) Interview with our CTO...SeniorContract workImmediate start
- ...Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform...SeniorLocal area
- ...democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of...SeniorFull timeTemporary workWork at officeWorldwideFlexible hours
- ...infra has to match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi... ...-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions...Senior
$150k - $220k
...Senior Site Reliability EngineerJob detailsDepartment / EngineeringRemoteFull-time$150,000 USD - $220,000 USD## About UsMetaRouter is a customer... ...architecture.## About The RoleAs a Senior Site Reliability Engineer, you own significant pieces of our infrastructure and...SeniorFull timeRemote work$150.4k - $277.6k
...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long... ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced...SeniorRelocationDay shift$180k - $200k
...Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options... ...equities traders rely on every market day. As our first Senior Site Reliability Engineer, you'll define what reliability means at tastytrade,...SeniorWork at office3 days per week- ...to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise.The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers...SeniorWork experience placementFlexible hours
$110k - $145k
...is global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role...SeniorFlexible hours$104k - $178k
## Sr. Site Reliability Engineer IApply: Hybrid: NYC Global HQ: Full time: Posted 12 Days Ago: JR00000779# ****Who We Are****DV is the leader in digital performance solutions, helping our advertiser and agency partners Verify the quality of their digital campaigns, Optimise...SeniorFull time$135.2k - $181.2k
...enhance electrical, mechanical, and sensor-based systems to ensure reliability and performance. Configure, calibrate, and validate... ...professional development, including an interest in emerging data engineering tools and methodologies. Preferred Qualifications: ~8+...SeniorWorldwide- ...Onshape is hiring a Principal Software Engineer (SRE) in a hybrid role based in Boston, MA. You will lead reliability initiatives, shape strategy, and act as a technical authority to ensure the platform is fast, resilient, and scalable for customers. You will drive...
- ...services, including Gen Insurance and Engine by Gen, delivering trusted digital and... ...consumers at scale. As a Principal Site Reliability Engineer, you will define and lead the... ...excellence strategy for our platform. This is a senior technical leadership role for an...
$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting...Work experience placementWork at officeRemote work$160k - $180k
...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical...Flexible hours$194k - $237k
## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on:... ...Purpose**The Principal Site Reliability Engineer partners with development teams by designing... ...techniques.* Be a thought leader: a senior point of expertise on site reliability...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$165k - $185k
...more at tandemdiabetes.com**A DAY IN THE LIFE:**The Principal Site Reliability Engineer (SRE) is responsible for the reliability, availability, and... ...coverage that works across time zones. Participates as a senior escalation tier for high-severity incidents.* Builds...Permanent employmentContract workLocal areaRemote workFlexible hoursShift work- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers... ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them...WorldwideHome officeFlexible hours
- ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and...Full timeRemote work
$107.9k - $195.05k
...The Digital Sector at Leidos currently has an opening for a Site Reliability Engineer (SRE) / Senior Cloud Engineer to work in our Baltimore, Maryland office. This is an exciting opportunity to use your experience helping the Center for Medicare and Medicaid Services...Contract workWork at office$110k - $145k
...operations. You will liaise with product and engineering teams to ensure applications and... ...feedback loop for platform and product reliability. The ideal candidate is a solutions-oriented... ...experience as a platform engineer, site reliability engineer, systems engineer...Work experience placement$114k - $148k
...Total compensation is based on experience, skills, and location using objective, job-related criteria. Summary As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant, and highly available. If...Work experience placement- ...personalized care faster. We are building AI agents to support the full arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and maintain the tooling, alerts and incident-response playbooks that keep...
$140k - $195k
...Improve reliability, observability, service health, incident response, and operational readiness. CodeVertex works across data... ..., secure systems, and operational clarity matter. The Site Reliability Engineer role helps turn business needs into reliable execution, whether...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Eastern, KY
- senior manager customer operations Eastern, KY
- senior product manager mobile Eastern, KY
- senior software engineer ruby on rails Eastern, KY
- sr finance manager Eastern, KY
- sr marketing manager Eastern, KY
- senior customer service Eastern, KY
- senior business manager Eastern, KY
- senior account executive Eastern, KY
- senior account director Eastern, KY

