Director, Site Reliability Engineering
$197.3k - $313.7kSalesforce
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce. Job Title: Director, Site Reliability Engineering Location: New York, NY; San Francisco, CA; Dallas, TX About the Role We are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities. In this role, you will transform our SRE function—moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch. As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide. Key Responsibilities SRE Strategy and Leadership Define and execute the long-term strategy and roadmap for Site Reliability Engineering. Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures. Build and develop a high-performing team of site reliability and operations engineers. Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices. Translate business priorities and customer impact into clear reliability investments and engineering outcomes. Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs. Reliability Engineering Establish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services. Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems. Define what it means for a service to be operationally and observably ready for production. Develop readiness reviews and certification practices for high-impact services and launches. Drive improvements in system availability, performance, resiliency, and recovery. Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment. Observability Define an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry. Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards. Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities. Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure. Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes. Establish governance and measurement to assess adoption and effectiveness of observability standards. Incident Management and Operational Excellence Improve incident detection, response, mitigation, communication, and learning. Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation. Reduce mean time to detect, acknowledge, mitigate, and recover. Improve on-call practices, escalation paths, runbooks, and operational ownership. Establish blameless post-incident review practices that produce measurable engineering improvements. Identify recurring sources of operational toil and create plans to eliminate or automate them. Partner with engineering leaders to ensure actions from incidents are prioritized and completed. Automation and AI-Enabled Operations Develop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation. Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support. Build automation that reduces cognitive load and improves the speed and consistency of operational decisions. Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight. Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction. Cross-Functional Partnership Partner with engineering, DevOps, and business stakeholders. Influence teams that do not directly report into SRE and build shared accountability for production outcomes. Create clear service ownership models and operational expectations across teams. Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination. Communicate reliability posture, risks, trends, and investments to executive and technical audiences. Measures of Success Improved availability and reliability of critical services. Reduced time to detect, diagnose, mitigate, and recover from incidents. Increased percentage of services meeting observability and production-readiness standards. Reduced alert noise, operational toil, and manual incident-management activity. Increased adoption of service-level objectives and measurable reliability practices. Improved quality and completion rate of post-incident corrective actions. Increased automation across detection, triage, remediation, and operational workflows. Stronger ownership of production reliability across engineering teams. Leadership Attributes Engineering-first and automation-oriented. Comfortable challenging legacy operating models and assumptions. Able to move between technical detail and executive-level strategy. Outcome-focused, pragmatic, and data-driven. Builds trust through clarity, accountability, and strong partnership. Develops leaders and creates an inclusive, high-performance engineering culture. Treats incidents as opportunities to improve systems rather than assign blame. Brings urgency to operational risks while maintaining focus on sustainable solutions. Minimum Qualifications Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master’s degree or MBA preferred. 10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams Proven experience building or transforming a reliability or operational engineering organization. Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery. Experience establishing observability, incident-management, service-level objective, and production-readiness practices. Demonstrated ability to improve reliability through engineering and automation rather than process alone. Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems. Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring. Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholders Ability to balance immediate operational needs with long-term engineering transformation. Strong written, verbal, and executive communication skills. Preferred Qualifications Strong experience operating large-scale systems in AWS or another major cloud environment. Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry. Demonstrated experience implementing OpenTelemetry or common instrumentation standards. Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities. Experience applying AI, machine learning, or agent-based automation to operational workflows. Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering. Software engineering experience and the ability to engage deeply in architecture and design discussions. Solid background in supporting high-profile launches, events, or systems with significant customer and business impact.
*LI-Y
Unleash Your Potential When you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future — but to redefine what’s possible — for yourself, for AI, and the world. Accommodations If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form. Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options. Posting Statement Salesforce is an equal opportunity employer and maintains a policy of non-discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that’s inclusive, and free from discrimination. Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications – without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law. This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between. Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education. In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records. At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions. The typical base salary range for this position is $197,300 - $313,700 annually. In select cities within the San Francisco and New York City metropolitan area, the base salary range for this role is $237,700 - $344,700 annually. The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable.- ...Karsun Solutions, LLC is seeking a Site Reliability Manager to lead a multi-disciplinary team responsible for reliability, security, and platform lifecycle across AWS-based services. The role emphasizes collaboration, observability, and continuous improvement in a client...Suggested
- Fun is seeking a Business Development professional to own enterprise deals end-to-end and accelerate commercial growth at the frontier of on-chain payments. This role is primarily in-person at our Midtown, NYC headquarters with a Monday–Thursday collaboration rhythm and...SuggestedWork from home
$93k - $160k
...Palantir Technologies is seeking a Site Reliability Operations Analyst in New York, NY. In this role, you will streamline workflows and reduce friction in deployments. Your responsibilities include supporting deployments, removing roadblocks, and managing multiple challenges...Suggested- ...Federal Reserve Bank of New York is seeking an experienced Cloud AWS Support Reliability Engineer (SRE) to build and maintain scalable AWS infrastructure and CI/CD pipelines. The role emphasizes observability, security, and resilience across enterprise cloud platforms...Suggested
- Gusto is hiring for a hands-on operations role focused on running and sharpening day-to-day reconciliation and loss detection. You will own execution, resolve issues, and automate manual parts using AI and data tooling. You’ll gain end-to-end understanding of money flow...Suggested
- ...optimize production infrastructure across CI/CD, cloud deployments, and security. You will collaborate with our internal product and engineering teams to keep services scalable, secure, and highly available. The role emphasizes GitHub Actions, Terraform, Vercel, AWS core...Remote work
$160k - $180k
...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience...$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating complex operational...$130k - $170k
...NBC Universal is looking for a Staff Software Engineer (SRE Lead) in New York, NY. This role involves overseeing day-to-day operations of SAP BTP CPI applications, managing incidents, leading offshore support teams, and ensuring high system performance. Candidates should...Remote work- ...researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At... ...new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform and customer...
$104.9k - $174.7k
...Management. You can learn more about LexisNexis Risk at the link below, About the Role: We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...Full timeWork at officeLocal areaRemote workFlexible hours- ...Karsun Solutions in the DMV area is seeking a Site Reliability Manager to ensure reliability, scalability, and performance of our systems. You will lead a team focusing on Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate has 1...
- ...A dynamic fintech company in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability of its fintech products. The role focuses on end-to-end monitoring of customer journeys, incident management, and collaboration with various teams. Candidates...
$141k - $208k
...be a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability...Local areaRemote workHome officeFlexible hours- ...human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You...Full timeWork at office
- ...A financial technology company based in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability and health of their fintech products. The role involves monitoring customer journeys, APIs, and incident management, requiring 5+ years of experience...
- ...Komodor, a remote-first company, is seeking a Solutions Engineer to connect customer business initiatives to the Komodor platform, understanding developers, DevOps and Incident response teams working with Kubernetes. You will identify customer pain points and communicate...Remote work
$100k - $250k
...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial... ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics...Local area- ...Applications Deployment Responsible for reliability and support of Container Platform on-... ...Perform blameless RCA, partner with engineering and operation teams across the... ...Additional Skills : Automation Process Engineer,Site Reliability Engineer,Full Stack DeveloperThis...
$123k - $165k
...Site Reliability Engineer II Our engineering fleet is a horizontal set of teams providing engineering services across the organization. Our specific team provides reliability engineering and operational support to backend service development teams. Technology is...- ...self-healing, deployment/rollback automation). Establish reliability standards: SLOs/SLIs, error budgets, production readiness reviews... ..., and release risk controls. Performance and reliability engineering: capacity planning, load/performance analysis, resilience...
$150k - $175k
...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we're guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed...Remote work- ...strategy sessions with other Optum Teams Require 2+ years of experience with Terraform Require 2+ years of experience with DevOps Solution Architect, DevOps, or System Engineer certification in one or more public cloud providers Terraform certification #J-18808-Ljbffr...
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership...
$189k - $283.6k
...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...of accountability * A strong desire to perform and grow as an engineer * 5+ years of software development experience Technologies...Full timeLocal areaRemote workRelocation packageFlexible hoursShift work- .... No one coasts. If you're driven by impact, pace, and raising the bar. This is the place. The Role As a Staff Site Reliability Engineer you'll play a lead role on the founding SRE team at our new NYC engineering hub. You'll own multi-team reliability and infrastructure...Work at office
$140k - $165k
...deployed services and infrastructure components, empowering cloud engineering teams to move fast without sacrificing stability is essential... ...CI/CD pipelines to ensure cloud software changes are deployed reliably and efficiently. Own and manage developer self-service...Full timeRemote work$169.88k
...breadth of their business needs. It delivers end-to-end services and solutions leveraging strengths from strategy and design to engineering, all fueled by its market leading capabilities in AI, generative AI, cloud and data, combined with its deep industry expertise and...Full timeLocal area$160k - $190k
The Director of Corporate Partnerships will own and build Sollis Health’s B2B revenue channel. This role will start as a team of one, with... ...Sollis for immediate access to ER-trained medical teams, on-site labs and imaging, expedited specialist appointments, and care navigation...Full timeContract workWork experience placementImmediate start$120k - $165k
...communities. This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family responsible for... ...Overview The Wealth Management Production Management Site Reliability Engineer position is a highly visible/...Temporary workWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director, Site Reliability Engineering. Be the first to apply!
- chief engineer New York, NY
- engineering director New York, NY
- project engineer assistant project manager New York, NY
- principal network engineer New York, NY
- senior director engineering New York, NY
- director of product engineering New York, NY
- director data engineering New York, NY
- senior chief engineer New York, NY
- hotel chief engineer New York, NY
- principal developer New York, NY

