Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, Site Reliability Engineering

$197.3k - $313.7k

Salesforce.Com Inc

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn't a buzzword - it's a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You're in the right place! Agentforce is the future of AI, and you are the future of Salesforce. Job Title: Director, Site Reliability Engineering Location: New York, NY; San Francisco, CA; Dallas, TX About the Role We are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities. In this role, you will transform our SRE function-moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch. As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide. Key Responsibilities SRE Strategy and Leadership Define and execute the long-term strategy and roadmap for Site Reliability Engineering. Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures. Build and develop a high-performing team of site reliability and operations engineers. Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices. Translate business priorities and customer impact into clear reliability investments and engineering outcomes. Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs. Reliability Engineering Establish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services. Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems. Define what it means for a service to be operationally and observably ready for production. Develop readiness reviews and certification practices for high-impact services and launches. Drive improvements in system availability, performance, resiliency, and recovery. Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment. Observability Define an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry. Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards. Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities. Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure. Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes. Establish governance and measurement to assess adoption and effectiveness of observability standards. Incident Management and Operational Excellence Improve incident detection, response, mitigation, communication, and learning. Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation. Reduce mean time to detect, acknowledge, mitigate, and recover. Improve on-call practices, escalation paths, runbooks, and operational ownership. Establish blameless post-incident review practices that produce measurable engineering improvements. Identify recurring sources of operational toil and create plans to eliminate or automate them. Partner with engineering leaders to ensure actions from incidents are prioritized and completed. Automation and AI-Enabled Operations Develop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation. Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support. Build automation that reduces cognitive load and improves the speed and consistency of operational decisions. Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight. Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction. Cross-Functional Partnership Partner with engineering, DevOps, and business stakeholders. Influence teams that do not directly report into SRE and build shared accountability for production outcomes. Create clear service ownership models and operational expectations across teams. Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination. Communicate reliability posture, risks, trends, and investments to executive and technical audiences. Measures of Success Improved availability and reliability of critical services. Reduced time to detect, diagnose, mitigate, and recover from incidents. Increased percentage of services meeting observability and production-readiness standards. Reduced alert noise, operational toil, and manual incident-management activity. Increased adoption of service-level objectives and measurable reliability practices. Improved quality and completion rate of post-incident corrective actions. Increased automation across detection, triage, remediation, and operational workflows. Stronger ownership of production reliability across engineering teams. Leadership Attributes Engineering-first and automation-oriented. Comfortable challenging legacy operating models and assumptions. Able to move between technical detail and executive-level strategy. Outcome-focused, pragmatic, and data-driven. Builds trust through clarity, accountability, and strong partnership. Develops leaders and creates an inclusive, high-performance engineering culture. Treats incidents as opportunities to improve systems rather than assign blame. Brings urgency to operational risks while maintaining focus on sustainable solutions. Minimum Qualifications Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master's degree or MBA preferred. 10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams Proven experience building or transforming a reliability or operational engineering organization. Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery. Experience establishing observability, incident-management, service-level objective, and production-readiness practices. Demonstrated ability to improve reliability through engineering and automation rather than process alone. Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems. Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring. Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholders Ability to balance immediate operational needs with long-term engineering transformation. Strong written, verbal, and executive communication skills. Preferred Qualifications Strong experience operating large-scale systems in AWS or another major cloud environment. Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry. Demonstrated experience implementing OpenTelemetry or common instrumentation standards. Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities. Experience applying AI, machine learning, or agent-based automation to operational workflows. Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering. Software engineering experience and the ability to engage deeply in architecture and design discussions. Solid background in supporting high-profile launches, events, or systems with significant customer and business impact. *LI-Y Unleash Your Potential When you join Salesforce, you'll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we'll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Accommodations If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form. Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates' resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options. Posting Statement Salesforce is an equal opportunity employer and maintains a policy of non-discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that's inclusive, and free from discrimination. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications - without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law. This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between. Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education. In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records.At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions. The typical base salary range for this position is $197,300 - $313,700 annually. In select cities within the San Francisco and New York City metropolitan area, the base salary range for this role is $237,700 - $344,700 annually.The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable. #J-18808-Ljbffr

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Director, Site Reliability Engineering in New York, NY vacancy
  • $120k - $142k

     ...on making journalism so good that it’s worth paying for. Mission Overview & Responsibilities: At The New York Times, our Site Reliability Engineering (SRE) team is central to how we design, test, and operate the systems that support our most critical customer... 
    Suggested
    Local area
    Flexible hours

    The New York Times

    New York, NY
    4 days ago
  • $151k - $297k

    As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA... 
    Suggested
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    6 hours ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Suggested
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    3 days ago
  • $167.7k - $245.2k

     ...requiring approximately 2 days per week on-site at Cisco offices in either San Francisco...  ...AI agents behave as intended, improving reliability and reducing risks. This unified...  ...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and continuously... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    New York, NY
    2 days ago
  • $165k - $241.4k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Suggested
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    2 days ago
  • $138.1k - $198.2k

     ...more intuitive with technology that simply works.  The SRE Engineering Enablement Team supports our CI Platforms, Developer Environments...  ...Our customers are all engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter of our engineering... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    New York, NY
    2 days ago
  • $141k - $216.6k

     ...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical... 
    Work experience placement
    Work at office

    Axon

    New York, NY
    1 day ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    4 days ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the... 
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    4 days ago
  • $110k - $120k

     ...largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading financial services and healthcare technology company based on... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    New York, NY
    6 hours ago
  • $123k - $165k

    Job Posting Title:Site Reliability Engineer IIReq ID:10143234Job Description:Department/Group OverviewOur engineering fleet is a horizontal set of teams providing engineering services across the organization. Our specific team provides reliability engineering and operational... 
    Full time

    Hulu

    New York, NY
    3 days ago
  •  ...developers on how to make things better. We collectively strive to build and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.ResponsibilitiesEKS & Karpenter Management: Manage, upgrade, and autoscale... 
    For contractors

    Varo Money

    New York, NY
    1 hour ago
  • $120k - $200k

     ...PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery...  ...practices (e.g., Chaos Engineering, resilience testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)
    Overseas

    Comrise

    New York, NY
    6 hours ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad... 
    Shift work

    JP Morgan Chase

    New York, NY
    4 days ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership... 

    JP Morgan Chase

    New York, NY
    6 hours ago
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  • $194k - $267k

     ...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    2 days ago
  • $131k - $164k

    Position OverviewWe are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation frameworks, to join our global Infrastructure & Operations team. This role is a hands-on senior engineering position responsible... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Diligent

    New York, NY
    3 days ago
  • $194k - $267k

     ...in on this mission. If you are too, let's talk.The TeamThe Site Reliability team is dedicated to architecting and owning the foundational...  ...durable, automated systems that maximize platform reliability and engineering velocity.The ideal candidate is someone who enjoys analyzing... 
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    1 day ago
  • $150k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people...  ...markets.Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience,... 
    Full time
    Temporary work
    Part time

    Goldman Sachs

    New York, NY
    2 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    6 hours ago
  • $186.9k - $267.7k

     ...requiring approximately 2 days per week on-site at Cisco offices in either San Francisco...  ...AI agents behave as intended, improving reliability and reducing risks. This unified...  ...and control.As a Staff Site Reliability Engineer (SRE), you will provide technical leadership... 
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    New York, NY
    3 days ago
  • Guide and shape the future of technology at a globally recognized firm, driven by pride in ownership.As a Senior Manager of Site Reliability Engineering at JPMorgan Chase within the Corporate Investment Bank, Markets team, you are the non-functional requirement owner and... 
    Bank staff
    Shift work

    JP Morgan Chase

    New York, NY
    6 hours ago
  • Site Reliability Engineer - Equity Trading PlatformLocation: New York | Practice Area: Capital Markets - Technology & Engineering | Type: PermanentKeep critical equity trading platforms resilient, reliable, and ready for the markets.The RoleWe are seeking a highly motivated... 
    Permanent employment
    Work at office
    Weekend work
    Afternoon shift

    Capco

    New York, NY
    2 days ago
  • $130k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people...  ...digital possibility? Begin your journey here.Securities Frontline Site Reliability Engineers (SREs) play a critical role in our fast-paced... 
    Full time
    Temporary work
    Part time
    Immediate start

    Goldman Sachs

    New York, NY
    1 day ago
  • $195k - $275k

     ...Management, Capital Markets Application & Data Services, Deployment Planning & Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the... 
    Temporary work
    Work at office
    Worldwide
    Night shift

    Morgan Stanley

    New York, NY
    1 day ago
  • $150k - $190k

    Senior Site Reliability Engineer, VPAt Morgan Stanley, we advise, originate, trade, manage and distribute capital for governments, institutions and individuals, and always do so with a standard of excellence. We are a leading global financial services firm that conducts... 
    Temporary work
    Worldwide
    Flexible hours
    Weekend work

    Morgan Stanley

    New York, NY
    3 days ago
  • $150k - $160k

    Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs... 
    Work at office
    Local area

    Haymarket Media Group

    New York, NY
    3 days ago
  • $140k - $215k

     ...operate at the intersection of our Core Platform and Embedded Reliability charters: building the foundational libraries, services, and...  ...product group depends on, while embedding directly with product engineering teams and their leadership to drive reliability outcomes at... 
    Full time
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    New York, NY
    2 days ago
  •  ...also has offices in New York, NY, Miami, FL, Zurich, Switzerland and Lisbon, Portugal. About the Role At Luma, our Site Reliability Engineer (SRE) team keeps our platform reliable, secure, and lightning fast. They own everything from AWS infrastructure and Kubernetes... 

    Luma Financial Technologies

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, Site Reliability Engineering. Be the first to apply!