Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, Site Reliability Engineering

$197.3k - $313.7k

Salesforce

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.Job CategorySoftware EngineeringJob DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.Job Title: Director, Site Reliability EngineeringLocation: New York, NY; San Francisco, CA; Dallas, TXAbout the RoleWe are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities.In this role, you will transform our SRE function—moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch.As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide.Key ResponsibilitiesSRE Strategy and LeadershipDefine and execute the long-term strategy and roadmap for Site Reliability Engineering.Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures.Build and develop a high-performing team of site reliability and operations engineers.Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.Translate business priorities and customer impact into clear reliability investments and engineering outcomes.Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.Reliability EngineeringEstablish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services.Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.Define what it means for a service to be operationally and observably ready for production.Develop readiness reviews and certification practices for high-impact services and launches.Drive improvements in system availability, performance, resiliency, and recovery.Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment.ObservabilityDefine an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards.Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities.Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure.Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes.Establish governance and measurement to assess adoption and effectiveness of observability standards.Incident Management and Operational ExcellenceImprove incident detection, response, mitigation, communication, and learning.Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation.Reduce mean time to detect, acknowledge, mitigate, and recover.Improve on-call practices, escalation paths, runbooks, and operational ownership.Establish blameless post-incident review practices that produce measurable engineering improvements.Identify recurring sources of operational toil and create plans to eliminate or automate them.Partner with engineering leaders to ensure actions from incidents are prioritized and completed.Automation and AI-Enabled OperationsDevelop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation.Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support.Build automation that reduces cognitive load and improves the speed and consistency of operational decisions.Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight.Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction.Cross-Functional PartnershipPartner with engineering, DevOps, and business stakeholders.Influence teams that do not directly report into SRE and build shared accountability for production outcomes.Create clear service ownership models and operational expectations across teams.Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination.Communicate reliability posture, risks, trends, and investments to executive and technical audiences.Measures of SuccessImproved availability and reliability of critical services.Reduced time to detect, diagnose, mitigate, and recover from incidents.Increased percentage of services meeting observability and production-readiness standards.Reduced alert noise, operational toil, and manual incident-management activity.Increased adoption of service-level objectives and measurable reliability practices.Improved quality and completion rate of post-incident corrective actions.Increased automation across detection, triage, remediation, and operational workflows.Stronger ownership of production reliability across engineering teams.Leadership AttributesEngineering-first and automation-oriented.Comfortable challenging legacy operating models and assumptions.Able to move between technical detail and executive-level strategy.Outcome-focused, pragmatic, and data-driven.Builds trust through clarity, accountability, and strong partnership.Develops leaders and creates an inclusive, high-performance engineering culture.Treats incidents as opportunities to improve systems rather than assign blame.Brings urgency to operational risks while maintaining focus on sustainable solutions.Minimum QualificationsBachelor’s degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master’s degree or MBA preferred.10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teamsProven experience building or transforming a reliability or operational engineering organization.Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery.Experience establishing observability, incident-management, service-level objective, and production-readiness practices.Demonstrated ability to improve reliability through engineering and automation rather than process alone.Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems.Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring.Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholdersAbility to balance immediate operational needs with long-term engineering transformation.Strong written, verbal, and executive communication skills.Preferred QualificationsStrong experience operating large-scale systems in AWS or another major cloud environment.Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry.Demonstrated experience implementing OpenTelemetry or common instrumentation standards.Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities.Experience applying AI, machine learning, or agent-based automation to operational workflows.Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering.Software engineering experience and the ability to engage deeply in architecture and design discussions.Solid background in supporting high-profile launches, events, or systems with significant customer and business impact.*LI-YUnleash Your PotentialWhen you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future — but to redefine what’s possible — for yourself, for AI, and the world.AccommodationsIf you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form.Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.Posting StatementSalesforce is an equal opportunity employer and maintains a policy of non-discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that’s inclusive, and free from discrimination. Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications – without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law. This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between. Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education.In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records.At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions. The typical base salary range for this position is $197,300 - $313,700 annually. In select cities within the San Francisco and New York City metropolitan area, the base salary range for this role is $237,700 - $344,700 annually. The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable.SummaryLocation: New York - New York; Texas - Dallas; California - San FranciscoType: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Director, Site Reliability Engineering in Dallas, TX vacancy
  • Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through...  ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform...  ...vendor resources Willingness to work on-site at stated location in the job openingDepartment... 
    Suggested
    Contract work
    For contractors
    Work experience placement

    Cedent Consulting

    Dallas, TX
    4 days ago
  • $138.4k - $173k

     ...infrastructure as well as help improve the reliability, quality of services and overall...  ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability...  ...about our locations by visiting our site.Compensation & BenefitsThe base salary that... 
    Suggested
    Full time
    Flexible hours

    AppFolio

    Dallas, TX
    4 days ago
  •  ...Talent Acquisition Team will reach out to help you navigate our interview process.Lantern is seeking an experienced Senior Site Reliability Engineer to champion the reliability, availability, and performance of our Azure-based healthcare platform. In this pivotal role,... 
    Suggested

    Lantern

    Dallas, TX
    2 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Dallas, TX
    4 days ago
  • $174k - $253k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Suggested

    Google

    Sunnyvale, TX
    2 days ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Full time

    Vanguard

    Dallas, TX
    2 days ago
  • $100k - $115k

     ...Internal Developer Platform (IDP) as a product, treating engineering teams as customers and optimizing for reliability, usability, and delivery velocity.Define and...  ....4+ years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Systems Engineering... 
    Temporary work

    Analytic Partners

    Dallas, TX
    12 hours ago
  • $147k - $211k

     ...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style...  ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations... 

    Google

    Sunnyvale, TX
    2 days ago
  • $207k - $301k

    Develop strong, influential relationships with multiple stakeholders across the Site Reliability Engineering and Developer organizations.Serve as an expert on particular fields of knowledge related to rate limiting or sharding.Develop plans and lead projects on evolving... 

    Google

    Sunnyvale, TX
    2 days ago
  • Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run scalable, massively distributed, fault-tolerant systems. At Goldman Sachs, SRE is responsible for improving... 

    Goldman Sachs

    Dallas, TX
    1 day ago
  •  ...Administrator / SRE in Dallas to own production Java environments, middleware, and cloud automation. You will optimize performance, drive reliability, and mentor teammates while aligning with enterprise security and AI-enabled integrations. You will work across Java apps, IBM... 

    Motion Recruitment Partners LLC

    Dallas, TX
    4 days ago
  • $114k - $148k

     ...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based... 
    Full time
    Temporary work
    Work experience placement
    Remote work

    GrabJobs

    Irving, TX
    1 day ago
  •  ...The Depository Trust & Clearing Corporation (DTCC) is seeking a Senior Application Support Engineer (SRE) to enhance reliability, scalability, and performance of mission-critical applications. You will apply SRE principles across engineering, infrastructure, and operations... 

    The Depository Trust & Clearing Corporation

    Dallas, TX
    4 days ago
  •  ...Senior Site Reliability Engineer (Enterprise Platform) Location: Remote - US - Open to Europe if happy to overlap with EST Compensation: Competitive We are a high-growth software company supporting the development of a premier open-source, EVM-compatible public ledger... 
    Contract work
    Currently hiring
    Remote work

    GrabJobs

    Irving, TX
    3 days ago
  •  ...improving platform infrastructure and applications with high reliability, resiliency, performance & quality, and faster time-to-market...  ...documentation, including runbooks/playbooks; and, Using Chaos Engineering to test the robustness of the systems and applications.... 

    Software Technology Inc

    Dallas, TX
    3 days ago
  • $50 - $53 per hour

     ...area onsite at the project, significantly reducing and/or eliminating the demands to travel. Key Responsibilities: Site Reliability Engineers are expected to be able to drive technology triage efforts to completion by assisting with restoral steps, identifying... 
    Hourly pay
    Live in
    Work at office
    Local area
    Flexible hours
    3 days per week

    Accenture

    Dallas, TX
    1 day ago
  • Compliance EngineeringWe are Compliance Engineering, a global team of more than 500 engineers and scientists who work on the most complex...  ...systems by pushing for changes that improve capacity and reliability.Practicing sustainable incident management in a blameless postmortem... 

    Goldman Sachs

    Dallas, TX
    12 hours ago
  • $207k - $301k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    Sunnyvale, TX
    2 days ago
  • $262k - $365k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    Sunnyvale, TX
    1 day ago
  • We are Compliance Engineering, a global team of more than 300 engineers and scientists who work on the most complex, mission-critical...  ...evolving systems by pushing for changes that improve capacity and reliability.Practicing sustainable incident management in a blameless... 

    Goldman Sachs

    Dallas, TX
    4 days ago
  • $160k - $210k

     ...departmental collaboration and a unified sense of purpose, making teamwork a cornerstone of our success. We are looking for a Senior Site Reliability engineer to work on expanding our global footprint of datacenters and improve service management across Cognitiv. Our immediate... 
    Work at office
    Local area
    Immediate start
    Remote work

    GrabJobs

    Garland, TX
    1 day ago
  •  ...Site Reliability Engineer- W2 Role* Technical proficiency: Strong Proficiency in Java, Strong understanding of Database concepts (Oracle, SQL, Dynamo DB etc.) Industry standard SRE Tools like Prometheus, Grafana, Data Dog Etc Good to have skills: Cloud Concepts / AWS,... 

    RSA Tech Group

    Dallas, TX
    4 days ago
  •  ...This RoleAs a Senior Application Support Engineer, you will help power DTCC's global...  ...markets infrastructure by ensuring the reliability, availability, and performance of Institutional...  ...processing and settlement.Leveraging Site Reliability Engineering (SRE) principles... 
    Flexible hours
    Afternoon shift

    DTCC- The Depository Trust & Clearing Corporation

    Dallas, TX
    3 days ago
  •  ...ensure applications are highly available, reliable, and performant at a global scale....  ...Bachelor of Computer Science or related Engineering field required. Master's Degree preferred...  ...Minimum of 1 year of lead experience of site reliability engineering team required.... 
    Contract work
    Work at office

    3B Staffing LLC

    Irving, TX
    2 days ago
  • $155k - $222.6k

     ...maintain automation solutions that improve the reliability, scalability, and operational efficiency...  ...cloud environments. Partner with other engineering teams, product management, and business...  ...~2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Webex Events (formerly Socio)

    Dallas, TX
    5 days ago
  • $50 - $53 per hour

     ...area onsite at the project, significantly reducing and/or eliminating the demands to travel. Key Responsibilities: Site Reliability Engineers are expected to be able to drive technology triage efforts to completion by assisting with restoral steps, identifying root... 
    Hourly pay
    Live in
    Work at office
    Local area
    Flexible hours
    3 days per week

    Accenture

    Irving, TX
    2 days ago
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    Dallas, TX
    3 days ago
  • $155k - $410k

    Industry/SectorNot ApplicableSpecialismSAPManagement LevelDirectorJob Description & SummaryThe OpportunityAs a SAP BRIM Consultant, Director, you will play a pivotal role in helping clients optimize their operational efficiency through specialized consulting services for... 
    Full time
    Temporary work
    H1b

    PwC

    Dallas, TX
    12 hours ago
  • $189k - $250k

     ...eligible for immigration sponsorship. WHAT YOU’LL DO: As a Director on a lean, cross-functional team, you will work...  ...performance. You are brought in to: * Determine data availability and reliability and design a structured process to aggregate relevant data sets... 
    Full time
    Work at office
    Local area
    Remote work

    Accordion

    Dallas, TX
    1 day ago
  • $127.5k - $187k

     ...offer visa sponsorship for this position. Role Summary The Director, Portfolio Management serves as a principal portfolio advisor responsible...  ...all National Life Group companies. Social Media Policy [ Site Disclosure and Privacy Policy [ National Life Group 1... 
    Hourly pay
    Full time
    Temporary work
    Work experience placement
    Work at office
    Flexible hours
    Shift work

    National Life Insurance Company

    Addison, TX
    9 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, Site Reliability Engineering. Be the first to apply!