Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, Site Reliability Engineering

$197.3k - $313.7k

Salesforce

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.Job CategorySoftware EngineeringJob DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.Job Title: Director, Site Reliability EngineeringLocation: New York, NY; San Francisco, CA; Dallas, TXAbout the RoleWe are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities.In this role, you will transform our SRE function—moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch.As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide.Key ResponsibilitiesSRE Strategy and LeadershipDefine and execute the long-term strategy and roadmap for Site Reliability Engineering.Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures.Build and develop a high-performing team of site reliability and operations engineers.Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.Translate business priorities and customer impact into clear reliability investments and engineering outcomes.Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.Reliability EngineeringEstablish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services.Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.Define what it means for a service to be operationally and observably ready for production.Develop readiness reviews and certification practices for high-impact services and launches.Drive improvements in system availability, performance, resiliency, and recovery.Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment.ObservabilityDefine an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards.Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities.Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure.Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes.Establish governance and measurement to assess adoption and effectiveness of observability standards.Incident Management and Operational ExcellenceImprove incident detection, response, mitigation, communication, and learning.Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation.Reduce mean time to detect, acknowledge, mitigate, and recover.Improve on-call practices, escalation paths, runbooks, and operational ownership.Establish blameless post-incident review practices that produce measurable engineering improvements.Identify recurring sources of operational toil and create plans to eliminate or automate them.Partner with engineering leaders to ensure actions from incidents are prioritized and completed.Automation and AI-Enabled OperationsDevelop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation.Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support.Build automation that reduces cognitive load and improves the speed and consistency of operational decisions.Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight.Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction.Cross-Functional PartnershipPartner with engineering, DevOps, and business stakeholders.Influence teams that do not directly report into SRE and build shared accountability for production outcomes.Create clear service ownership models and operational expectations across teams.Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination.Communicate reliability posture, risks, trends, and investments to executive and technical audiences.Measures of SuccessImproved availability and reliability of critical services.Reduced time to detect, diagnose, mitigate, and recover from incidents.Increased percentage of services meeting observability and production-readiness standards.Reduced alert noise, operational toil, and manual incident-management activity.Increased adoption of service-level objectives and measurable reliability practices.Improved quality and completion rate of post-incident corrective actions.Increased automation across detection, triage, remediation, and operational workflows.Stronger ownership of production reliability across engineering teams.Leadership AttributesEngineering-first and automation-oriented.Comfortable challenging legacy operating models and assumptions.Able to move between technical detail and executive-level strategy.Outcome-focused, pragmatic, and data-driven.Builds trust through clarity, accountability, and strong partnership.Develops leaders and creates an inclusive, high-performance engineering culture.Treats incidents as opportunities to improve systems rather than assign blame.Brings urgency to operational risks while maintaining focus on sustainable solutions.Minimum QualificationsBachelor’s degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master’s degree or MBA preferred.10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teamsProven experience building or transforming a reliability or operational engineering organization.Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery.Experience establishing observability, incident-management, service-level objective, and production-readiness practices.Demonstrated ability to improve reliability through engineering and automation rather than process alone.Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems.Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring.Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholdersAbility to balance immediate operational needs with long-term engineering transformation.Strong written, verbal, and executive communication skills.Preferred QualificationsStrong experience operating large-scale systems in AWS or another major cloud environment.Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry.Demonstrated experience implementing OpenTelemetry or common instrumentation standards.Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities.Experience applying AI, machine learning, or agent-based automation to operational workflows.Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering.Software engineering experience and the ability to engage deeply in architecture and design discussions.Solid background in supporting high-profile launches, events, or systems with significant customer and business impact.*LI-YUnleash Your PotentialWhen you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future — but to redefine what’s possible — for yourself, for AI, and the world.AccommodationsIf you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form.Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.Posting StatementSalesforce is an equal opportunity employer and maintains a policy of non-discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that’s inclusive, and free from discrimination. Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications – without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law. This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between. Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education.In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records.At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions. The typical base salary range for this position is $197,300 - $313,700 annually. In select cities within the San Francisco and New York City metropolitan area, the base salary range for this role is $237,700 - $344,700 annually. The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable.SummaryLocation: New York - New York; Texas - Dallas; California - San FranciscoType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Director, Site Reliability Engineering in San Francisco, CA vacancy
  • $350k

     ...and novel use-cases. We’re hiring to grow the platform alongside the Tinker community. About the Role We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research... 
    Suggested
    Full time
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    2 days ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Suggested
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    4 days ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Suggested
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    5 hours ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Suggested
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    2 days ago
  • $165k - $241.4k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Suggested
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    4 days ago
  • $139.76k - $287.75k

     ...their business.We are seeking a Senior Site ReliabilityEngineer to help operate, scale...  ...will be instrumental in advancing the reliability, scalability, automation, observability,...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    3 days ago
  •  ...let’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ..., working together to build scalable, reliable, and secure products that empower businesses...  ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    5 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 

    Alembic

    San Francisco, CA
    5 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

     ...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical... 
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Temporary work

    TextNow

    San Francisco, CA
    3 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $194k - $267k

     ...defining work. We're all in on this mission. If you are too, let's talk.We are seeking a highly technical ObservabilitySite Reliability Engineer with a specialty in Google Cloud, to own and expand our Observability ecosystem into GCP. In this role, you will move beyond... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    5 hours ago
  • $194k - $267k

     ...in on this mission. If you are too, let's talk.The TeamThe Site Reliability team is dedicated to architecting and owning the foundational...  ...durable, automated systems that maximize platform reliability and engineering velocity.The ideal candidate is someone who enjoys analyzing... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $204k - $306k

     ...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from...  ...in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    4 days ago
  • $160k - $200k

    Job Purpose:BTIG seeks a DevOps/Site Reliability Engineer to join our technology team. This role is central to improving developer velocity by handling production application support escalations, managing and evolving our infrastructure stack, and providing operational... 
    Full time

    BTIG

    San Francisco, CA
    3 days ago
  • $220k - $235k

     ...SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role, you will pair deep...  ...Provide technical leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud PlatformDefine and... 
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    4 days ago
  • $195k - $257.5k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and operate the infrastructure that powers our blockchain platform at... 
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $186.9k - $267.7k

     ...requiring approximately 2 days per week on-site at Cisco offices in either San Francisco...  ...AI agents behave as intended, improving reliability and reducing risks. This unified...  ...and control.As a Staff Site Reliability Engineer (SRE), you will provide technical leadership... 
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Francisco, CA
    1 hour ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $215k - $275k

     ...Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing... 
    Work at office

    Anyscale

    San Francisco, CA
    2 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    5 days ago
  • $157k - $239k

    San Francisco, CA / Golden, COInfrastructure - Cloud Infrastructure /Full time /On-siteWanna join the adventure?As a Site Reliability Engineer with strong networking skills in our Cloud Infrastructure (SRE) team, you help the team own the networks that keep Loft running... 
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    3 days ago
  • $165k - $241.4k

     ...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale,... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    1 day ago
  •  ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'... 

    Alembic

    San Francisco, CA
    2 days ago
  •  ...Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to...  ...still expect us to hit our SLAs. What? We're looking for a Site level Reliability Engineer as part of Infrastructure team to: Own... 
    Contract work
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    IVO Inc

    San Francisco, CA
    3 days ago
  •  ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient... 

    TechChain Talent

    San Francisco, CA
    5 days ago
  •  ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it... 
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, Site Reliability Engineering. Be the first to apply!