Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas

Goldman Sachs

Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run scalable, massively distributed, fault-tolerant systems. At Goldman Sachs, SRE is responsible for improving the availability and reliability of the firm’s most critical platform services and ensures they meet the requirements of our internal and external users. It is also responsible for firmwide policies and standards focused on firm’s digital resilience. We are looking for engineers who are motivated to collaborate with our businesses to build and run sustainable production systems, which can evolve and adapt to changes in our fast-paced, global business environment.The SRE team develops and maintains platforms and tools which help other Engineering teams in Goldman Sachs to build and operate reliable and resilient systems. These systems span on-premises datacenters and multiple public cloud environments. The platforms we offer include central logging, monitoring, agents and alerting and we provide tools to drive adoption and improvements to capacity planning, operational readiness assessments, production incident postmortems, SLIs / SLOs, and deployment automation including canary releases.The products and services we provide to our internal customers are used by thousands of engineers every day. We believe that reliability is the most important feature of any system, and we are devoted to giving our engineers the platforms and tools they need to build and operate reliable products.Role OverviewAs a Site Reliability Engineer (SRE) at Goldman Sachs, you will be a pivotal leader in ensuring the availability, reliability, and scalability of the firm's most critical platform applications and services. You will combine deep software and systems engineering expertise to architect, build, and run large-scale, massively distributed, fault-tolerant systems. This role involves providing technical leadership, mentoring senior engineers, and collaborating closely with internal teams and executive stakeholders to build and operate sustainable production systems that can adapt to our dynamic global business environment. You will drive a culture of continuous improvement, championing the adoption of advanced SRE principles and best practices across the organization. ResponsibilitiesStrategic Reliability & Performance: Drive the strategic direction for availability, scalability, and performance of mission-critical applications and platform services, ensuring alignment with firm-wide objectives.Architectural Leadership: Lead the design, build, and implementation of highly available, resilient, and scalable infrastructure and application architectures.Advanced Automation & Tooling: Architect and develop sophisticated platforms, tools, and automation solutions to eliminate toil, optimize operational workflows, and enhance deployment processes across the enterprise.Complex Incident Management & Post-Mortem Analysis: Lead critical incident response, conduct in-depth root cause analysis for systemic issues, and implement long-term preventative measures to significantly enhance system stability and resilience.System Design & Capacity Planning: Partner with development teams to embed reliability into application design from inception, provide expert system design consulting, and lead comprehensive capacity planning initiatives for future growth.Observability & Insights: Define and implement advanced monitoring, high volume logging with multi-user query capabilities, and tracing strategies to provide deep, actionable insights into application performance, infrastructure health, and user experience.Technical Vision & Mentorship: Provide technical vision, lead complex technical projects, conduct rigorous code reviews, enforce SDLC best practices, and actively mentor and develop senior and staff-level engineers.Technology Evaluation & Adoption: Stay at the forefront of industry trends and advancements, evaluating and integrating cutting-edge tools and frameworks to significantly improve operational efficiency and reliability.On-Call Leadership: Participate in and lead on-call rotations, providing expert guidance and hands-on support for critical system incidents.QualificationsExperience: Minimum of 6+ years of hands-on experience in Site Reliability Engineering, with a proven track record in architecting, designing, building, and maintaining highly available, scalable, and fault-tolerant systems at an enterprise level.Technical Proficiency:Exceptional programming skills in one or more major languages such as Java, Python, Go with a focus on building robust, scalable software.Extensive hands-on experience with cloud platforms (e.g., AWS, GCP) and deep expertise in containerization and orchestration technologies (e.g., Docker, Kubernetes).Mastery of Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation) and configuration management tools (e.g., Puppet, Chef, Ansible).Advanced proficiency in Prompt Engineering and Retrieval-Augmented Generation (RAG) architectures to automate complex SRE workflows, such as the generation of Infrastructure as Code (IaC), dynamic runbooks, and incident response summaries.Profound understanding of Linux internals, networking, distributed systems, and advanced system performance tuning.Expertise in designing and implementing comprehensive monitoring, alerting, logging and tracing solutions (e.g., Prometheus, Grafana, ELK stack, Datadog, PagerDuty).Deep experience with CI/CD tools and practices (e.g., Jenkins, GitLab, Maven).Strong foundation in databases and distributed systems.Exceptional problem-solving abilities and analytical skills, with a track record of resolving complex technical challenges.Preferred Experience:Experience with Distributed Databases like Elastic SearchExperience with working on GCP Big QueryExperience with messaging Systems Like KafkaEducation: Advanced degree (Bachelor’s or Mas ter's or PhD) in Computer Science or a related technical field involving coding and/or systems engineering, or equivalent practical experience.Soft Skills: Superior communication, collaboration, and interpersonal skills, with the ability to influence technical direction, lead cross-functional initiatives, and effectively engage with global teams and executive leadership. Proven ability to work independently, manage multiple complex stakeholders, and drive significant organizational change.Posting Date: 2026-04-13

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas in Dallas, TX vacancy
  • Site Reliability Engineering (SRE) is an engineering discipline that combines software development and systems engineering to build and run large-...  ...availability and reliability of our firm's most critical platform services and ensures they meet the requirements of our... 
    Suggested

    Goldman Sachs

    Dallas, TX
    23 days ago
  • WHAT WE DO Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering...  ...reliable, observable, and resilient platforms that support critical business...  ...teams more effective, and championing SRE principles (such as SLOs, error budgets... 
    Suggested

    Goldman Sachs

    Dallas, TX
    4 days ago
  •  ...fulfill travel worldwide.SRE Software Systems Engineer IV - Data...  ...artificial intelligence platforms that power smarter travel...  ...will drive platform reliability, auto-scaling and...  ...role requires strong Site Reliability...  ...DA1SummaryLocation: Dallas-Fort Worth MetroplexType... 
    Suggested
    Full time
    Worldwide
    Flexible hours
    Weekend work

    Sabre Holdings

    Dallas, TX
    7 days ago
  • $85 - $90 per hour

     ...Role:  Senior SRE Engineer  Location: Dallas / Fort Worth, Texas Rate: up to $85-$90 per hour INC...  ...Structure: 8 Month contract *** 4 days on-site *** -- We have a great new...  ...Experience with container orchestration platforms such as Kubernetes. ~ Experience using... 
    Suggested
    Hourly pay
    Contract work
    Work experience placement

    CorGTA

    Dallas, TX
    more than 2 months ago
  •  ...Role: Vice President – AI Engineer Division: Risk Engineering...  ...Risk Location: Dallas, Americas About Goldman...  ...risk metrics, our platform is continuously growing...  ..., scalability, and reliability in distributed and...  ...state‑of‑the‑art on‑site health centers in... 
    Suggested
    Full time
    Temporary work
    Work at office

    Jobleads-US

    Dallas, TX
    4 days ago
  •  ...outline CORPORATE TITLE Vice President language OFFICE LOCATION(S) Dallas account_balance...  ..., Software Engineering withGoldman Sachs...  ...availability and reliability of database systems...  ...enhance our data platform.Design, implement...  ...state‑of‑the‑art on‑site health centers in... 
    Full time
    Temporary work
    Work at office

    Jobleads-US

    Dallas, TX
    1 day ago
  •  ...Observability/SRE /Performance EngineerLocation: Minneapolis, MN; Dallas, TX; Brookfield, WI; Columbia, SC - Hybrid Duration: 4+-month Possibility of extension...  ..., dashboards). Overall looking for a good Reliability Engineer that will support our environments by setting... 

    InterSources

    Dallas, TX
    4 days ago
  • Job Duties: Vice President, Systems Engineering with Goldman Sachs Services LLC in Dallas, Texas. Apply expertise across all Network Engineering...  ...firm’s infrastructure and platforms.Oversee the identification,...  ...and ensure network service reliability, performance and... 
    Work experience placement
    Work at office
    Remote work

    Goldman Sachs

    Dallas, TX
    7 days ago
  •  ...identify compliance, conduct, and reputational risks, and refine firm controls as appropriate. CTG’s global team (with locations in Dallas, New York, Salt Lake City, London, Warsaw, Tokyo, Singapore, and Hong Kong) is comprised of individuals with varying backgrounds... 
    Work at office

    Goldman Sachs

    Dallas, TX
    a month ago
  •  ...Bengaluru, Amsterdam, London, New York, and Dallas. All our offices work closely together...  ....YOUR IMPACT We are seeking a Vice President to join our Operations Enablement team...  ...senior Operations partner to Product, Engineering, Risk, Legal, Compliance, and functional... 
    Work at office
    Worldwide

    Goldman Sachs

    Dallas, TX
    5 days ago
  •  ...the ability to think outside the box? We are seeking a Vice President based in Dallas to report to the COO of Wealth Solutions and partner with...  ...cross-divisional businesses such as HCM, CWS, Operations, Engineering, Compliance and Legal to develop and execute project... 
    Work experience placement

    Goldman Sachs

    Dallas, TX
    a month ago
  • Corporate Insurance & Advisory Vice President - Corporate Insurance Risk Manager (Dallas) Corporate Insurance & Advisory manages the Firm’s global commercial insurance needs and advises its investing businesses on insurance-related risk. The team identifies operational... 
    Contract work
    Local area

    Goldman Sachs

    Dallas, TX
    a month ago
  •  ...Audit, Data Analytics, Technology Audit, Vice President, Dallas The Goldman Sachs Group, Inc. is a...  ...effective controls by assessing the reliability of financial reports, monitoring the...  ...cyber-security and technology risk, and engineering.RESPONSIBILITIESDevelop and maintain... 

    Goldman Sachs

    Dallas, TX
    a month ago
  •  ...every day while also partnering with Product, Digital, Sales, and Engineering to build out the next generation capabilities in Transaction...  ...in Tokyo, Singapore, Bengaluru, London, New York, and Dallas. All our offices work closely together as a single global team... 
    Work at office
    Local area

    Goldman Sachs

    Dallas, TX
    a month ago
  •  ...Vice President, Manufacturing Operations - Sarnova - Dallas, TX Job Description Posted Friday, September 4, 2026 at...  ...manufacturing relationships, ensuring reliable sourcing of medical product...  ..., Product Management, Finance, Engineering, and Commercial teams to ensure... 
    Contract work

    Jobleads-US

    Dallas, TX
    5 days ago
  •  ...difference, come join our team and help shape the future of convenience.The SRE RunOps Engineer 2 is responsible for ensuring the reliability, availability, and performance of the 7NOW delivery platform and associated services through proactive monitoring, incident response... 
    Hourly pay
    Work experience placement

    7 eleven

    Irving, TX
    7 days ago
  • $290k - $320k

    Vice President, Quality and Commissioning (Dallas, Tulsa or Kansas City) Remote (Strong preference...  ...technology, Beale's platform is designed to address...  ...communities to secure reliable power solutions and enable...  ...during commissioning are engineered out of future project... 
    Local area
    Remote work
    Shift work

    Beale Infrastructure Group

    Dallas, TX
    5 days ago
  •  ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes...  ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Dallas, TX
    27 days ago
  • $172k - $300k

     ...Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable,...  ...on heroics.If you are an expert in SRE practices who loves building the...  ...premises environments and foundational platform services.Experience with data or ML platforms... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    1 day ago
  •  ...Management, Strategic Sourcing (Spend Management)Location: Dallas, TXLevel: Associate or Vice President (title will be assigned appropriate to years of...  ...Management, Operational Risk and Resilience, and CPM Engineering teams to deliver business planning and analytics, expense... 
    Contract work
    Temporary work
    For contractors
    Work at office

    Goldman Sachs

    Dallas, TX
    a month ago
  •  ...governing transactions in the financial markets. The Fund Accounting Team within the Controllers Group is looking to add a Vice President in Dallas, TX.We're a team of specialists charged with managing the firm’s liquidity, capital and risk, and providing the overall... 

    Goldman Sachs

    Dallas, TX
    24 days ago
  •  ...Private Wealth Advisors (PWAs) deliver an unparalleled investment platform inclusive of the full product and service offerings of Goldman...  ...locations in the United States: Atlanta, Boston, Chicago, Dallas, Denver, Detroit, Houston, Los Angeles, New York, Miami (including... 

    Goldman Sachs

    Dallas, TX
    7 days ago
  • $111.61k - $131.3k

     ...support, product/project management, or application developmentPreferred Skills/ExperienceStrongexpertiseinSiteReliabilityEngineering(SRE),DevOps,ProductionSupport,PlatformEngineering,andDistributedSystemsOperations.Experienceleadingtechnicalteams,incidentresponseefforts... 
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Irving, TX
    1 day ago
  • YOUR IMPACTWe are looking for an experienced Technical Program Manager to join our Data Engineering team. Data is the lifeblood of Goldman Sachs, and our data platform is critical in delivering commercial success. In this role, you will partner directly with the leadership... 
    Flexible hours
    Shift work

    Goldman Sachs

    Dallas, TX
    7 days ago
  • The Core Engineering The Core Engineering builds and operates the platforms, applications, data solutions, models, and analytics that power critical processes for...  ...to meet production-grade standards of quality, reliability, and maintainability.Hands-on proficiency with... 

    Goldman Sachs

    Dallas, TX
    14 days ago
  • Build/Release Engineer Client offering an impressive work environment, and cutting edge technology stack, is seeking a Build/Release Engineer...  ...opportunity, starting with a 6 month assignment in the Dallas area. Duties will include: Installation and Configuration of MSBuild... 
    Long term contract

    Georgia IT Inc

    Dallas, TX
    6 days ago
  • About the RoleThe Enterprise Platforms team within our Data & Agentic Platform Solutions...  ...and DevOps capabilities. As an IT Site Reliability Engineer within the Enterprise Platforms team,...  ...strategy — while also providing platform SRE support across a broader portfolio of... 
    Local area

    Texas Instruments

    Dallas, TX
    24 days ago
  • Cloud BC Labs in Dallas is seeking an experienced OpenShift Engineer to design, deploy, administer, and support enterprise container platforms using OpenShift and Kubernetes. The role emphasizes production reliability, scaling, and cross-functional collaboration. You will... 

    CloudBC Labs

    Dallas, TX
    2 days ago
  • What We DoAt Goldman Sachs, our Engineers don’t just make things - we make things possible. Change the world by connecting people and...  ...part of WM Engineering at Goldman Sachs, the WM Cloud Enablement Platform team is responsible for enabling the use of public cloud... 

    Goldman Sachs

    Dallas, TX
    a month ago
  •  ...Our client is looking Dallas, TX / Austin, TX (Hybrid) project Dallas, TX / Austin...  ...requirements. Position: Senior Agentic AI Engineers (JAVA, Resiliency & Observability...  ...too Experience on AI/Agentic AI, SRE, Reliability Engineering, Observability, and... 

    Lorven technologies

    Dallas, TX
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas. Be the first to apply!