Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer (Cloud, Observability & Automation)

$95k - $180k
Full-time

DTCC

Are you ready to make an impact at DTCC?

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.

Pay and Benefits:

  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits, based on location
  • Pension / Retirement benefits
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).

The Impact You Will Have in This Role

The Enterprise Application Support (EAS) team supports critical applications across the ITP and ECS business lines, ensuring the reliability, scalability, and performance of enterprise platforms.

As a Principal Site Reliability Engineer (SRE) , you will drive operational excellence across mission-critical systems. You will lead reliability initiatives, champion observability and automation, drive major incident response, and partner with engineering, infrastructure, and security teams to build resilient, highly available applications.

This is a hands-on technical leadership role focused on improving system performance, reducing operational risk, accelerating recovery, and advancing SRE best practices through modern cloud, observability, automation, and AI-powered technologies.

Your Primary Responsibilities :

  • Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications.
  • Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms.
  • Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs.
  • Lead major incident response, root cause analysis, and continuous service improvement initiatives.
  • Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies.
  • Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle.
  • Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives.
  • Identify operational risks and deliver strategic reliability improvements across the technology ecosystem.
  • Collaborate with technical and business stakeholders to improve service reliability and operational outcomes.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 8+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines.

Talent Needed for Success

  • Strong hands-on experience with AWS and cloud-native architectures.
  • Proficiency in Python, Java, Go, or similar programming languages.
  • Strong Linux/Unix systems administration and troubleshooting experience.
  • Expertise in observability and monitoring platforms including Splunk, Grafana, Dynatrace, and ITSI.
  • Experience leading major incident management and root cause investigations in complex production environments.
  • Strong understanding of distributed systems, resiliency engineering, performance tuning, automation, and operational excellence.
  • Excellent communication and stakeholder management skills with the ability to influence technical and business partners.

Preferred Qualifications

  • Experience with AI-assisted engineering tools such as Amazon Q, Kiro, or similar technologies.
  • Experience designing and measuring SLOs, SLIs, and operational KPIs.
  • Experience supporting large-scale enterprise applications in financial services or other highly regulated environments.

The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.


With over 50 years of experience, DTCC is the premier post-trade market infrastructure for the global financial services industry. From 20 locations around the world, DTCC, through its subsidiaries, automates, centralizes, and standardizes the processing of financial transactions, mitigating risk, increasing transparency, enhancing performance and driving efficiency for thousands of broker/dealers, custodian banks and asset managers. Industry owned and governed, the firm innovates purposefully, simplifying the complexities of clearing, settlement, asset servicing, transaction processing, trade reporting and data services across asset classes, bringing enhanced resilience and soundness to existing financial markets while advancing the digital asset ecosystem. In 2024, DTCC’s subsidiaries processed securities transactions valued at U.S. $3.7 quadrillion and its depository subsidiary provided custody and asset servicing for securities issues from over 150 countries and territories valued at U.S. $99 trillion. DTCC’s Global Trade Repository service, through locally registered, licensed, or approved trade repositories, processes more than 25 billion messages annually. To learn more, please visit us at or connect with us on LinkedIn, X, YouTube, Facebook and Instagram.

DTCC proudly supports Flexible Work Arrangements favoring openness and gives people freedom to do their jobs well, by encouraging diverse opinions and emphasizing teamwork. When you join our team, you’ll have an opportunity to make meaningful contributions at a company that is recognized as a thought leader in both the financial services and technology industries. A DTCC career is more than a good way to earn a living. It’s the chance to make a difference at a company that’s truly one of a kind.

Learn more about Clearance and Settlement by clicking here.


Serves as a dedicated technology resource for advancing DTCC’s business opportunities and providing industry thought leadership for leveraging new technology. The goal of this new department is to partner internally with IT, our business and regulatory divisions and externally with clients, regulators, and fintech vendors, to help build new platforms and business models to advance DTCC’s mission to support the financial markets.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer (Cloud, Observability & Automation) in Jersey City, NJ vacancy
  •  ...Principal Software Engineer This role drives the evolution of the SCM platform...  ...-added core services and automation that streamline and...  ...practices and standards for reliability, quality, and operational...  ...Experience implementing observability, SLO/SLA frameworks, and... 
    Principal

    Chase

    Jersey City, NJ
    1 day ago
  •  ...DESCRIPTION Elevate your engineering prowess to...  ...among the top echelon in site reliability. As a Sr Lead Site Reliability...  ...work through automation. You'll join a team of...  ...create and implement observability and reliability designs...  ...Experience with cloud-based data and analytics... 
    Suggested

    J.P. Morgan

    Jersey City, NJ
    1 day ago
  •  ...itD is seeking a Site Reliability Engineer to develop and enhance automation solutions that improve the reliability, scalability...  ...efficiency of large-scale cloud infrastructure. The ideal...  ...environments. Knowledge of monitoring, observability, and site reliability... 
    Suggested
    Work experience placement
    Remote work

    GrabJobs

    Jersey City, NJ
    1 day ago
  •  ...DESCRIPTION Elevate your engineering prowess to...  ...among the top echelon in site reliability. As an Associate Site...  ...reducing work through automation. You'll join a team of...  ...machine learning, and cloud development. Our $9.5...  ...services meet reliability, observability, and operability... 
    Suggested
    Worldwide

    J.P. Morgan

    Jersey City, NJ
    1 day ago
  •  ...you a seasoned engineer ready to solve complex...  ....As a Senior Principal Software...  ...of large-scale, cloud-native systems that...  ...meet performance, reliability, and security requirementsDrive...  ...as code, CI/CD automation, and observability frameworks...  ...coverage, on-site health and... 
    Principal
    Work at office

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  •  ...the right place.As a Principal Software Engineer at JPMorgan Chase...  ...— ensuring safety, observability, and reproducibility...  ...assisted development and automation capabilities, to...  ...disciplines (e.g., cloud, AI/ML, data engineering...  ...care coverage, on-site health and wellness... 
    Principal

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  • $210k - $220k

     ...workflow platform applies AI, automation, and integration with...  ...with security, IT, engineering, finance, and other...  ...on our journey. Senior Site Reliability Engineer - Government Cloud You'll join the team responsible...  ...as repeatable, observable, and assessment-ready as... 
    Work at office
    Remote work

    GrabJobs

    Newark, NJ
    4 days ago
  •  ...Sr. Lead Software Engineer at JPMorgan Chase within...  ...the foundational cloud infrastructure that...  ..., driving platform reliability, scalability, and automation while collaborating...  ..., including CI/CD, observability, and automated...  ...care coverage, on-site health and wellness... 

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  •  ...Innovation team, the Lead DevOps Engineer will play a key role in driving the...  ...best practices, including CI/CD automation, infrastructure reliability, observability, and secure deployment patterns....  ...similar roleStrong knowledge of cloud platforms such as Azure, GCP or AWSBachelor... 
    Work experience placement
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Jersey City, NJ
    3 days ago
  •  ...Senior Devops Engineer Location: Durham, NC & Jersey City, NJ...  ...capabilities (Version Control, CI/CD, Automation, Experimentation, Lean, Observability, etc.) Develop CI/CD...  ...years’ experience in Software, Cloud, or Site Reliability Engineering fields BS or MS... 
    Long term contract

    Samprasoft

    Jersey City, NJ
    4 days ago
  •  ...mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase...  ...solutions. Through code and cloud infrastructure, you will...  ...reliability approaches using automated continuous integration and...  ...recoveryExperience in observability such as white and black box... 
    Work at office

    JP Morgan Chase

    Jersey City, NJ
    1 day ago
  •  ...solutions.As a Senior Lead Site Reliability Engineer at JPMorganChase within...  ...principles, implementing robust observability, and shaping the target...  ...(e.g., testing/validation automation and production readiness),...  ...deep expertise in public cloud and modernization journey.... 

    JP Morgan Chase

    Jersey City, NJ
    1 day ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available...  ...reliability, scalability, performance, automation, and operational excellence....  ...Knowledge of monitoring, observability, alerting, and performance... 
    Local area

    2T Consulting

    Jersey City, NJ
    5 days ago
  •  ...Senior Lead Software Engineer at JPMorgan Chase within...  ...coding, peer review, automated testing) and promoting...  ...Ability to build with cloud‑native microservices at...  ....Exposure with observability tooling and practices...  ...health care coverage, on-site health and wellness centers... 
    Contract work
    For contractors

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  • $120k - $175k

     ...highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate...  ...propose solutions. Create and support observability/monitoring tools and vendor...  ...application development in these areas: Cloud computing such as AWS, Azure, and/or... 
    Full time
    Remote work
    Work visa
    Flexible hours

    GrabJobs

    Jersey City, NJ
    3 days ago
  •  ...realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the...  ...quality checks, test/validation automation, and operational readiness), ensuring...  ...knowledge and experience in observability such as white and black box... 
    Local area

    JP Morgan Chase

    Jersey City, NJ
    2 days ago
  • $113.3k - $205.52k

     ...What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development...  ...learn from production to build the automation and agentic tooling that improves reliability...  ...). (Required) Experience utilizing observability tools (i.e. Grafana, Prometheus,... 
    Work at office
    Remote work
    Worldwide
    Flexible hours

    GrabJobs

    Jersey City, NJ
    1 day ago
  •  ...newly formed platform engineering product line at a leading...  ..., resiliency, observability, and compliance plumbing...  ...enterprise.As a Senior Principal Software Engineer on a...  ...disciplinesExtensive practical cloud-native...  ...health care coverage, on-site health and wellness centers... 
    Principal

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  •  ...are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available...  ...reliability, scalability, performance, automation, and operational excellence....  ..., alerting, logging, and observability capabilities to detect and... 

    Longfinch Technologies

    Kearny, NJ
    20 hours ago
  • $142.32k - $213.48k

     ...excellence through secure, reliable, and efficient...  ...Senior Full-stack Java Engineer - Futures & Derivatives...  ...-on Java development, cloud-native architecture, and...  ...development cycles, automate processes, and improve...  ...maintain CI/CD, DevOps, and observability practices to enhance... 
    Full time
    Work at office

    Citigroup

    Jersey City, NJ
    3 days ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at...  ...solutions. Through code and cloud infrastructure, you will configure...  ...approaches using automated continuous integration and...  ...Terraform. Experience in observability including white and black... 

    Chase

    Jersey City, NJ
    3 days ago
  •  ...Director of Software Engineering at JPMorganChase...  ...using cloud-native architecture...  ...engineering and SDLC/TLM automation (using enterprise...  ..., scalability, reliability, and cost-to-serve...  ..., performance, observability, disaster recovery...  ...care coverage, on-site health and... 

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  •  ...quality solutions. As a Lead Site Reliability Engineer at JPMorganChase within...  ...quality checks, test/validation automation, and operational readiness...  ...services in a public cloud environment (AWS, Azure,...  ...familiarity with cloud-native observability, auto-scaling, and... 
    Work at office

    J.P. Morgan

    Jersey City, NJ
    1 day ago
  • $113.1k - $232.3k

     ...Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied...  ...and advanced proficiency across cloud platform engineering, observability, and performance and reliability...  ...release on error budgets and automated reliability checks, and owning the... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    Jersey City, NJ
    3 days ago
  •  ...with leaders across engineering and technology to...  ...define objective reliability goals for...  ...include composing observability designs through instrumentation...  ...include automating services to...  ...The Senior GCP Site Reliability Engineer...  ...of experience in cloud infrastructure engineering... 
    Work at office
    Shift work
    Day shift

    Bank of America Corporation

    Jersey City, NJ
    a month ago
  •  ...Lead Infrastructure Engineer at JPMorgan Chase within...  ...within delivery and automation routines to identify...  ...guidance.Partner with SRE/Observability and Cyber teams to...  ...for automation and reliability outcomes.Practical experience...  ...care coverage, on-site health and wellness... 

    JP Morgan Chase

    Jersey City, NJ
    2 days ago
  •  ...to take your software engineering career to the next...  ...architecture decisions, build cloud-native services on...  ...evaluation and observability. You will help raise...  ...better outcomes through reliable automation and decision support....  ...health care coverage, on-site health and wellness... 

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  •  ...companies.As Senior Principal Software Engineer at JPMorganChase...  ..., release, and observability phases....  ...engineering and SDLC/TLM automation (using...  ...speed, scalability, reliability, and cost-to-serve...  ...with on-prem , cloud-native ecosystems...  ...care coverage, on-site health and... 
    Principal

    JP Morgan Chase

    Jersey City, NJ
    2 days ago
  • $200k - $220k

     ...seeking a hands-on Lead Software Engineer / Principal Software Engineer with...  ...and building large-scale cloud solutions. This individual...  ...recommend solutions that improve reliability, scalability, performance,...  ...monitoring, logging, and observability tools such as Datadog,... 
    Principal
    Full time
    Local area

    Broadridge

    Newark, NJ
    2 days ago
  •  ...continuous delivery pipelines to automate software testing and...  ...Code IaC Provision and manage cloud resources AWS Azure GCP using...  ...degree in Computer Science Engineering or a related technical field...  ...5 years working in a DevOps site reliability SRE or cloud engineering rol

    LTM

    Jersey City, NJ
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer (Cloud, Observability & Automation). Be the first to apply!