Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer, Infrastructure Observability

Full-time

Private

Our client is seeking a Principal Site Reliability Engineer, Infrastructure Observability to help build and advance a world-class SRE function focused on observability, reliability, scalability, resilience, and automation across complex cloud and on-premises environments. This role will be instrumental in driving operational excellence through modern engineering practices, automation, and best-in-class observability tooling.

The ideal candidate brings deep expertise in cloud infrastructure, DevOps, SRE methodologies, incident management, automation, and infrastructure monitoring. They will serve as a strategic leader and hands-on technical expert, helping shape reliability practices across a highly distributed enterprise environment.

Responsibilities:

  • Lead the design and implementation of reliability-focused solutions that improve system availability and minimize service disruptions.
  • Drive SRE best practices across the organization, promoting a culture of automation, operational excellence, and continuous improvement.
  • Champion observability initiatives, including monitoring, alerting, logging, and application performance management (APM).
  • Conduct incident trend analysis and lead initiatives to reduce recurring technology failures.
  • Facilitate blameless post-mortems and reliability reviews to strengthen operational maturity.
  • Develop automated solutions to proactively prevent incidents and accelerate remediation efforts.
  • Create unified visibility across technology platforms to identify risks, redundancies, and optimization opportunities.
  • Partner with engineering, infrastructure, security, and business stakeholders to improve service reliability and operational performance.
  • Contribute to target-state architecture and the long-term evolution of the technology ecosystem.
  • Mentor engineers and help establish standards, frameworks, and operational practices across the organization.

Requirements:

  • Bachelor's degree or equivalent combination of education and experience.
  • 10+ years of experience designing, building, and operating enterprise infrastructure solutions with significant organizational impact.
  • 5+ years of hands-on experience with Amazon Web Services (AWS).
  • 5+ years building, leading, or supporting Site Reliability Engineering (SRE) and/or DevOps functions.
  • Experience implementing and operating chaos engineering practices at scale.
  • Proven success leading strategic technology and transformation initiatives.
  • Strong scripting, systems administration, and infrastructure automation experience.
  • Demonstrated ability to leverage automation to improve reliability and reduce operational risk.
  • Proficiency in multiple programming languages, including Python, Java, Go, Node.js, and/or .NET Core.
  • Strong database experience with SQL Server, PostgreSQL, MySQL, or similar platforms.
  • Deep understanding of Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability metrics, and reliability measurement frameworks.
  • Experience implementing and managing Error Budgets.
  • Strong incident response, root cause analysis, and service recovery expertise.
  • Experience standardizing observability, monitoring, logging, and dashboarding across enterprise environments.
  • Hands-on experience with tools such as New Relic, Splunk, Elastic Stack, Prometheus, Grafana, SolarWinds, and cloud-native monitoring platforms.
  • Experience with infrastructure automation and cloud management tools including Terraform, Ansible, Vault, and Vagrant.
  • Ability to influence stakeholders across technical and business teams.
  • Strong communication, leadership, and mentoring skills.
  • Willingness to participate in on-call rotations and support critical production environments.

Preferred Qualifications

  • Cloud, DevOps, or Site Reliability Engineering certifications.
  • Working knowledge of Microsoft Azure.
  • Experience within highly regulated, large-scale enterprise environments.

Why Apply?

This is an opportunity to play a pivotal role in shaping the reliability, observability, and operational excellence strategy for a leading enterprise technology organization. You'll work alongside senior engineering leaders, influence large-scale technology initiatives, and drive the adoption of modern SRE practices across a complex global environment.

Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer, Infrastructure Observability in Owings Mills, MD vacancy
  • $159k - $272k

     ...the opportunity to grow and make a difference in ways that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on... 
    Principal
    Full time
    Private practice
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    4 days ago
  • $159k - $272k

     ...Role SummaryCloud Reliability operates as a...  ...and reliability engineering, delivering secure...  ...the firm. The Principal Cloud Reliability...  ...on reliability, observability, automation, and...  ...design authority, site reliability...  ...Build and maintain infrastructure as code using Terraform... 
    Principal
    Full time
    Local area
    Remote work
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    10 hours ago
  •  ...understanding of how reliability, availability, recoverability...  ..., roadmaps, and engineering priorities.* Makes...  ...and operating cloud infrastructure with senior-level impact...  ..., resilience, observability, recovery, incident response...  ....* Cloud or site reliability engineering... 
    Principal

    Jobleads-US

    Owings Mills, MD
    10 hours ago
  • $159k - $272k

     ...SummaryCloud Storage Platform Engineering provides secure, scalable,...  ...The Cloud Storage Platform Principal is a senior individual...  ...architecture, engineering, and reliability outcomes for enterprise...  ...automation while partnering with infrastructure, cloud engineering and... 
    Principal
    Full time
    Local area
    Remote work
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    10 hours ago
  • $110k - $188k

     ...Lead Workplace Experience Engineer - Power Platform is responsible...  ...platform and AI service reliability, regulatory alignment, and...  ...consultants, and represents Infrastructure Operations in architecture,...  ...monitoring, telemetry, and observability strategy for Power Platform... 
    Suggested
    Full time
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    4 days ago
  •  ...T. Rowe Price is seeking a Principal Cloud Reliability Engineer to lead enterprise cloud foundations with...  ...guiding teams across applications and infrastructure to deliver scalable cloud...  ...and CloudFormation, and advancing observability, incident response, and platform... 

    Jobleads-US

    Owings Mills, MD
    5 days ago
  • $121k - $206k

     ...you. About the TeamThe AI Engineering and Application Development...  ...This team delivers secure, reliable, and forward-looking engineering...  ..., build, and implement infrastructure and software solutions for...  ...software, tools, and related observability) is sufficiently robust,... 
    Full time
    Work experience placement
    Local area
    Remote work
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    10 hours ago
  • $121k - $206k

     ...the opportunity to grow and make a difference in ways that matter to you. Role Summary We are seeking a hands-on Senior Software Engineer with deep expertise in Oracle Cloud technologies, including ERP, EPM, Oracle Integration Cloud (OIC), OTBI, and APEX. This role is... 
    Full time
    Local area
    Remote work
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    4 days ago
  • $170k - $220k

    Job DescriptionDewberry is expanding its Energy Market Sector and seeking a Physical Electrical Engineer to lead the growth of our medium and high voltage substation design practice. This is a unique opportunity to build and mentor a high-performing team while delivering... 
    Principal

    Dewberry

    Owings Mills, MD
    4 days ago
  • $122k - $209k

     ...automation, and cloud-native security controls. Working across infrastructure, security, and application teams, you will establish the...  ...enterprise standards, assess and mitigate risk, guide critical engineering decisions, and mentor technical talent across the... 
    Full time
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    1 day ago
  •  ...Computer Science, Mathematics, Engineering or a related field.Masters...  ...must be willing to work on-site in Woodlawn, MD 5 days a week...  ...for performance and reliability.Demonstrate a strong understanding...  ...using existing IBM DataPower infrastructure. Ensure interoperability with... 
    Principal
    Temporary work

    Global CI

    Windsor Mill, MD
    4 days ago
  •  ...Description Apply now: DevOps Engineer, location is Owings Mills,...  ...the cloud-native infrastructure that powers quantitative research...  .... Implement monitoring, observability, and alerting using Prometheus...  ...teams to improve platform reliability, scalability, and developer... 
    Contract work
    Work at office
    Immediate start
    2 days per week

    Mondo

    Owings Mills, MD
    a month ago
  • $110.6k - $178k

     ...The Job: As a Senior Manager, Customer Identity & Platform Engineering, you’ll be part of our IT - Customer Engagement team...  ...best practices for DevSecOps, CI/CD, automated testing, observability, reliability, and platform performance.Partner with Product Management... 
    Full time
    H1b
    Local area
    Remote work

    Stanley Black & Decker

    Towson, MD
    4 days ago
  •  ...T. Rowe Price is seeking a Cloud Storage Platform Principal to own the architecture, automation, and operational direction for enterprise cloud storage. You will work across NetApp, Nasuni, and AWS storage services to enable CloudNext and data center exit while mentoring... 

    Jobleads-US

    Owings Mills, MD
    10 hours ago
  • $145k - $247k

     ...leader to shape how AI-enabled software engineering evolves across our mobile organization...  ...processes Improve software quality, observability, telemetry, and release confidence through...  ...experiences, including performance, reliability, usability, telemetry, and release... 
    Full time
    Local area
    Remote work
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    2 days ago
  •  ...Omm IT Solutions in Woodlawn, MD is seeking a Principal Software Engineer to design and develop scalable Java-based applications. You will lead microservices using Spring Boot, RESTful APIs, and modern front-end frameworks while aligning with Twelve-Factor App principles... 
    Principal

    Jobleads-US

    Woodlawn, MD
    2 days ago
  •  ...T. Rowe Price Hong Kong is seeking a senior cloud/site reliability engineer to design, build, and operate enterprise-scale cloud platforms. The role emphasizes reliability, security, and performance across AWS with knowledge of Azure and related tooling. Experience... 

    Jobleads-US

    Owings Mills, MD
    10 hours ago
  •  ...Computer Science, Mathematics, Engineering or a related field.Masters...  ...must be willing to work on-site in Woodlawn, MD 5 days a week...  ...for performance and reliability.Demonstrate a strong understanding...  ...using existing IBM DataPower infrastructure. Ensure interoperability with... 
    Principal
    Temporary work

    Global CI

    Windsor Mill, MD
    1 day ago
  •  ...Senior AI Software Engineer at T. Rowe Price will design, build, and scale production-grade AI agents within a Salesforce-centric ecosystem. You will lead technical workstreams, incubate AI products, and partner with tech teams to enable broad adoption at scale. This... 

    Jobleads-US

    Owings Mills, MD
    3 days ago
  •  ...Our client, a leader in cloud infrastructure and enterprise solutions, is seeking a Cloud Engineer III to join their team. As a Cloud Engineer III, you will be part of the Cloud Operations Department supporting cross-functional teams. The ideal candidate will demonstrate... 
    Weekly pay
    Temporary work
    Flexible hours

    Experis Technology Group

    Owings Mills, MD
    2 days ago
  • $84.84k - $153.55k

     ...directed by management and senior staff. This position will provide Software solutions delivery support and mentoring for Software Engineers.  Lead the design and implementation of software solutions that meet business requirements and technical specifications.... 
    Full time
    Temporary work
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Covista

    Pikesville, MD
    4 days ago
  • $105.5k - $168.8k

     ...of us—from design and engineering to the manufacturing...  ...implementing monitoring, observability, systems management,...  ...optimal performance, reliability, and maintainability...  ...) Experience with Infrastructure as code (e.g....  ...learn more on our career site under "Our Commitment... 
    Hourly pay

    BD

    Cockeysville, MD
    2 days ago
  •  ...Principal Software Engineer Java Woodlawn, MD 5 days onsite Core Tech Stack: Java | OpenShift / AWS | Angular/React| JavaScript | Web Services | Spring Boot | Spring Batch | Typescipt | Microservices | REST/SOAP | OpenShift | Docker | PostgreSQL | DB2 |... 
    Principal

    Akaasa Technologies

    Woodlawn, MD
    2 days ago
  • $121k - $206k

     ...ways that matter to you. Role SummaryAs a Senior AI Software Engineer, you will design, build, and scale production-grade AI agents...  ...technical debt and drive ongoing improvements in AI platforms and infrastructure.Proactively seek opportunities to apply Agentic AI... 
    Full time
    Local area
    Remote work
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    4 days ago
  • $130k - $140k

     ...Full-time Description Software Systems Engineer Position Summary Trust Consulting Services, Inc. is seeking a Software...  ...engineering changes, and corrective actions that strengthen software reliability, maintainability, performance, and product quality Perform... 
    Full time
    Temporary work
    Local area
    Remote work
    Monday to Friday

    GrabJobs

    Randallstown, MD
    1 day ago
  • $121k - $206k

     ...investment operations and portfolio of investment applications through key long-term technology initiatives. As a Senior Software Engineer, you will play a critical role in designing and delivering next-generation, cloud-native applications that solve complex business,... 
    Full time
    Local area
    Remote work
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    4 days ago
  • $77.5k - $132k

     ...difference in ways that matter to you. Role Summary Join T. Rowe Price's Global Product Technology team as an Associate Software Engineer to help build the next generation of product data platforms and client experiences. This entry-level role is ideal for early-career... 
    Local area
    Remote work
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    2 hours ago
  • $159k - $272k

     ...that matter to you. Role SummaryThe Principal Desktop Engineer is a hands-on engineering leader responsible...  ...macOS platforms, virtual desktop infrastructure (VDI), and modern endpoint...  ...OS lifecycle compliance, deployment reliability, automation maturity (including zero... 
    Principal
    Full time
    Contract work
    Local area
    Remote work
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    2 days ago
  • Job-ID27127872Reference26-01207Information Technology - Engineer, Software Sr PURPOSE: Performs complex analysis, design, development...  ...applications. Works with cross functional teams to develop highly reliable software that runs at scale. Provides recommendations to infuse... 
    Work experience placement

    Mindlance

    Owings Mills, MD
    2 days ago
  • $143k - $156k

     ...Shift5 Shift5 is the observability platform for onboard...  ..., and other critical infrastructure. Come join us....  ...is seeking a Systems Engineer to join our team. In...  ...to support customer site integrations and testing...  ...aerospace, defense, or high-reliability electronics sectors.... 
    Contract work
    For contractors
    Remote work
    Flexible hours

    GrabJobs

    Stevenson, MD
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer, Infrastructure Observability. Be the first to apply!