Principal Site Reliability Engineer, Infrastructure Observability
Private
Our client is seeking a Principal Site Reliability Engineer, Infrastructure Observability to help build and advance a world-class SRE function focused on observability, reliability, scalability, resilience, and automation across complex cloud and on-premises environments. This role will be instrumental in driving operational excellence through modern engineering practices, automation, and best-in-class observability tooling.
The ideal candidate brings deep expertise in cloud infrastructure, DevOps, SRE methodologies, incident management, automation, and infrastructure monitoring. They will serve as a strategic leader and hands-on technical expert, helping shape reliability practices across a highly distributed enterprise environment.
Responsibilities:
- Lead the design and implementation of reliability-focused solutions that improve system availability and minimize service disruptions.
- Drive SRE best practices across the organization, promoting a culture of automation, operational excellence, and continuous improvement.
- Champion observability initiatives, including monitoring, alerting, logging, and application performance management (APM).
- Conduct incident trend analysis and lead initiatives to reduce recurring technology failures.
- Facilitate blameless post-mortems and reliability reviews to strengthen operational maturity.
- Develop automated solutions to proactively prevent incidents and accelerate remediation efforts.
- Create unified visibility across technology platforms to identify risks, redundancies, and optimization opportunities.
- Partner with engineering, infrastructure, security, and business stakeholders to improve service reliability and operational performance.
- Contribute to target-state architecture and the long-term evolution of the technology ecosystem.
- Mentor engineers and help establish standards, frameworks, and operational practices across the organization.
Requirements:
- Bachelor's degree or equivalent combination of education and experience.
- 10+ years of experience designing, building, and operating enterprise infrastructure solutions with significant organizational impact.
- 5+ years of hands-on experience with Amazon Web Services (AWS).
- 5+ years building, leading, or supporting Site Reliability Engineering (SRE) and/or DevOps functions.
- Experience implementing and operating chaos engineering practices at scale.
- Proven success leading strategic technology and transformation initiatives.
- Strong scripting, systems administration, and infrastructure automation experience.
- Demonstrated ability to leverage automation to improve reliability and reduce operational risk.
- Proficiency in multiple programming languages, including Python, Java, Go, Node.js, and/or .NET Core.
- Strong database experience with SQL Server, PostgreSQL, MySQL, or similar platforms.
- Deep understanding of Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability metrics, and reliability measurement frameworks.
- Experience implementing and managing Error Budgets.
- Strong incident response, root cause analysis, and service recovery expertise.
- Experience standardizing observability, monitoring, logging, and dashboarding across enterprise environments.
- Hands-on experience with tools such as New Relic, Splunk, Elastic Stack, Prometheus, Grafana, SolarWinds, and cloud-native monitoring platforms.
- Experience with infrastructure automation and cloud management tools including Terraform, Ansible, Vault, and Vagrant.
- Ability to influence stakeholders across technical and business teams.
- Strong communication, leadership, and mentoring skills.
- Willingness to participate in on-call rotations and support critical production environments.
Preferred Qualifications
- Cloud, DevOps, or Site Reliability Engineering certifications.
- Working knowledge of Microsoft Azure.
- Experience within highly regulated, large-scale enterprise environments.
Why Apply?
This is an opportunity to play a pivotal role in shaping the reliability, observability, and operational excellence strategy for a leading enterprise technology organization. You'll work alongside senior engineering leaders, influence large-scale technology initiatives, and drive the adoption of modern SRE practices across a complex global environment.
$159k - $272k
...the opportunity to grow and make a difference in ways that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on...PrincipalFull timePrivate practiceLocal areaRemote workWork from home3 days per week$159k - $272k
...Role SummaryCloud Reliability operates as a... ...and reliability engineering, delivering secure... ...the firm. The Principal Cloud Reliability... ...on reliability, observability, automation, and... ...design authority, site reliability... ...Build and maintain infrastructure as code using Terraform...PrincipalFull timeLocal areaRemote work3 days per week- ...understanding of how reliability, availability, recoverability... ..., roadmaps, and engineering priorities.* Makes... ...and operating cloud infrastructure with senior-level impact... ..., resilience, observability, recovery, incident response... ....* Cloud or site reliability engineering...Principal
$159k - $272k
...SummaryCloud Storage Platform Engineering provides secure, scalable,... ...The Cloud Storage Platform Principal is a senior individual... ...architecture, engineering, and reliability outcomes for enterprise... ...automation while partnering with infrastructure, cloud engineering and...PrincipalFull timeLocal areaRemote work3 days per week$110k - $188k
...Lead Workplace Experience Engineer - Power Platform is responsible... ...platform and AI service reliability, regulatory alignment, and... ...consultants, and represents Infrastructure Operations in architecture,... ...monitoring, telemetry, and observability strategy for Power Platform...SuggestedFull timeLocal areaRemote workWork from home3 days per week- ...T. Rowe Price is seeking a Principal Cloud Reliability Engineer to lead enterprise cloud foundations with... ...guiding teams across applications and infrastructure to deliver scalable cloud... ...and CloudFormation, and advancing observability, incident response, and platform...
$121k - $206k
...you. About the TeamThe AI Engineering and Application Development... ...This team delivers secure, reliable, and forward-looking engineering... ..., build, and implement infrastructure and software solutions for... ...software, tools, and related observability) is sufficiently robust,...Full timeWork experience placementLocal areaRemote work3 days per week$121k - $206k
...the opportunity to grow and make a difference in ways that matter to you. Role Summary We are seeking a hands-on Senior Software Engineer with deep expertise in Oracle Cloud technologies, including ERP, EPM, Oracle Integration Cloud (OIC), OTBI, and APEX. This role is...Full timeLocal areaRemote work3 days per week$170k - $220k
Job DescriptionDewberry is expanding its Energy Market Sector and seeking a Physical Electrical Engineer to lead the growth of our medium and high voltage substation design practice. This is a unique opportunity to build and mentor a high-performing team while delivering...Principal$122k - $209k
...automation, and cloud-native security controls. Working across infrastructure, security, and application teams, you will establish the... ...enterprise standards, assess and mitigate risk, guide critical engineering decisions, and mentor technical talent across the...Full timeLocal areaRemote workWork from home3 days per week- ...Computer Science, Mathematics, Engineering or a related field.Masters... ...must be willing to work on-site in Woodlawn, MD 5 days a week... ...for performance and reliability.Demonstrate a strong understanding... ...using existing IBM DataPower infrastructure. Ensure interoperability with...PrincipalTemporary work
- ...Description Apply now: DevOps Engineer, location is Owings Mills,... ...the cloud-native infrastructure that powers quantitative research... .... Implement monitoring, observability, and alerting using Prometheus... ...teams to improve platform reliability, scalability, and developer...Contract workWork at officeImmediate start2 days per week
$110.6k - $178k
...The Job: As a Senior Manager, Customer Identity & Platform Engineering, you’ll be part of our IT - Customer Engagement team... ...best practices for DevSecOps, CI/CD, automated testing, observability, reliability, and platform performance.Partner with Product Management...Full timeH1bLocal areaRemote work- ...T. Rowe Price is seeking a Cloud Storage Platform Principal to own the architecture, automation, and operational direction for enterprise cloud storage. You will work across NetApp, Nasuni, and AWS storage services to enable CloudNext and data center exit while mentoring...
$145k - $247k
...leader to shape how AI-enabled software engineering evolves across our mobile organization... ...processes Improve software quality, observability, telemetry, and release confidence through... ...experiences, including performance, reliability, usability, telemetry, and release...Full timeLocal areaRemote work3 days per week- ...Omm IT Solutions in Woodlawn, MD is seeking a Principal Software Engineer to design and develop scalable Java-based applications. You will lead microservices using Spring Boot, RESTful APIs, and modern front-end frameworks while aligning with Twelve-Factor App principles...Principal
- ...T. Rowe Price Hong Kong is seeking a senior cloud/site reliability engineer to design, build, and operate enterprise-scale cloud platforms. The role emphasizes reliability, security, and performance across AWS with knowledge of Azure and related tooling. Experience...
- ...Computer Science, Mathematics, Engineering or a related field.Masters... ...must be willing to work on-site in Woodlawn, MD 5 days a week... ...for performance and reliability.Demonstrate a strong understanding... ...using existing IBM DataPower infrastructure. Ensure interoperability with...PrincipalTemporary work
- ...Senior AI Software Engineer at T. Rowe Price will design, build, and scale production-grade AI agents within a Salesforce-centric ecosystem. You will lead technical workstreams, incubate AI products, and partner with tech teams to enable broad adoption at scale. This...
- ...Our client, a leader in cloud infrastructure and enterprise solutions, is seeking a Cloud Engineer III to join their team. As a Cloud Engineer III, you will be part of the Cloud Operations Department supporting cross-functional teams. The ideal candidate will demonstrate...Weekly payTemporary workFlexible hours
$84.84k - $153.55k
...directed by management and senior staff. This position will provide Software solutions delivery support and mentoring for Software Engineers. Lead the design and implementation of software solutions that meet business requirements and technical specifications....Full timeTemporary workWork at officeVisa sponsorshipWork visaFlexible hours$105.5k - $168.8k
...of us—from design and engineering to the manufacturing... ...implementing monitoring, observability, systems management,... ...optimal performance, reliability, and maintainability... ...) Experience with Infrastructure as code (e.g.... ...learn more on our career site under "Our Commitment...Hourly pay- ...Principal Software Engineer Java Woodlawn, MD 5 days onsite Core Tech Stack: Java | OpenShift / AWS | Angular/React| JavaScript | Web Services | Spring Boot | Spring Batch | Typescipt | Microservices | REST/SOAP | OpenShift | Docker | PostgreSQL | DB2 |...Principal
$121k - $206k
...ways that matter to you. Role SummaryAs a Senior AI Software Engineer, you will design, build, and scale production-grade AI agents... ...technical debt and drive ongoing improvements in AI platforms and infrastructure.Proactively seek opportunities to apply Agentic AI...Full timeLocal areaRemote work3 days per week$130k - $140k
...Full-time Description Software Systems Engineer Position Summary Trust Consulting Services, Inc. is seeking a Software... ...engineering changes, and corrective actions that strengthen software reliability, maintainability, performance, and product quality Perform...Full timeTemporary workLocal areaRemote workMonday to Friday$121k - $206k
...investment operations and portfolio of investment applications through key long-term technology initiatives. As a Senior Software Engineer, you will play a critical role in designing and delivering next-generation, cloud-native applications that solve complex business,...Full timeLocal areaRemote work3 days per week$77.5k - $132k
...difference in ways that matter to you. Role Summary Join T. Rowe Price's Global Product Technology team as an Associate Software Engineer to help build the next generation of product data platforms and client experiences. This entry-level role is ideal for early-career...Local areaRemote work3 days per week$159k - $272k
...that matter to you. Role SummaryThe Principal Desktop Engineer is a hands-on engineering leader responsible... ...macOS platforms, virtual desktop infrastructure (VDI), and modern endpoint... ...OS lifecycle compliance, deployment reliability, automation maturity (including zero...PrincipalFull timeContract workLocal areaRemote workWork from home3 days per week- Job-ID27127872Reference26-01207Information Technology - Engineer, Software Sr PURPOSE: Performs complex analysis, design, development... ...applications. Works with cross functional teams to develop highly reliable software that runs at scale. Provides recommendations to infuse...Work experience placement
$143k - $156k
...Shift5 Shift5 is the observability platform for onboard... ..., and other critical infrastructure. Come join us.... ...is seeking a Systems Engineer to join our team. In... ...to support customer site integrations and testing... ...aerospace, defense, or high-reliability electronics sectors....Contract workFor contractorsRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer, Infrastructure Observability. Be the first to apply!
- chief engineer Owings Mills, MD
- principal developer Owings Mills, MD
- general engineer Owings Mills, MD
- engineering director Owings Mills, MD
- hotel chief engineer Owings Mills, MD
- principal engineer Owings Mills, MD
- data center chief engineer Owings Mills, MD
- infrastructure engineer Owings Mills, MD
- infrastructure developer Owings Mills, MD
- senior principal cloud computing engineer Owings Mills, MD



