Manager Site Reliability Engineering
$51.9 per hourHighmark Health
Company :
Allegheny Health Network
Job Description :
GENERAL OVERVIEW:
This job is responsible for the reliability, availability, and performance of critical healthcare IT systems, principally in the Environment of Care (EOC), enabling seamless access to essential services for patients, providers, and the people we serve. Proactively identifies and mitigates potential disruptions to maintain the highest standards of care and operational efficiency. This role blends software engineering, clinical engineering, and security principles with a deep understanding of healthcare operations to minimize downtime, improve system resilience, and to support clinical workflows and continuity of hospital operations. Works cross-functionally with AHN site leaders and teams to navigate and to monitor and support building automation and facility systems, clinical engineering / IoT, healthcare delivery technology architecture, infrastructure and platform operations, and cybersecurity. Fosters a culture of automation, continuous improvement, collaboration, and patient safety. Develops core metrics for monitoring and maintaining system health for SRE practitioners (e.g., latency, traffic, errors, and saturation) leveraging industry practices, manufacturer guidance, and other service delivery metrics.
ESSENTIAL RESPONSIBILITIES
Perform management responsibilities to include, but are not limited to: involved in hiring and termination decisions, coaching and development, rewards and recognition, performance management and staff productivity.Plan, organize, staff, direct and control the day-to-day operations of the department; develop and implement policies and programs as necessary; may have budgetary responsibility and authority. (25%)
Oversees the partnership with clinical engineering, cybersecurity, device manufacturers, suppliers, and Information Technology SMEs to oversee and to implement strategies for managing, monitoring, and securing a diverse range of clinical devices and other technology equipment (e.g., IoT), ensuring compliance with HIPAA and other relevant regulations (e.g., FDA, TJC, PCI). Keeps current on healthcare IT trends, including AI, security patching, and best practices for device hardening. Oversees and assists with network segmentation and access controls to isolate and to protect clinical and other critical devices. Automates monitoring tasks to improve efficiency and reduce errors. Identifies and remediates vulnerabilities in clinical devices and related infrastructure. Manages and reports issues with assets, devices, integration services, and other equipment. Engages the appropriate parties to develop and deploy a fix/solution or oversees ownership of resolution actions. Utilizes observability practices to gain deep insights into system behavior, enabling faster identification and resolution of issues. (15%)
Oversees the SRE partnership with Clinical Engineering and Cybersecurity Engineering to troubleshoot technical issues related to medical equipment and systems. Participates in the medical device technology lifecycle – from product/device evaluation, discovery, to implementation, maintenance, and through retirement. Develops the framework and structure to maintain documentation related to the IT infrastructure supporting clinicaland other critical devices. Participates in the planning and oversees the execution of preventative maintenance activities. Provides direction and guidance to team members on how to analyze complex problems and develop effective solutions, how to troubleshoot system outages and performance issues, and how to work collaboratively with other IT, cybersecurity, facility, AI and application teams to resolve issues and to conduct root cause analyses. (15%)
Oversees the SRE partnership with facility leaders to optimize the performance and monitoring of building automation systems (BAS), including HVAC, lighting, fire suppression, security systems, etc. Manages processes and procedures used to monitor BAS performance metrics and proactively identifies potential issues. Works with facilities management to implement improvements to the BAS infrastructure. Works with cybersecurity, vendors/manufacturers, et. al. to ensure the security of building automation systems and oversees monitoring of performance, service delivery, and support. (15%)
Oversees the SRE partnership with IT teams including, but not limited to platform / product management, disaster recovery services, infrastructure and architecture, storage management, and release management. Participates in the planning and execution of downtime drills and system / device recovery exercises. Supports other emergency preparedness drills and exercises, as needed. Leads or participates in post-incident reviews to identify root causes and implement corrective actions. Works with cross-functional stakeholders to Implement and to maintain redundant systems and failover mechanisms to minimize downtime. Reviews and provides feedback on emergency operations plans and other materials which are used to respond to emergency situations (e.g., Continuity of Operations Plans, Incident Response Guides, Downtime Procedures). Manages team members who are supporting the planning and execution of system migrations, releases, and upgrades to ensure minimal disruption to clinical operations. Oversees detailed migration or installation plans, including risk assessments, rollback procedures, and communication strategies. Assists local site leaders with navigating shared services (e.g., AI, IT, Information Security, Clinical Engineering, Platform Operations, Technology Acquisition). (15%)
Establishes core metrics for monitoring and maintaining system health for SRE practitioners (e.g., latency, traffic, errors, and saturation). Manages the processes and procedures used for documentation and knowledge sharing including maintaining detailed documentation of systems, device inventories, processes, and procedures.Leads by example by sharing knowledge and best practices with other staff and cross-functional teams. Provides training and mentorship to junior or less experienced team members. Stays current with the latest technologies and trends in site reliability engineering. Leads or participates in briefings with cross-functional stakeholders to manage priorities and team assignments, support ticket queues, etc. (10%)
Other duties as assigned or requested. (5%)
Q UALIFICATIONS:
Required
Bachelor’s degree in Computer Science, Engineering, Management Information Systems, IT, or related field or relevant experience and/or education as determined by the company in lieu of bachelor's degree.
3 years with Management or leadership role
Preferred
Master's degree in Computer Science, Engineering, Management Information Systems, IT, or related field
5 years of experience with Site Reliability Engineering (SRE), Systems Administration, or DevOps particularly in healthcare IT
5 years of experience in Medical device management lifecycle, network / device segmentation, vulnerability and patch management
5 years of experience in Healthcare IT experience in architecture, automation, IoT, telemetry, telehealth, security, system development lifecycle, capacity planning, networking, continuous integration / continuous delivery pipelines (CI/CD), incident management, scripting, metrics, monitoring, redundancy, etc.
3 years of experience working in highly regulated environments
3 years of experience with Progressive leadership roles, preferably inclinical engineering, IT, business continuity, backup and storage management, building automation, or cybersecurity discipline in healthcare
SKILLS:
Problem-Solving: Excellent analytical and troubleshooting skills; High capacity to think analytically, interpret information / observations, apply judgment and to assist with making effective, strategic decisions.
Collaboration: Ability to work effectively in a team environment; demonstrated ability to support multiple sites and locations while maintaining consistency in service delivery processes and procedures.
Communication: Strong written and verbal communication skills.
Flexibility: Willingness to participate in activities or incidents which may occur outside of regular work schedules.
Leadership: Demonstrated resource and project planning capabilities, decision making skills, history of results-oriented delivery, and effective team building across multiple locations and a diverse team of staff, partners, and stakeholders.
Security Awareness: Understanding of security best practices and how to apply them in a healthcare IT environment.
Delivery and Execution: Demonstrated competency in the execution of multiple projects, including managing resources across multiple projects to meet goals.
Relationships: Strong relationship building skills and ability to influence with and without authority in a matrixed organization.
Disclaimer: The job description has been designed to indicate the general nature and essential duties and responsibilities of work performed by employees within this job title. It may not contain a comprehensive inventory of all duties, responsibilities, and qualifications required of employees to do this job.
Compliance Requirement : This job adheres to the ethical and legal standards and behavioral expectations as set forth in the code of business conduct and company policies.
As a component of job responsibilities, employees may have access to covered information, cardholder data, or other confidential customer information that must be protected at all times. In connection with this, all employees must comply with both the Health Insurance Portability Accountability Act of 1996 (HIPAA) as described in the Notice of Privacy Practices and Privacy Policies and Procedures as well as all data security guidelines established within the Company’s Handbook of Privacy Policies and Practices and Information Security Policy.
Furthermore, it is every employee’s responsibility to comply with the company’s Code of Business Conduct. This includes but is not limited to adherence to applicable federal and state laws, rules, and regulations as well as company policies and training requirements.
Pay Range Minimum:
$51.90
Pay Range Maximum:
$83.84
Base pay is determined by a variety of factors including a candidate’s qualifications, experience, and expected contributions, as well as internal peer equity, market, and business considerations. The displayed salary range does not reflect any geographic differential Highmark may apply for certain locations based upon comparative markets.
Highmark Health and its affiliates prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities and prohibit discrimination against all individuals based on any category protected by applicable federal, state, or local law.
We endeavor to make this site accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact the email below.
For accommodation requests, please contact HR Services Online at View email address on click.appcast.io
California Consumer Privacy Act Employees, Contractors, and Applicants Notice
Req ID: J280531
$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's... ...inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...integrating into Akamai's existing incident management processes Contributing to SLO tracking...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$146.4k - $263.6k
...edge technology? Do you enjoy working with a diverse multi-national team of engineering talents? Join our highly skilled Site Reliability team Our team designs, develops, and manages applications and infrastructure that support Akamai's Compute products and services...SuggestedWork experience placementWork at office$121.5k - $306.4k
...and provides input on best practices for reliability and functionality. Establishes direction... ...the impact of changes, mentoring other managers on what to communicate. Defines approaches... ..., executing improvements, building site reliability knowledge, and providing clear...SuggestedTemporary workFlexible hours$51.9 per hour
...Company: Allegheny Health Network Job Title: Site Reliability Engineering – Clinical & Facility Services General Overview This role ensures the... ...and support clinical workflows. Essential Responsibilities Manage the department through hiring, coaching, performance management...SuggestedLocal area$103.71k - $138.28k
...and experience in system architecture and engineering disciplines. Specific technical... ...Supports due diligence activities including site surveys, design, design review, bill of... ...experience to include indexing, clustering, managing, and troubleshooting. 5+ years with automation...SuggestedTemporary workRemote work$127.1k - $198.58k
...position contributes to team efforts in engineering, analytics, and technical planning, applying... ...bring the best of scientific thought, management, and engineering expertise together in... ...visiting the Benefits ( page on our Careers ( site. Compensation at Noblis Compensation at...Permanent employmentFull timeContract workPart timeLocal areaRemote work$125k - $191.7k
...categorized as hybrid/Remote Role: As a Senior Software Systems Engineer on the Software Validation team within the AV organization,... ...Industry experience in system engineering and requirements management including system analysis, requirements authoring, test generation...Local areaRemote workWork from homeFlexible hours$80k
...Duties and Responsibilities: Provide Tier‑3 engineering support for Microsoft 365 GCC, Exchange... ...availability, performance, and security. Manage, monitor, restore, and optimize... ...SharePoint Online platform operations, including site collections, permissions, integrations, and...Contract work$94.1k - $150k
...The Platform Engineer (Ops Technology Lead) is responsible for designing, implementing, and... ...within the CASTLE-NET program, ensuring reliability, scalability, and security. This role supports application deployment and management, ensures compliance with CASTLE-NET policies...Contract workWork at office$75k - $110k
...We are looking for a Customer Solutions Engineer to provide business and technical support... ...operations. The role includes project management, implementation, and support of software... ...detailing customer requirements. Perform on‑site presentations, product demonstrations,...Work at officeRemote workWork from home- ...more technical specializations. As a technical leader, they mentor others and share their knowledge within the Red River Solution Engineering and Sales teams. They provide architectural guidance for customers across the Red River product portfolio, including products and...Work experience placementWork at office
- Our Benefits - Designed with You in Mind Comprehensive Health & Well-being Coverage From your very first day, you’ll have access to medical, dental, vision, and prescription drug coverage - ensuring you and your family stay healthy and protected. Generous Paid Time...Full timeImmediate start
- ...components including vSphere, NSX, and SDDC Manager. Assist with security configuration... ...control changes with guidance from senior engineers. Document security configurations, standard... ...security issues. Schedule & Presence: This on‑site role supports 24/7 operations through...Full timeTemporary workMonday to FridayShift work
- ...The Systems Administrator, Senior manages and optimizes complex enterprise infrastructure spanning Windows and Linux servers, VMware environments, and cloud services that support mission‑critical government systems. The role leads major changes, automation, and incident...Contract workWork at office
$180k - $220k
...to realize our bold vision for healthcare. Senior Software Engineer The Role As a Senior Software Engineer, you will lead major... ...initiatives that advance Datavant’s platform scalability and reliability. You’ll drive technical design, coach peers, and ensure system...$84.98k - $111.5k
...system. This position provides journey-level expertise to all software development functions. Reporting to an Information Technology Manager, Information Technology Supervisor, or equivalent, this is a journey-level position that works within a team supporting the...Full timeWork at officeRemote workNight shift2 days per week- ...within enterprise platforms and IT Service Management (ITSM). This position offers structured... ...while supporting the delivery of reliable, high‑quality IT services. In this role,... ...applications on the ServiceNow platform using App Engine Studio. Support implementation and...Full timeContract workPart timeInternshipFlexible hours
$112.3k - $140k
...looking for a stimulating and challenging career where your engineering expertise can be leveraged to create, enhance, and maintain our... ...solutions in automatic tank gauging, dispensing, and fuel management systems driven by innovation and break-through thinking. We are...Local area$114.6k - $234.6k
...operational goals, sharing results with manager upon completion. -Adheres to and improves... ...; provides guidance and coaching to engineers to drive improvements. -Utilizes advanced... ...availability, health, support, and reliability. Core Responsibilities Planning &...Temporary workFlexible hoursShift work- ...Overview of Job Function As a Software Engineer, you will be a core contributor to Verint... ...systems, and collaborate daily with Product Managers, Designers, QA Engineers, and globally... ...Actions, or Azure DevOps — ensuring reliable, automated build-test-deploy workflows....Local areaWorldwideShift work
$197.4k - $232k
...Type: FullTime Location Type: Remote Department Engineering Compensation: $197.4K – $232K • Offers Equity At... ...environment. Make architecture and technical decisions that balance reliability, scalability, performance, and operability, and clearly...Full timeRemote work- ...documentation of test results. Correspond with Software and Template Engineering, give feedback to any questions they may have, and facilitate... ...product line. Additional responsibilities as defined by management. Comply with all applicable U.S. Food and Drug...For contractorsWork experience placementWork at officeLocal areaRemote workFlexible hours
$89.2k - $209.5k
...Role Summary Oracle Health Platform Engineering builds core platform capabilities that... ...productivity and strengthen platform security and reliability. Responsibilities Key... ...Qualifications • Identity and Access Management (IAM) concepts: authentication, authorization...Temporary workFlexible hours$83.43k - $222.48k
...world. Currently, we are seeking a Senior Software Development Engineer with deep expertise in Java/JEE, Spring Boot, RESTful... ...hands‑on experience with Data Migration, ETL pipelines, and data management best practices, ideally in a Healthcare/EHR integration context...Hourly payFull timeTemporary workLocal area$30 per hour
...unique opportunities for smart, hands-on engineers with the expertise and passion to solve... ...Being empowered with the flexibility, reliability, and scalability of Virtual Networking,... ...technical and non-technical audiences (management, peers) Understanding of agile...Hourly payTemporary workInternshipFlexible hours$135.2k - $306.4k
...technical and business challenges. Oracle Kubernetes Engine (OKE) is OCI's managed Kubernetes service. OKE enables customers to create, run,... ...cluster lifecycle management, orchestration, scalability, reliability, performance, automation, observability, security, and integration...Temporary workRemote workFlexible hours$130k - $180k
...TIC) is a leading provider of engineering and consulting services for... ...surveyors, architects, project managers, and environmental... ...~ Construction Staking (site development and highway) ~... ...attendance, punctuality, and reliability. Willingness and ability...Work at officeLocal areaRemote work$79.4k - $135k
...Position Overview The Incident Manager, Mid leads the full lifecycle of IT incidents and service requests to restore normal operations quickly and minimize disruption to mission‑critical systems. This role oversees day‑to‑day execution of the incident management process...Contract workWork experience placementWork at office$62.2k - $105.7k
...Position Overview The Incident Manager oversees the end‑to‑end lifecycle of IT incidents in an enterprise environment, ensuring rapid restoration of normal service with minimal disruption to mission‑critical systems. The role coordinates cross‑functional technical teams...Contract workWork experience placementWork at office$169.8k - $355.4k
...and debug software programs for databases, applications, tools, networks etc. Responsibilities As a member of the software engineering division, you will take an active role in the definition and evolution of standard practices and procedures. Suggest and justify...Temporary workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager Site Reliability Engineering. Be the first to apply!
- IT site lead Olympia, WA
- junior website developer Olympia, WA
- site safety Olympia, WA
- site leader Olympia, WA
- on-site clinical research associate (traveling/remote) Olympia, WA
- site reliability engineer
- junior site reliability engineer
- site reliability engineer remote
- site reliability engineering manager
- site reliability engineer sre


