Site Reliability Engineer
Compunnel
div:has([data-free-thinking-preview-answer=true])+:is(.text-message,.relative:has(> .text-message))]:-mt-2 grow">
Job SummaryThe Site Reliability Engineer will improve the availability, resiliency, recoverability, performance, and operational supportability of critical technology services across on-premises, cloud, network, database, VDI, storage, and vendor-managed environments. The role applies automation, observability, infrastructure engineering, and ITSM practices to reduce recurring incidents, eliminate manual operational work, strengthen recovery readiness, and improve service ownership and accountability.
Key Responsibilities
• Define availability targets, error budgets, and reliability metrics for critical services.
• Partner with Service Owners and technical teams to document ownership, dependencies, escalation paths, runbooks, and recovery requirements.
• Identify single points of failure, fragile dependencies, capacity constraints, and operational risks before they cause outages.
• Prioritize reliability improvements based on business impact, risk, technical feasibility, and cost.
• Implement actionable monitoring, alerting, dashboards, service health indicators, and event correlation across infrastructure, networks, applications, databases, cloud resources, storage, and VDI.
• Reduce alert noise through improved thresholds, dependencies, routing, escalation logic, and baselines.
• Develop early-warning indicators for storage exhaustion, routing instability, database integrity issues, certificate expiration, capacity saturation, and service availability risks.
• Use incident and performance data to identify recurring patterns and reliability improvement opportunities.
• Provide technical leadership during P1 and P2 incidents involving multiple technologies, vendors, or support teams.
• Drive technical workstreams, validate hypotheses, identify dependencies, and maintain focus on service restoration.
• Lead problem investigations for recurring, high-impact, or complex incidents and coordinate root cause analysis.
• Translate root cause findings into corrective-action backlogs with owners, target dates, measurable outcomes, and documented risks.
• Replace repetitive, manual, error-prone, or inefficient operational processes with sustainable automation.
• Build scripts, integrations, workflows, and infrastructure automation using PowerShell, Python, REST APIs, Terraform, Ansible, or comparable technologies.
• Automate diagnostics, service validation, recovery checks, configuration reviews, capacity checks, and remediation workflows.
• Design automation with appropriate testing, security, logging, rollback, and change-control safeguards.
• Define and validate recovery requirements, RTOs, and RPOs for critical business services.
• Design and execute failover, fault-tolerance, recovery, and disaster recovery tests.
• Validate clustering, replication, backups, high availability, and recovery procedures against business requirements.
• Maintain recovery runbooks, dependency maps, test evidence, lessons learned, and remediation actions.
• Partner with DBA teams to improve Microsoft SQL Server reliability, integrity, recoverability, and observability.
• Validate monitoring for database availability, replication, backups, storage, capacity, clustering, and integrity.
• Support testing of SQL Failover Clusters, Availability Groups, secondary-site replication, point-in-time recovery, and database recovery procedures.
• Participate in root cause analysis for database corruption, infrastructure failures, recovery delays, and data consistency events.
• Partner with Network Engineering to improve WAN, SD-WAN, switching, wireless, VPN, routing, Internet connectivity, and remote-access reliability.
• Develop monitoring and validation for BGP route advertisements, convergence, path selection, failover behavior, traffic flow, and asymmetric routing.
• Support fault-tolerance and failover testing for critical network paths and services.
• Develop network diagrams, routing documentation, dependency maps, troubleshooting procedures, and escalation runbooks.
• Apply strong knowledge of TCP/IP, DNS, DHCP, routing, switching, VLANs, VPN, firewalls, wireless, WAN, Internet, and hybrid-cloud connectivity.
• Support SD-WAN environments, preferably Cisco Meraki, as well as enterprise networking platforms such as Cisco and Juniper.
• Use packet captures, flow data, SNMP, syslog, API telemetry, synthetic testing, and path analysis to troubleshoot and monitor network performance.
• Support SASE, ZTNA, and remote-access platforms such as Zscaler and Citrix NetScaler.
• Correlate network behavior with application, database, storage, VDI, and cloud-service impact.
• Support reliability architecture and operational readiness for Azure, hybrid-cloud services, and Azure Virtual Desktop.
• Define monitoring, capacity, availability, security, and recovery requirements for VDI and cloud services.
• Analyze VDI compute, memory, profile, session, storage, and application performance.
• Ensure cloud migrations include operational acceptance criteria, support documentation, observability, recovery testing, and defined ownership.
• Identify unsupported, end-of-life, and approaching-end-of-support technologies across operating systems, virtualization, hardware, databases, network platforms, and supporting services.
• Support modernization planning for Windows Server, Red Hat Enterprise Linux, VMware vSphere, VMware vCenter, Cisco UCS, Citrix, and Microsoft SQL Server.
• Prioritize vulnerabilities based on exploitability, exposure, service criticality, and business impact.
• Support SLA-based vulnerability remediation, exception tracking, and risk acceptance through Ivanti Neurons for ITSM or a comparable ITSM platform.
• Participate in architecture, design, change, and production-readiness reviews for critical services.
• Define reliability acceptance criteria for production implementations.
• Verify that changes include monitoring, testing, rollback, recovery, ownership, documentation, and vendor-support plans.
• Support post-implementation validation and confirm that new services meet availability and support requirements.
• Coordinate reliability initiatives across internal technology teams and external service providers.
• Establish vendor escalation paths, support expectations, communication procedures, and technical accountability.
• Lead technical discussions with cloud, infrastructure, networking, virtualization, storage, security, and application vendors.
• Support RACI models for service ownership, incident response, vulnerability remediation, lifecycle management, and disaster recovery.
Required Qualifications
• Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent professional experience.
• 5+ years of experience in site reliability engineering, infrastructure engineering, DevOps, systems engineering, production engineering, or enterprise operations.
• Experience supporting business-critical production services in hybrid on-premises and cloud environments.
• Hands-on experience with monitoring, observability, alerting, performance analysis, and capacity planning.
• Experience automating operational processes using PowerShell, Python, REST APIs, Terraform, Ansible, or comparable technologies.
• Experience with incident management, root cause analysis, problem management, and corrective-action tracking.
• Knowledge of availability targets, disaster recovery, high availability, and reliability reporting.
• Experience developing technical documentation, dependency maps, operational runbooks, and recovery procedures.
• Working knowledge of Windows Server, Linux, virtualization, cloud, networking, storage, or enterprise database platforms.
• Knowledge of ITIL and ITSM practices, including Incident, Problem, Change, Configuration, and Knowledge Management.
• Strong analytical, troubleshooting, communication, and cross-functional collaboration skills.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Blue Bell, PA vacancy
- ...Senior Site Reliability EngineerLocation: Exton or Philadelphia, PA (Hybrid – 3 times a week in-office)Position SummaryAre you ready to start... ...looking for you!We are looking for a Senior Site Reliability Engineer to take on the responsibility of automating cloud-based...SuggestedCasual workWork at officeWorldwide
- ...production processes reduce manual intervention and increase platform reliability Identify vulnerabilities and continuous improvement... ...reporting requirements This role requires three days per week on-site. Basic Qualifications: Minimum 5 years of experience lead...SuggestedHourly payLive inWork at officeLocal areaFlexible hours3 days per week
$95.3k - $158.8k
...Senior Site Reliability Engineer Are you passionate about building resilient, scalable systems that power mission-critical applications? Do you thrive on automating operations, improving reliability, and ensuring exceptional system performance? About the team:...SuggestedFull timeLocal areaRemote work$145k - $160k
...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep...SuggestedTemporary workRemote workFlexible hours- ...include global situational awareness, counter drone systems, edge network sensors and RF communications. What You’ll Get to Do: Engineer complex multidisciplinary solutions for customers that span hardware interfaces, mobile applications, web applications, and edge...SuggestedFull timeTemporary workImmediate start
- ...solutions for the consulting, architecture, engineering, project controls, procurement,... ...project. You may be assigned to a client site for an extended period. Overnight travel... ...safety rules. Must have access to reliable transportation. Must have the ability...For contractorsNight shift
- ...change people's lives. We are the leading provider of sustainable Engineering, Architecture, Construction and Consulting solutions to the... ...executed by or under the supervision of the Senior Lead include site selection, master planning, commissioning and startup, load calculations...Work at office
$1,000 per month
...technology and customer-centric solutions. Overview As a Senior Backend Engineer on the Trust Platform team, you'll play a pivotal role in... ...coding practices. Experience with designing resilient and reliable systems that meet high security and regulatory standards. Nice...Temporary workWork at officeImmediate startRemote workFlexible hours- ...development as required. Participate in technical design discussions, code reviews, and architecture decisions. Improve AI-assisted engineering workflows and development standards. Collaborate with architects, engineers, product teams, and adjacent teams to deliver...Contract work
$126.23k - $210.38k
...organization, apply now.We are currently seeking a Lead Security Engineer to join our team in Fort Washington, Pennsylvania (US-PA),... ...tools healthy and effective.This is a hybrid role with 3 days on-site in the office and 2 days remote. Only local candidates will be considered...Temporary workWork at officeLocal areaRemote workFlexible hours- ...demand. We invite you to bring your bold ideas and big dreams and become part of a global team of over 50,000 planners, designers, engineers, scientists, digital innovators, program and construction managers and other professionals delivering projects that create a...Full timeWork experience placementWork at officeLocal areaRemote workWork from homeWorldwideRelocationHome officeFlexible hours
$90k - $110k
Piper Companies is seeking a Platform Engineer to support a company focused on modernizing enterprise infrastructure and strengthening... ...will lead the engineering, automation, and operational reliability of core Active Directory and certificate services. Join a mission...- Merck & Co. is seeking a Senior Specialist, Cybersecurity Engineering to secure our strategic enterprise platforms through hands-on engineering, automation, and risk reduction. You will partner with platform teams to embed security controls and sustain trusted platform...
$112.5k - $202.9k
...T-Mobile Solutions Engineer At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we... ...Solutions Engineer will also participate in customer meetings, site visits, technical presentations, and other field activities as needed...Full timeTemporary workPart timeWork experience placementWork at officeLocal areaFlexible hours2 days per week$110k - $130k
Senior Software Engineer Position Summary The role plays a critical part in advancing Morgan Properties’ technology capabilities by... ...emphasis on database efficiency, query governance, and system reliability. Responsibilities include diagnosing production issues, analyzing...Contract workTemporary workWork at office$91.7k - $163.7k
...help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. As a Senior Software Engineer specializing in data engineering and intelligent solutions, you will design, build, and deploy enterprise-scale software, cloud...Minimum wageFull timeTemporary workWork experience placementLocal area- Software Engineer We are looking for a new Software Engineer to build the new Novovu workshop. Responsibilities: Familiar with the software development life cycle (SDLC) from analysis to deployment. Comply with coding standards and technical design. Believes in systematic...Remote job
- Title: Software Engineer Location: San Antonio, TX - On-Site the hiring manager is willing to take a candidate who has LabVIEW experience from coursework during college or someone who has not worked with LabVIEW professionally but had college projects using it. They would...Currently hiring
$70k - $80k
...proven leadership and best‑in‑class technical expertise to deliver innovative real estate solutions. Job Description The Building Engineer addresses and resolves tenant/customer concerns regarding the building or their space. Essential Duties & Responsibilities Perform...Temporary workFor contractorsWork at officeLocal areaFlexible hoursShift workWeekend work- ...Cross-functional influence The person willlikely needto drive alignment across business units, IT, governance, privacy, HR, and engineering without formal authority. Business-value and ROI thinking A candidate must be able to distinguish between usage metrics and business...Full timeLocal area
- ...Description This position is hybrid and will require 3 days on site at one of the following Quest sites: Secaucus, NJ, Schaumburg,... ...Career advancement opportunities and so much more! DevSecOps Engineer II plays a critical role in implementing application solutions...Full timeTemporary workPart timeWork experience placementFlexible hoursShift work3 days per week
- ...surgery so patients can resume their lives as quickly as possible. Position Summary: The Senior Systems Engineer will design, develop, and deliver high-reliability Class III implantable neuromodulation systems and associated external accessories used in spinal cord...Full time
- ...in a meaningful career. Embrace the chance to drive change with M3 USA. Due to our continued growth, we are hiring for an AI Engineer. Job Description The AI Engineer will build AI-powered systems that automate and improve workflows for M3 USA’s healthcare...Work experience placementLocal areaRemote workFlexible hours
- ...Germany, Brazil, Sweden, China, USA, and South Korea, as well as India. Due to our continued growth, we are hiring for a AI Engineering Intern to join M3 USA. Job Description Build and ship ~ Own a project end to end: understand the problem, build a...InternshipRemote work
$102k - $137k
...Software Developer to join our growing AI Engineering team and help shape the future of... ...Evaluate and improve model performance, reliability, accuracy, and observability through monitoring... ...! Click here to be directed to our site that is dedicated to veterans and transitioning...Work at officeLocal area- Merck is seeking a Senior Specialist, Cybersecurity Engineering to secure enterprise platforms through hands-on security controls, automation, and continuous risk reduction. You will partner with platform teams to embed security in delivery, improve resilience, and standardize...
- ...discretionary assets, as of December 31, 2025.**The Opportunity:**The Head of Application Engineering will serve as a senior execution leader responsible for the continued growth, delivery, reliability, quality, and engineering operating discipline of Hamilton Lane’s Cobalt...Work experience placementWork at office
- ...Forward Deployed AI EngineerThe Forward Deployed AI Engineer is responsible for designing, developing, and deploying data and AI-driven... ..., and financial applications.Ensure data quality, lineage, reliability, observability, and performance across production environments...Work at officeLocal area
- ...Hamilton Lane Incorporated is seeking a Head of Application Engineering in Conshohocken, PA to lead a global engineering organization across application engineering, DevOps, and QA for the Cobalt platform. You will shape engineering strategy, drive end-to-end delivery,...
- Java Developer With Scala Role: Jr. & Sr. Java with Scala Developer Location: Plymouth Meeting, PA Duration: 6 Months Contract (will continually extend) Mode of Interview: Face To Face After Phone Minimum Education, Experience, & Specialized Knowledge Required: • Computer...Contract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- on-site clinical research associate (traveling/remote) Blue Bell, PA
- construction site safety Blue Bell, PA
- junior site reliability engineer
- site reliability engineering manager
- site reliability engineer
- lead site reliability engineer
- site reliability engineer remote
- site reliability engineer sre
- website qa testing
- site buyer





