Site Reliability Engineer
RIT Solutions, Inc.
Site Reliability Engineer
hybird - malvern, pa
needs at least 8 years experience within the US Job Description
The Site Reliability Engineer (SRE) is responsible for improving the reliability, resiliency, observability, and operational excellence of Client's Cash & Money Movement ecosystem. This role serves as the reliability leader for the department, partnering with product and engineering teams to identify gaps, prevent incidents, improve recovery, and ensure critical client journeys remain highly available and resilient.
Key Responsibilities
Observability & Monitoring
This is a Hub-and-Spoke SRE model, SRE defines what "good" looks like and drives continuous improvement while engineering teams remain accountable for execution and results.
SRE owns
hybird - malvern, pa
needs at least 8 years experience within the US Job Description
The Site Reliability Engineer (SRE) is responsible for improving the reliability, resiliency, observability, and operational excellence of Client's Cash & Money Movement ecosystem. This role serves as the reliability leader for the department, partnering with product and engineering teams to identify gaps, prevent incidents, improve recovery, and ensure critical client journeys remain highly available and resilient.
Key Responsibilities
Observability & Monitoring
- Own and maintain a single-pane-of-glass dashboard for application, platform, dependency, and client journey health.
- Improve SLOs, SLIs, alerts, dashboards, and monitoring standards.
- Ensure proactive detection of client-impacting issues using logs, metrics, traces, and synthetic monitoring.
- Improve MTTD, MTTR, and overall service reliability.
- Maintain incident response playbooks and alerting standards.
- Facilitate blameless postmortems, root cause analysis, and track corrective actions through closure.
- Analyze trends and recurring failure patterns to prevent repeat incidents.
- Lead FMEA assessments for critical applications and journeys.
- Identify single points of failure and partner with teams on remediation plans.
- Conduct Game Days, chaos testing, failover testing, and recovery exercises.
- Validate multi-region, multi-AZ, and disaster recovery capabilities.
- Define reliability standards and operational guardrails.
- Review production readiness of high-risk changes.
- Drive adoption of safe deployment practices such as canary releases, feature flags, and automated rollback mechanisms.
- Build and lead the Cash & Money Movement SRE Community of Practice.
- Drive engagement, knowledge sharing, and reliability culture across the organization.
- Identify and mentor application-level SRE champions/POCs.
- Facilitate weekly reliability forums, office hours, and operational reviews.
- Educate teams on SRE best practices, observability, incident management, resilience testing, and safe change principles.
- Partner closely with Danlin Hibay's SRE and operational excellence organizations to stay aligned with enterprise standards, emerging tools, lessons learned, and engineering best practices.
- Act as the liaison between Cash & Money Movement and enterprise SRE communities to bring recommendations, standards, and innovations back to product teams
- Unified Cash & Money Movement Reliability Dashboard
- Journey Health Dashboard (Add Bank, Transfers, Wires, ACH, Direct Deposit, Cash Plus, etc.)
- SLO/SLI Framework and Alert Standards
- FMEA Library and Resiliency Test Plans
- Incident Playbooks and Postmortem Reviews
- Reliability Community of Practice
- Reliability Maturity Assessments and Executive Reporting
- Reduced Sev 1/2/3 incidents
- Reduced MTTD and MTTR
- 100% critical applications with SLOs, dashboards, and actionable alerts
- Completion of FMEA and resiliency testing for critical journeys
- Timely closure of postmortem action items
- Improved reliability, availability, and client experience across Cash & Money Movement.
- Active and engaged reliability community across Cash & Money Movement
This is a Hub-and-Spoke SRE model, SRE defines what "good" looks like and drives continuous improvement while engineering teams remain accountable for execution and results.
SRE owns
- Reliability standards and best practices
- Observability and dashboards
- Assessments, FMEA, and resilience testing
- Incident reviews and postmortems
- Community of Practice
- Education, coaching, and governance
- Reliability backlog execution
- Remediation and implementation
- Operational outcomes
- Service health and reliability improvements
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Malvern, PA vacancy
- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to...SuggestedFull time
- ...Platform Operations Senior Analyst to oversee daily operations of specific applications. This role involves ensuring availability and reliability of applications, integrating operation standards, and working collaboratively in a 24/7 environment. The ideal candidate will...Suggested
$112.5k - $187.5k
...NoticePersonal Information We CollectYour Privacy ChoicesTeam OverviewAt TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering standards, and owns some of the most complex and consequential...SuggestedFull timeTemporary workWork experience placementWork at officeFlexible hours2 days per week- ...We're Hiring: SRE Production Support Engineer Malvern, PA · Onsite Contract Experience: 5+ yrs Skills: Shell, Bash, AWS,... ...Collaborate with development teams to improve application reliability. Support production deployments and release activities....SuggestedPermanent employmentContract workTemporary work
- ...carefully selected external vendor data provider solutions, ensuring reliable, performant, and well‑governed use of data.Collaborate with... ...systems, data pipelines, and compute‑intensive workloadsDrive engineering best practices around resilience, observability, performance,...SuggestedFull timeWork experience placement
- ...Senior Reliability Engineer Hybrid - Malvern, PA Job Description As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements...
- The Platform Engineering Consultant - Cloud and Database Services designs, builds, and operates cloud data infrastructure as a product—delivering secure, reliable, and scalable database services through automation and self-service. This role blends cloud engineering (Infrastructure...
- ...Analytics (CJA), and Adobe Journey Optimizer (AJO)Partner with other teams & departments including Marketing Analytics, Digital, IT/Data Engineering, and Enterprise Architecture to align technical solutions with business objectivesQualifications Tech Stack Adobe Experience...Relocation3 days per week
- ImpactAs a cloud engineer, you will get to work with a variety of cloud native AWS services to develop full stack solutions, that will service global cross functional teams across the Vanguard enterprise. The tech stack requires an engineer very familiar with Python which...Full time
- ...of Vanguard's Chief Technology Office (CTO), the Cloud Platform Engineer will provide technical expertise to design, build, and support... ...scalable cloud platforms and CI/CD solutions that enable secure, reliable, and efficient software delivery. This role will help drive...Full timeWork experience placementWork at office
$140k - $170k
We are looking for a Senior Site Reliability Engineer to work as part of a lean, product‑focused engineering organization. This role is about building and operating reliable cloud‑based systems by writing code, automating infrastructure and delivery workflows, and reducing...Full timeWork experience placementFlexible hours- ...colleague video. An exciting career as an integral part of a world-leading software company providing solutions for architecture, engineering, and construction - watch this short documentary about how we got our start. An attractive salary and benefits package. A...Worldwide
- Software Engineer IILocation: Hybrid - Exton/Philadelphia, PA, US Position Summary:We are looking for a passionate Software Engineer II to help design and build the web-based viewing and collaboration experiences at the heart of Bentley Infrastructure Cloud (BIC). Our...For contractorsWork experience placementWorldwide
$120k - $202.5k
...experience in encryption, key management, cybersecurity, and platform engineering who is passionate about protecting data and enabling secure... ..., and automation initiatives that enhance service reliability, efficiency, and customer experience.Serve as a trusted partner...Full timeTemporary workWork experience placementFlexible hours- The Manager, Platform Engineering, is responsible for leading platform development, maintenance, and support across enterprise systems. This role involves managing technical teams, resolving complex platform issues, implementing infrastructure projects and developing technical...Full timeWork experience placement
- ...includes comprehensive health and wellness care, work-life balance, and an investment in your future at its core.Elasticsearch Lead Engineer - SIEM Platform:Architect and maintain high-availability Elasticsearch clusters supporting large-scale security event...Full timeWork experience placement
- Shape the future of enterprise identity by engineering, securing, and automating mission-critical workforce identity services that enable... ...of Okta capabilities that improve security, resiliency, reliability, and user experience.This role combines identity architecture...Permanent employmentFull time
- Envestnet is seeking an experienced software engineer to design and evolve scalable fintech platforms supporting wealth management and... ...collaborate with product, data science and compliance teams to deliver reliable, low‑latency services in a regulated environment. #J-18808-...
- ...automated testing.Strong knowledge of TDD/BDD methodologies, API testing practices, and validation frameworks.Experience with chaos engineering and large-scale configuration testing preferred.Observability & OperationsExpertise tracing requests across distributed systems...Full timeContract work
- ...portfolio managers, traders, product managers, architects, and fellow engineers to develop innovative, scalable, and resilient solutions that... ...design decisions, ensuring solutions are secure, scalable, reliable, and aligned with enterprise technology standards. Partner...Full timeWork experience placement
- Job ID: 25947711Reference Number: 25-00828Title: Python DeveloperLocation: Malvern, PAPosted Date: 2025-07-22Contact: Mahi MahiContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN Staffing We're looking for a Senior Python Developer with 8-1...
- Job ID: 26778827Reference Number: 25-01602Title: Python developerPosted Date: 2025-12-03Contact: Ekta ChaudharyContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN Staffing Role :: Python Developer with JavaLocation:: Malvern, PA70%on Python...Work experience placement
- ...queues, metadata, bulk actions).Design and maintain REST/CMIS API integrations with internal and external systems.Collaborate with engineering, cloud operations, and product teams to deliver stable, scalable content solutions.Qualifications What We're Looking For a...Contract work
- Job ID: 26747862Reference Number: 25-01558Title: Java Developer :: Hybrid Malvern PAPosted Date: 2025-11-25Contact: Ekta ChaudharyContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN Staffing Responsibilities Build/enhance comparison tool...Work experience placement
- ...technical infrastructure support services for issues elevated from the Support Center and other Technical Services groups. Ensures reliable operation of production. Diagnoses and troubleshoots availability interruptions and other production issues. Plans and...
- ...Platform Engineer Provide technical and cloud engineering support across all Collaboration Platform applications and platforms including Writer, Figma, Adobe Creative Cloud/Firefly, etc. Qualifications: Robust understanding of web development technologies...Shift work
$192.3k - $248.1k
...partners in Pennsylvania. Our team functions as the primary technical engine within the account organization, collaborating closely with... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible to...Full timeTemporary workLocal areaRemote workFlexible hours- ...Senior Software Engineer, AI Code Modernization Location: Hybrid – Exton, PA/Philadelphia Position Summary Bentley is... .... Success is measured through production impact, scalability, reliability, and the ability to accelerate modernization across Bentley's...Worldwide
- ...educators in every school, every day. We keep our product engineering team small and incredibly talent-rich. We believe this helps... ...fast, and earn the trust of our customers through predictable, reliable, and thoughtful interactions. Work at all levels of the stack...Remote workFlexible hours
- ...Job Description Job Description Location: Malvern, PA THE ROLE As a full-time Senior Software Engineer, we are looking for a detail-oriented individual with a strong background in JavaScript, modern React packages, and front and back-end coding. This individual...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- on-site clinical research associate (traveling/remote) Malvern, PA
- construction site safety Malvern, PA
- site safety Malvern, PA
- site reliability engineer remote
- site reliability engineer sre
- site reliability engineering manager
- site reliability engineer
- lead site reliability engineer
- junior site reliability engineer
- site activation specialist


