SRE
RIT Solutions
SRE Hybrid - Malvern, PA
As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance innovative thinking with pragmatism and understand the long-term impacts of technical decisions. You communicate complex ideas clearly and collaborate effectively to deliver scalable solutions.
Core Responsibilities
Team is focused on automating incident response and infrastructure management. While Java and Python receive a stronger emphasis, candidates with solid programming fundamentals in any language and the ability to adapt will be considered. Experience with AWS and event-driven architectures is also valuable. From a technical standpoint, familiarity with observability concepts (e.g., distributed tracing) and tools like Prometheus or Grafana is beneficial, though not mandatory. More important is an understanding of the underlying principles, such as instrumentation and monitoring strategies.
- Improve resiliency engineering practices across platforms and applications, including resilient application design patterns, system observability and deployment strategies
- Incident detection, troubleshooting, and resolution.
- Develop automation for incident response and infrastructure management
- Develop and support OpenTelemetry integrations for multiple application platforms (browser, ECS, lambda, etc) and languages (JavaScript, Java)
- Contribute to architectural decisions and support implementation of solutions.
Skills and Qualifications
* Deep knowledge of Java or Javascript. Practical experience developing and operating software in distributed systems environments. * Problem-solving and analytical thinking: ability to diagnose complex issues and propose efficient solutions. Strong debugging and optimization skills for performance and scalability. * Cloud platforms: Hands-on experience with AWS services and cloud infrastructure * System architecture and design: ability to design scalable, secure, and maintainable systems. * Working knowledge of Python (or similar scripting language). * Strong knowledge of resiliency engineering techniques for both platforms and applications. * Experience troubleshooting complex production issues and implementing effective mitigations. * Familiarity with OpenTelemetry specification and core APIs. From a screening perspective, we recommend focusing on: · How candidates approach software releases and validate functionality · Their understanding of system dependencies and fault tolerance · Experience with diagnosing and resolving production issues · Their ability to reflect on past incidents and identify improvements · Evidence of systems thinking and architectural awareness
- Job ID: 20139475Reference Number: 23-01136Title: Devops/SRELocation: Malvern, PA, 07512Posted Date: 2023-07-12Company: HAN Staffing Migrated key systems from on-prem hosting to AWS Worked in Agile and Scrum methodologies/practices Worked on designing and developing a multitude...Suggested
- ...We're Hiring: SRE Production Support Engineer Malvern, PA · Onsite Contract Experience: 5+ yrs Skills: Shell, Bash, AWS, Azure, Docker, Kubernetes, Terraform, Ansible, Jenkins, CI/CD, Nginx, Grafana Roles & Responsibilities: Provide L2/L3 production support...SuggestedPermanent employmentContract work
- ...modernization for our channel ecosystem (e.g., Genesys, Sierra, Text/SMS, and Email), driving best‑in‑class Site Reliability Engineering (SRE) practices, observability, and operational excellence. You’ll partner across product, engineering, security/risk, analytic and...SuggestedFull time
- ...operability is designed in, not bolted on—shifting left on reliability, monitoring, and remediation.Champion Site Reliability Engineering (SRE) and ITSM best practices within IAM operations.Global Delivery & Service OwnershipOversee incident, problem, change, and service...SuggestedFull timeLocal areaShift work
- ...Kubernetes & containerized identity servicesAutomate provisioning, deployment, monitoring, and drift detection for identity platforms.Support SRE‑style operational maturity: SLIs/SLOs, alerting, incident response, and runbooks for identity services.Security, Risk &...SuggestedFull timeVisa sponsorship
$140k - $170k
...from you: BA/BS, in a related technical field; or the equivalent in education and work experience 8+ years of experience in DevOps, SRE, platform engineering, or similar roles supporting application teams running production services Strong CI/CD experience (Jenkins and...Full timeWork experience placementFlexible hours- ...PostResponsibilitiesOperate in a highly collaborative team emphasizing best practices in software development, automation, DevSecOps and SRE.Build cloud-first, consumer-focused and applying lean agile methodologies.Ability to learn and build in third party application that...
- ...Site Reliability Engineer Hybrid - Malvern, PA needs at least 8 years experience within the US The Site Reliability Engineer (SRE) is responsible for improving the reliability, resiliency, observability, and operational excellence of Client's Cash & Money Movement...Work at office
- ...infrastructure; distinguish product defects from test/environment issues and drive root-cause analysis.Partner with Software Engineering, SRE/DevOps, Security, and Product to define software quality reports, quality gates, release criteria, and automation standards....Full timeContract workTemporary workH1bWork at officeMonday to FridayShift work
- ...upgrade programs aligned with vendor roadmaps.Build and lead a Platform Reliability Engineering (PRE) or Site Reliability Engineering (SRE) function focused on proactive monitoring, automation, and resilience.Implement enterprise monitoring, observability, and event...
- ...frameworks, including GDPR, CCPA, HIPAA, or similar standards. Experience working in Agile, DevOps, or Site Reliability Engineering (SRE) environments.Familiarity with scripting or automation technologies such as PowerShell, Python, or similar languages. Success Factors...Full timeLocal area
$152.07k - $202.76k
...EKS, S3, CloudFront). Ensure compliance with security, cost, and operational governance policies. Collaborate with DevOps and SRE teams to embed reliability and observability into architectural decisions. Continuous Improvement Identify opportunities for performance...Full timeTemporary workRemote work$135.7k - $251.9k
...Development experience with any of the following programming languages: Java or Python• Recent on-program Site Reliability Engineering (SRE) experience (e.g., automation, incident management, monitoring, optimization, programming/scripting, metrics, security, etc.)•...Full timeTemporary workWork experience placementCasual workFlexible hours- ...Clearance ~ BS in Software Engineering or related field ~2-10 years of experience in CI/CD and Kubernetes-focused engineering or SRE/Platform Engineering ~ Strong Kubernetes skills: core objects, scheduling, RBAC, networking, storage, upgrades, troubleshooting ~...Flexible hours
$68 - $85 per hour
...nexs-riven qestions re rotine prt of the work, not n nnl exception. Yo will work cross or service, project, n procrement tems to mke sre trnsctions re recore ccrtely n on time, n tht billing reches the cstomer clen the first time. Key Responsibilities Csh n...- ...basic observability (logging, monitoring, alerting) for owned services. Supports production troubleshooting in collaboration with DevOps/SRE teams. Designs and delivers scalable fintech product services and distributed systems with a focus on performance, resilience, and...Temporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!


