SRE
RIT Solutions
SRE Hybrid - Malvern, PA
As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance innovative thinking with pragmatism and understand the long-term impacts of technical decisions. You communicate complex ideas clearly and collaborate effectively to deliver scalable solutions.
Core Responsibilities
Team is focused on automating incident response and infrastructure management. While Java and Python receive a stronger emphasis, candidates with solid programming fundamentals in any language and the ability to adapt will be considered. Experience with AWS and event-driven architectures is also valuable. From a technical standpoint, familiarity with observability concepts (e.g., distributed tracing) and tools like Prometheus or Grafana is beneficial, though not mandatory. More important is an understanding of the underlying principles, such as instrumentation and monitoring strategies.
- Improve resiliency engineering practices across platforms and applications, including resilient application design patterns, system observability and deployment strategies
- Incident detection, troubleshooting, and resolution.
- Develop automation for incident response and infrastructure management
- Develop and support OpenTelemetry integrations for multiple application platforms (browser, ECS, lambda, etc) and languages (JavaScript, Java)
- Contribute to architectural decisions and support implementation of solutions.
Skills and Qualifications
* Deep knowledge of Java or Javascript. Practical experience developing and operating software in distributed systems environments. * Problem-solving and analytical thinking: ability to diagnose complex issues and propose efficient solutions. Strong debugging and optimization skills for performance and scalability. * Cloud platforms: Hands-on experience with AWS services and cloud infrastructure * System architecture and design: ability to design scalable, secure, and maintainable systems. * Working knowledge of Python (or similar scripting language). * Strong knowledge of resiliency engineering techniques for both platforms and applications. * Experience troubleshooting complex production issues and implementing effective mitigations. * Familiarity with OpenTelemetry specification and core APIs. From a screening perspective, we recommend focusing on: · How candidates approach software releases and validate functionality · Their understanding of system dependencies and fault tolerance · Experience with diagnosing and resolving production issues · Their ability to reflect on past incidents and identify improvements · Evidence of systems thinking and architectural awareness
- Job ID: 20139475Reference Number: 23-01136Title: Devops/SRELocation: Malvern, PA, 07512Posted Date: 2023-07-12Company: HAN Staffing Migrated key systems from on-prem hosting to AWS Worked in Agile and Scrum methodologies/practices Worked on designing and developing a multitude...Suggested
- ...We're Hiring: SRE Production Support Engineer Malvern, PA · Onsite Contract Experience: 5+ yrs Skills: Shell, Bash, AWS, Azure, Docker, Kubernetes, Terraform, Ansible, Jenkins, CI/CD, Nginx, Grafana Roles & Responsibilities: Provide L2/L3 production support...SuggestedPermanent employmentContract work
- ...modernization for our channel ecosystem (e.g., Genesys, Sierra, Text/SMS, and Email), driving best‑in‑class Site Reliability Engineering (SRE) practices, observability, and operational excellence. You’ll partner across product, engineering, security/risk, analytic and...SuggestedFull time
- ...connect to enterprise infrastructure (CI/CD, CMDB, ITSM, identity providers).Partners with engineering leadership across Platform, DevOps, SRE, and application security teams to align on shared interfaces, data contracts, and remediation workflows that reduce friction at...SuggestedFull timeWork experience placementImmediate startShift work
- ...operability is designed in, not bolted on—shifting left on reliability, monitoring, and remediation.Champion Site Reliability Engineering (SRE) and ITSM best practices within IAM operations.Global Delivery & Service OwnershipOversee incident, problem, change, and service...SuggestedFull timeLocal areaShift work
- ...Kubernetes & containerized identity servicesAutomate provisioning, deployment, monitoring, and drift detection for identity platforms.Support SRE‑style operational maturity: SLIs/SLOs, alerting, incident response, and runbooks for identity services.Security, Risk &...Full timeVisa sponsorship
$140k - $170k
...from you: BA/BS, in a related technical field; or the equivalent in education and work experience 8+ years of experience in DevOps, SRE, platform engineering, or similar roles supporting application teams running production services Strong CI/CD experience (Jenkins and...Full timeWork experience placementFlexible hours- Job Title Required Skills: Has exceptional coding skills— Java, python Has AWS knowledge and thinks out of the box SRE knowledge would be a nice to have Somebody who has a drive and willing to learn SRE and can do automation
- ...Site Reliability Engineer Hybrid - Malvern, PA needs at least 8 years experience within the US The Site Reliability Engineer (SRE) is responsible for improving the reliability, resiliency, observability, and operational excellence of Client's Cash & Money Movement...Work at office
- ...Responsibilities Operate in a highly collaborative team emphasizing best practices in software development, automation, DevSecOps and SRE. Build cloud-first, consumer-focused and applying lean agile methodologies. Ability to learn and build in third party...
- ...infrastructure; distinguish product defects from test/environment issues and drive root-cause analysis.Partner with Software Engineering, SRE/DevOps, Security, and Product to define software quality reports, quality gates, release criteria, and automation standards....Full timeContract workTemporary workH1bWork at officeMonday to FridayShift work
- ...upgrade programs aligned with vendor roadmaps.Build and lead a Platform Reliability Engineering (PRE) or Site Reliability Engineering (SRE) function focused on proactive monitoring, automation, and resilience.Implement enterprise monitoring, observability, and event...
- ...in-depth understanding and knowledge of Power’s data center infrastructure & network standards and coding practices, and helps shape SRE standards & operating procedures, and coding practices. Provides technical leadership to assigned primary SRE team, as well as other...For contractorsWork at officeRemote workHome office
- ...discussions.What you will bring to the role:Required Skills & ExperienceTechnical SkillsProven experience in cloud/platform engineering, SRE, or database/data infrastructure engineering roles.Strong understanding of database platform fundamentals (availability, consistency,...
- ...Development experience with any of the following programming languages: Java or Python• Recent on-program Site Reliability Engineering (SRE) experience (e.g., automation, incident management, monitoring, optimization, programming/scripting, metrics, security, etc.)•...Full timeWork at officeRemote workFlexible hours
- ...basic observability (logging, monitoring, alerting) for owned services. Supports production troubleshooting in collaboration with DevOps/SRE teams. Designs and delivers scalable fintech product services and distributed systems with a focus on performance, resilience, and...Temporary work
$113.89k - $187.1k
...requirements into new OpenStack product releases. Continue to evolve and enhance our existing deployment automation capabilities. Support SRE team for L3 escalations Required Qualifications: Advanced Linux knowledge and hands on experience Expertise in...Work experience placementWorldwide- ...Clearance ~ BS in Software Engineering or related field ~2-10 years of experience in CI/CD and Kubernetes-focused engineering or SRE/Platform Engineering ~ Strong Kubernetes skills - core objects, scheduling, RBAC, networking, storage, upgrades, troubleshooting...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!


