SRE
RIT Solutions
SRE Hybrid - Malvern, PA
As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance innovative thinking with pragmatism and understand the long-term impacts of technical decisions. You communicate complex ideas clearly and collaborate effectively to deliver scalable solutions.
Core Responsibilities
Team is focused on automating incident response and infrastructure management. While Java and Python receive a stronger emphasis, candidates with solid programming fundamentals in any language and the ability to adapt will be considered. Experience with AWS and event-driven architectures is also valuable. From a technical standpoint, familiarity with observability concepts (e.g., distributed tracing) and tools like Prometheus or Grafana is beneficial, though not mandatory. More important is an understanding of the underlying principles, such as instrumentation and monitoring strategies.
- Improve resiliency engineering practices across platforms and applications, including resilient application design patterns, system observability and deployment strategies
- Incident detection, troubleshooting, and resolution.
- Develop automation for incident response and infrastructure management
- Develop and support OpenTelemetry integrations for multiple application platforms (browser, ECS, lambda, etc) and languages (JavaScript, Java)
- Contribute to architectural decisions and support implementation of solutions.
Skills and Qualifications
* Deep knowledge of Java or Javascript. Practical experience developing and operating software in distributed systems environments. * Problem-solving and analytical thinking: ability to diagnose complex issues and propose efficient solutions. Strong debugging and optimization skills for performance and scalability. * Cloud platforms: Hands-on experience with AWS services and cloud infrastructure * System architecture and design: ability to design scalable, secure, and maintainable systems. * Working knowledge of Python (or similar scripting language). * Strong knowledge of resiliency engineering techniques for both platforms and applications. * Experience troubleshooting complex production issues and implementing effective mitigations. * Familiarity with OpenTelemetry specification and core APIs. From a screening perspective, we recommend focusing on: · How candidates approach software releases and validate functionality · Their understanding of system dependencies and fault tolerance · Experience with diagnosing and resolving production issues · Their ability to reflect on past incidents and identify improvements · Evidence of systems thinking and architectural awareness
- ...We're Hiring: SRE Production Support Engineer Malvern, PA · Onsite Contract Experience: 5+ yrs Skills: Shell, Bash, AWS, Azure, Docker, Kubernetes, Terraform, Ansible, Jenkins, CI/CD, Nginx, Grafana Roles & Responsibilities: Provide L2/L3 production support...SuggestedPermanent employmentContract work
- ...SRE Role SRE role is a combination between architect-digital, full stack developer-digital, cloud engineering and system engineer. 9 to 12 months duration in Malvern PA Responsibilities Specific skillsets in addition to expert level full stack developer profile...SuggestedWork experience placement
- ...discussions, and operational initiatives. Contributes to improving engineering practices and organizational effectiveness across the SRE group. Participates in special projects and performs other duties as assigned. Qualifications Minimum 10 years of experience...Suggested
- ...containerized identity services Automate provisioning, deployment, monitoring, and drift detection for identity platforms. Support SRE‑style operational maturity: SLIs/SLOs, alerting, incident response, and runbooks for identity services. Security, Risk &...SuggestedVisa sponsorship
- Job Title Required Skills: Has exceptional coding skills— Java, python Has AWS knowledge and thinks out of the box SRE knowledge would be a nice to have Somebody who has a drive and willing to learn SRE and can do automationSuggested
- ...Site Reliability Engineer Hybrid - Malvern, PA needs at least 8 years experience within the US The Site Reliability Engineer (SRE) is responsible for improving the reliability, resiliency, observability, and operational excellence of Client's Cash & Money Movement...Work at office
- ...Responsibilities Operate in a highly collaborative team emphasizing best practices in software development, automation, DevSecOps and SRE. Build cloud-first, consumer-focused and applying lean agile methodologies. Ability to learn and build in third party...
- ...operability is designed in, not bolted on—shifting left on reliability, monitoring, and remediation. Champion Site Reliability Engineering (SRE) and ITSM best practices within IAM operations. Global Delivery & Service Ownership Oversee incident, problem, change, and service...Local areaShift work
$120k - $150k
..., API‑first and cloud architectures. We continue to enhance reliability and accelerate engineering productivity by strengthening our SRE and AI practices. This is a large investment in innovation to continue to drive operational excellence at our facilities. If you want...Permanent employmentTemporary workWork at officeFlexible hours$160k - $190k
..., API-first and cloud architectures. We continue to enhance reliability and accelerate engineering productivity by strengthening our SRE and AI practices. This is a large investment in innovation to continue to drive operational excellence at our facilities. If you want...Permanent employmentTemporary workWork at officeImmediate startFlexible hours- ...Clearance ~ BS in Software Engineering or related field ~2-10 years of experience in CI/CD and Kubernetes-focused engineering or SRE/Platform Engineering ~ Strong Kubernetes skills - core objects, scheduling, RBAC, networking, storage, upgrades, troubleshooting...Flexible hours
- ...infrastructure; distinguish product defects from test/environment issues and drive root-cause analysis. Partner with Software Engineering, SRE/DevOps, Security, and Product to define software quality reports, quality gates, release criteria, and automation standards....Contract workTemporary workH1bWork at officeMonday to Friday
- ...upgrade programs aligned with vendor roadmaps.Build and lead a Platform Reliability Engineering (PRE) or Site Reliability Engineering (SRE) function focused on proactive monitoring, automation, and resilience.Implement enterprise monitoring, observability, and event...Part time
$274k
...error budgets, on-call processes, incident response, and blameless post-mortems, in coordination with the CoE Platform Reliability / SRE function.Drive platform cost transparency and FinOps practices ensuring compute and storage costs are visible, attributed, and continuously...Contract workPart timeH1bLocal areaVisa sponsorshipWork visaRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!




