SRE
RIT Solutions
SRE Hybrid - Malvern, PA
As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance innovative thinking with pragmatism and understand the long-term impacts of technical decisions. You communicate complex ideas clearly and collaborate effectively to deliver scalable solutions.
Core Responsibilities
Team is focused on automating incident response and infrastructure management. While Java and Python receive a stronger emphasis, candidates with solid programming fundamentals in any language and the ability to adapt will be considered. Experience with AWS and event-driven architectures is also valuable. From a technical standpoint, familiarity with observability concepts (e.g., distributed tracing) and tools like Prometheus or Grafana is beneficial, though not mandatory. More important is an understanding of the underlying principles, such as instrumentation and monitoring strategies.
- Improve resiliency engineering practices across platforms and applications, including resilient application design patterns, system observability and deployment strategies
- Incident detection, troubleshooting, and resolution.
- Develop automation for incident response and infrastructure management
- Develop and support OpenTelemetry integrations for multiple application platforms (browser, ECS, lambda, etc) and languages (JavaScript, Java)
- Contribute to architectural decisions and support implementation of solutions.
Skills and Qualifications
* Deep knowledge of Java or Javascript. Practical experience developing and operating software in distributed systems environments. * Problem-solving and analytical thinking: ability to diagnose complex issues and propose efficient solutions. Strong debugging and optimization skills for performance and scalability. * Cloud platforms: Hands-on experience with AWS services and cloud infrastructure * System architecture and design: ability to design scalable, secure, and maintainable systems. * Working knowledge of Python (or similar scripting language). * Strong knowledge of resiliency engineering techniques for both platforms and applications. * Experience troubleshooting complex production issues and implementing effective mitigations. * Familiarity with OpenTelemetry specification and core APIs. From a screening perspective, we recommend focusing on: · How candidates approach software releases and validate functionality · Their understanding of system dependencies and fault tolerance · Experience with diagnosing and resolving production issues · Their ability to reflect on past incidents and identify improvements · Evidence of systems thinking and architectural awareness
- ...We're Hiring: SRE Production Support Engineer Malvern, PA · Onsite Contract Experience: 5+ yrs Skills: Shell, Bash, AWS, Azure, Docker, Kubernetes, Terraform, Ansible, Jenkins, CI/CD, Nginx, Grafana Roles & Responsibilities: Provide L2/L3 production support...SuggestedPermanent employmentContract work
- Job ID: 20139475Reference Number: 23-01136Title: Devops/SRELocation: Malvern, PA, 07512Posted Date: 2023-07-12Company: HAN Staffing Migrated key systems from on-prem hosting to AWS Worked in Agile and Scrum methodologies/practices Worked on designing and developing a multitude...Suggested
- ...SRE Hybrid - Malvern, PA ASSESSMENT REQUIRED As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance...Suggested
- ...SRE Role SRE role is a combination between architect-digital, full stack developer-digital, cloud engineering and system engineer. 9 to 12 months duration in Malvern PA Responsibilities Specific skillsets in addition to expert level full stack developer profile...SuggestedWork experience placement
- ...focused practices, including monitoring, observability, automation, and root cause analysis Oversee Site Reliability Engineering (SRE) functions to improve system stability and performance Reduce recurring incidents and operational risk through proactive improvements...SuggestedRemote workMonday to Friday
- ...connect to enterprise infrastructure (CI/CD, CMDB, ITSM, identity providers).Partners with engineering leadership across Platform, DevOps, SRE, and application security teams to align on shared interfaces, data contracts, and remediation workflows that reduce friction at...Full timeWork experience placementImmediate startShift work
- Job Title Required Skills: Has exceptional coding skills— Java, python Has AWS knowledge and thinks out of the box SRE knowledge would be a nice to have Somebody who has a drive and willing to learn SRE and can do automation
- ...Site Reliability Engineer Hybrid - Malvern, PA needs at least 8 years experience within the US The Site Reliability Engineer (SRE) is responsible for improving the reliability, resiliency, observability, and operational excellence of Client's Cash & Money Movement...Work at office
- ...PostResponsibilitiesOperate in a highly collaborative team emphasizing best practices in software development, automation, DevSecOps and SRE.Build cloud-first, consumer-focused and applying lean agile methodologies.Ability to learn and build in third party application that...
- ...Kubernetes & containerized identity servicesAutomate provisioning, deployment, monitoring, and drift detection for identity platforms.Support SRE‑style operational maturity: SLIs/SLOs, alerting, incident response, and runbooks for identity services.Security, Risk &...Full timeVisa sponsorship
- ...observability, resilience, and operational excellence. This role is well suited for a senior engineer who thrives at the intersection of SRE, cloud architecture, platform engineering, DevOps, and systems engineering, and who can independently drive complex technical...Work at officeRemote work
- As an Automation Engineer II, you will be a member of the Site Reliability Engineering (SRE) team and play a key role in designing, developing, and maintaining automation solutions that support enterprise infrastructure and operational capabilities. Working closely with...Work experience placement
- ...Clearance ~ BS in Software Engineering or related field ~2-10 years of experience in CI/CD and Kubernetes-focused engineering or SRE/Platform Engineering ~ Strong Kubernetes skills: core objects, scheduling, RBAC, networking, storage, upgrades, troubleshooting ~...Flexible hours
- ...upgrade programs aligned with vendor roadmaps.Build and lead a Platform Reliability Engineering (PRE) or Site Reliability Engineering (SRE) function focused on proactive monitoring, automation, and resilience.Implement enterprise monitoring, observability, and event...Work at officeRelocationRelocation package
- ...frameworks, including GDPR, CCPA, HIPAA, or similar standards. Experience working in Agile, DevOps, or Site Reliability Engineering (SRE) environments.Familiarity with scripting or automation technologies such as PowerShell, Python, or similar languages. Success Factors...Full timeLocal area
- ...Architectures Cloud Engineering & Infrastructure as Code (IaC) DevOps, CI/CD & Engineering Productivity Site Reliability Engineering (SRE) & Operational Excellence Monitoring, Observability & Platform Reliability Security, Governance & Engineering Controls The...Work experience placement
$68 - $85 per hour
...nexs-riven qestions re rotine prt of the work, not n nnl exception. Yo will work cross or service, project, n procrement tems to mke sre trnsctions re recore ccrtely n on time, n tht billing reches the cstomer clen the first time. Key Responsibilities Csh n...$135.7k - $251.9k
...Development experience with any of the following programming languages: Java or Python• Recent on-program Site Reliability Engineering (SRE) experience (e.g., automation, incident management, monitoring, optimization, programming/scripting, metrics, security, etc.)•...Full timeTemporary workWork experience placementCasual workFlexible hours- ...basic observability (logging, monitoring, alerting) for owned services. Supports production troubleshooting in collaboration with DevOps/SRE teams. Designs and delivers scalable fintech product services and distributed systems with a focus on performance, resilience, and...Temporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!



