SRE
RIT Solutions, Inc.
SRE
Hybrid - Malvern, PA
ASSESSMENT REQUIRED
As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance innovative thinking with pragmatism and understand the long-term impacts of technical decisions. You communicate complex ideas clearly and collaborate effectively to deliver scalable solutions.
Core Responsibilities
Team is focused on automating incident response and infrastructure management. While Java and Python receive a stronger emphasis, candidates with solid programming fundamentals in any language and the ability to adapt will be considered. Experience with AWS and event-driven architectures is also valuable.
From a technical standpoint, familiarity with observability concepts (e.g., distributed tracing) and tools like Prometheus or Grafana is beneficial, though not mandatory. More important is an understanding of the underlying principles, such as instrumentation and monitoring strategies.
* Improve resiliency engineering practices across platforms and applications, including resilient application design patterns, system observability and deployment strategies
* Incident detection, troubleshooting, and resolution.
* Develop automation for incident response and infrastructure management
* Develop and support OpenTelemetry integrations for multiple application platforms (browser, ECS, lambda, etc) and languages (JavaScript, Java)
* Contribute to architectural decisions and support implementation of solutions.
Skills and Qualifications
* Deep knowledge of Java or Javascript. Practical experience developing and operating software in distributed systems environments.
* Problem-solving and analytical thinking: ability to diagnose complex issues and propose efficient solutions. Strong debugging and optimization skills for performance and scalability.
* Cloud platforms: Hands-on experience with AWS services and cloud infrastructure
* System architecture and design: ability to design scalable, secure, and maintainable systems.
* Working knowledge of Python (or similar scripting language).
* Strong knowledge of resiliency engineering techniques for both platforms and applications.
* Experience troubleshooting complex production issues and implementing effective mitigations.
* Familiarity with OpenTelemetry specification and core APIs.
From a screening perspective, we recommend focusing on:
• How candidates approach software releases and validate functionality
• Their understanding of system dependencies and fault tolerance
• Experience with diagnosing and resolving production issues
• Their ability to reflect on past incidents and identify improvements
• Evidence of systems thinking and architectural awareness
- ...We're Hiring: SRE Production Support Engineer Malvern, PA · Onsite Contract Experience: 5+ yrs Skills: Shell, Bash, AWS, Azure, Docker, Kubernetes, Terraform, Ansible, Jenkins, CI/CD, Nginx, Grafana Roles & Responsibilities: Provide L2/L3 production support...SuggestedPermanent employmentContract work
- ...Nice to Have Skills Building JQL-based APIs. Front-end development (React/TypeScript). Site Reliability Engineering (SRE) practices. qualifications: Required Technical Skills Core Languages: Robust, demonstrable experience in both Java and Python...SuggestedHourly payTemporary workWork experience placementNight shift
$66.48 per hour
...and ensuring alignment with product priorities. The Scrum Leader will partner closely with product owners, engineering leads, and SRE teams to improve delivery predictability, remove impediments, and strengthen quality, operational readiness, and continuous improvement...SuggestedHourly payContract workTemporary workWork experience placement- ...AWS microservices, serverless, infrastructure as code / cloud formation, security architecture, NoSQL and relational databases, SRE and DevOps best practices for continuous delivery AWS Services - Eventbridge - Lambda - Step Function - SQS,...SuggestedFull time
- ...enterprise infrastructure (CI/CD, CMDB, ITSM, identity providers). Partners with engineering leadership across Platform, DevOps, SRE, and application security teams to align on shared interfaces, data contracts, and remediation workflows that reduce friction at organizational...SuggestedWork experience placementImmediate startShift work
- ...with exceptional professionals for this role. Join Personal Investor Tech's Site Reliability Engineering team and lead cutting-edge SRE initiatives that impact hundreds of applications and millions of investors. You'll architect and build enterprise-scale resiliency...
- ...operability is designed in, not bolted on—shifting left on reliability, monitoring, and remediation. Champion Site Reliability Engineering (SRE) and ITSM best practices within IAM operations. Global Delivery & Service Ownership Oversee incident, problem, change, and...Local areaShift work
- ...containerized identity services Automate provisioning, deployment, monitoring, and drift detection for identity platforms. Support SRE‑style operational maturity: SLIs/SLOs, alerting, incident response, and runbooks for identity services. Security, Risk &...Visa sponsorship
- ...implementations with a keen eye toward the future state of technology and the industry. This team member works closely with the Security team, the SRE team, and Development team to build the frameworks that will take our technology into the future. This team member is future-focused,...
- ...Experience developing and integrating APIs, including working with a Supergraph, and JQL API protocol–based ecosystem. Nice to Have Skills Building JQL-based APIs. Front-end development (React/TypeScript). Site Reliability Engineering (SRE) practices....Night shift
- ...proactive engineer who excels at the intersection of infrastructure, automation, and system reliability, blending responsibilities across SRE, DevOps, and Cloud Engineering. At CubeSmart, we're intentional about culture. You can experience it everywhere from our mission...Work at officeRemote work
- ...operational problems. You are curious and take a proactive approach to identifying problems and Approx desired start: ASAP Role: Expert SRE UI Engineer Expert-level proficiency in JavaScript, spanning both client-side and server-side execution environments Robust...Immediate start
- ...Responsibilities Operate in a highly collaborative team emphasizing best practices in software development, automation, DevSecOps and SRE. Build cloud-first, consumer-focused and applying lean agile methodologies. Ability to learn and build in third party...
$120k - $150k
..., API‑first and cloud architectures. We continue to enhance reliability and accelerate engineering productivity by strengthening our SRE and AI practices. This is a large investment in innovation to continue to drive operational excellence at our facilities. If you want...Permanent employmentTemporary workWork at officeFlexible hours$160k - $190k
..., API-first and cloud architectures. We continue to enhance reliability and accelerate engineering productivity by strengthening our SRE and AI practices. This is a large investment in innovation to continue to drive operational excellence at our facilities. If you want...Permanent employmentTemporary workWork at officeImmediate startFlexible hours- ...Clearance ~ BS in Software Engineering or related field ~2-10 years of experience in CI/CD and Kubernetes-focused engineering or SRE/Platform Engineering ~ Strong Kubernetes skills - core objects, scheduling, RBAC, networking, storage, upgrades, troubleshooting...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!


