Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Cubesmart

Job Description

Job Description

Overview

This is a hybrid role - 2 days remote and 3 days in the Malvern, PA office.

 

We are seeking a highly experienced Senior Site Reliability & Cloud Systems Engineer to architect, build, automate, and operate scalable, secure, resilient, and highly available cloud platforms in AWS. This role combines hands-on reliability engineering with cloud architecture and automation expertise, with a strong emphasis on building immutable infrastructure and improving system resilience.

 

You will play a critical role in evolving our AWS ecosystem into a fully automated, self-service, “push-button” platform, minimizing manual operational intervention while improving reliability, security, performance, scalability, and engineering velocity. You will establish and champion engineering standards for immutable infrastructure, automation, observability, resilience, and operational excellence.

 

This role is well suited for a senior engineer who thrives at the intersection of SRE, cloud architecture, platform engineering, DevOps, and systems engineering, and who can independently drive complex technical initiatives from architecture and design through implementation and production operation.

 

Who we are:

At CubeSmart, we’re intentional about culture. You can experience it everywhere from our mission statement of “genuine care” to our “It’s What’s Inside That Counts” tagline to calling each other “teammates” rather than employees. This spirit fosters a fun and collaborative environment that has resulted in our rapid growth and being recognized amongst the top in our industry. 

 

CubeSmart’s award-winning team is made up of people who genuinely care. Teammates care about our customers and the life events and/or business needs they are facing. Teammates are passionate, responsible and understanding. The CubeSmart team is made up of people who have a can-do attitude, are committed to their own success and the success of the company, and lead by example. 

If this sounds like a team and culture that matches your personal values and motivations, we want to hear from you.

Responsibilities

Reliability, Performance & Production Operations

  • Own the reliability, availability, scalability, performance, and operational health of AWS-hosted, Linux-based production platforms and associated lower environments.
  • Define and drive SRE practices, reliability standards, SLOs, SLIs, error budgets, and operational maturity across critical systems.
  • Architect and implement comprehensive observability using platforms such as Datadog, Amazon CloudWatch, Prometheus/Grafana, and PagerDuty.
  • Establish proactive monitoring, alerting, capacity planning, performance engineering, and predictive reliability practices.
  • Lead complex production incident response, including technical triage, mitigation, service restoration, root cause analysis, and executive/stakeholder communication.
  • Drive blameless post-incident reviews and ensure corrective actions are translated into measurable reliability improvements.
  • Participate in and provide leadership during 24/7 on-call operations, including escalation management for high-severity production incidents.
  • Identify systemic reliability risks and proactively eliminate single points of failure, operational bottlenecks, and sources of technical debt.
  • Develop and implement disaster recovery, business continuity, backup, restoration, and resilience strategies.
  • Perform capacity and performance analysis for distributed applications and infrastructure at scale.

Cloud Architecture & Automation

  • Architect and implement highly automated, ephemeral, immutable, and reproducible AWS environments across production and non-production workloads.
  • Lead the design of scalable, fault-tolerant, secure distributed systems using AWS Well-Architected principles.
  • Establish infrastructure patterns that enable engineering teams to provision and manage environments through self-service and “push-button” automation.
  • Eliminate manual infrastructure operations through Infrastructure as Code using Terraform, Ansible, Packer, and related technologies.
  • Design and maintain reusable infrastructure modules, automation frameworks, and engineering standards.
  • Build and evolve CI/CD and GitOps workflows using technologies such as Jenkins, GitHub Actions, GitLab CI, ArgoCD, and Flux.
  • Develop sophisticated automation and operational tooling using Python and Bash.
  • Identify opportunities to reduce operational toil through automation, platform engineering, and intelligent operational workflows.
  • Establish engineering patterns for immutable infrastructure, automated provisioning, blue/green deployments, canary releases, and automated rollback.

Infrastructure & Platform Engineering

  • Architect, deploy, and operate AWS services including:
  • EKS, ECS, Fargate, Lambda
  • RDS and Aurora PostgreSQL
  • OpenSearch
  • Redis and ElastiCache
  • Load balancing, networking, storage, compute, and supporting AWS services
  • Design and manage enterprise AWS networking architectures, including Transit Gateways, VPCs, routing, security groups, network segmentation, load balancers, and service meshes.
  • Architect Kubernetes and container platforms for highly available production workloads.
  • Establish standardized platform capabilities that enable development teams to deploy applications safely and independently.
  • Design and implement caching, asynchronous processing, microservices, service discovery, and other distributed-system patterns.
  • Evaluate emerging AWS and cloud-native technologies and make recommendations based on reliability, security, scalability, operational complexity, and cost.
  • Provide technical leadership for major infrastructure modernization and cloud transformation initiatives.

Security, Risk & Governance

  • Design and implement zero-trust cloud security architectures across AWS environments.
  • Establish secure identity and access patterns using IAM, Organizations, SCPs, OIDC, KMS, Secrets Manager, and related AWS security services.
  • Implement least-privilege access models and automated security controls across infrastructure and deployment pipelines.
  • Embed security into CI/CD and Infrastructure as Code workflows using SAST, DAST, SCA, and vulnerability-management tools, including technologies such as Snyk.
  • Implement automated compliance auditing, configuration validation, security scanning, backup verification, and governance controls.
  • Partner with security and compliance teams to establish cloud security standards and ensure infrastructure meets organizational and regulatory requirements.
  • Secure secrets, credentials, certificates, and sensitive configuration using technologies such as HashiCorp Vault and Ansible Vault.
  • Continuously assess infrastructure for security vulnerabilities and proactively remediate risks.

Platform Engineering, Leadership & Strategy

  • Serve as a senior technical authority for cloud infrastructure, reliability engineering, automation, and platform architecture.
  • Mentor engineers and establish best practices for cloud engineering, SRE, DevOps, Infrastructure as Code, observability, and operational excellence.
  • Partner closely with development, security, architecture, and operations teams to design reliable and scalable application platforms.
  • Influence architectural decisions and provide technical guidance on complex infrastructure and distributed-system challenges.
  • Establish and maintain engineering standards, reference architectures, reusable patterns, and operational guidelines.
  • Create and maintain comprehensive architecture documentation, runbooks, troubleshooting guides, disaster recovery procedures, and operational standards.
  • Drive FinOps initiatives, including cloud cost optimization, resource utilization, capacity forecasting, rightsizing, and infrastructure efficiency.
  • Integrate infrastructure and operational changes into ITIL-compliant change-management and service-management processes, including platforms such as Freshservice.
  • Identify and prioritize opportunities to improve reliability, automation, security, developer experience, and operational efficiency.
  • Lead cross-functional technical initiatives and drive complex projects from requirements and architecture through implementation and production adoption.

Qualifications

  • 10+ years of professional experience in Site Reliability Engineering, DevOps, Cloud Engineering, Infrastructure Engineering, Platform Engineering, or a closely related discipline.
  • Demonstrated experience designing, implementing, and operating enterprise-scale AWS cloud environments.
  • Deep hands-on expertise with AWS architecture, services, networking, security, and operational best practices.
  • Extensive Linux systems engineering experience, preferably with Ubuntu.
  • Advanced experience with Infrastructure as Code, particularly Terraform and Ansible.
  • Proven ability to design and implement highly automated, immutable, and reproducible infrastructure.
  • Extensive experience designing and maintaining enterprise CI/CD pipelines and deployment automation.
  • Strong programming and automation capabilities using Python and Bash.
  • Extensive experience with production observability, monitoring, logging, alerting, and incident-management platforms.
  • Strong understanding of distributed systems, scalability, fault tolerance, high availability, and performance engineering.
  • Demonstrated experience leading production incident response and conducting root cause analysis.
  • Strong understanding of cloud security principles, including IAM, KMS, secrets management, least privilege, OIDC, and zero-trust architectures.
  • Experience implementing security controls and automated governance within cloud infrastructure and CI/CD pipelines.
  • Demonstrated ability to independently lead complex technical initiatives and make sound architectural decisions.
  • Strong written and verbal communication skills, including the ability to communicate technical risks and decisions to both engineering and non-engineering stakeholders.
  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
  • Candidates must be authorized to work in the U.S. without the need for current or future sponsorship.

Preferred Qualifications

  • Extensive experience with Docker, Kubernetes, Amazon EKS, ECS, and Fargate.
  • Experience designing and operating Kubernetes platforms at enterprise scale.
  • Advanced experience with GitOps, including ArgoCD or Flux.
  • Experience implementing SAST, DAST, SCA, and secure SDLC practices.
  • Strong knowledge of distributed systems, microservices, caching, messaging, and asynchronous architectures.
  • Experience designing highly available and geographically resilient cloud architectures.
  • Experience with service meshes and cloud-native networking.
  • Experience implementing automated disaster recovery and multi-region resilience strategies.
  • Experience with FinOps, cloud cost optimization, and capacity forecasting.
  • Experience with ITIL processes, change management, and enterprise service-management platforms.
  • Experience mentoring engineers and establishing technical standards across engineering teams.
  • AWS professional-level certification or equivalent demonstrated expertise.
  • Experience leading large-scale cloud migrations, infrastructure modernization, or platform transformation initiatives.

#LI-MT1

 

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Malvern, PA vacancy
  •  ...Senior Site Reliability Engineer Bentley Systems Location Exton or Philadelphia, PA (Hybrid - 3 times a week in-office) Position Summary Are you ready to start a new journey with a team of energized professionals advancing and connecting the world's infrastructure... 
    Senior
    Casual work
    Work at office
    Worldwide

    Bentley Systems

    Exton, PA
    4 days ago
  •  ...Site Reliability Engineer Hybrid - Malvern, PA needs at least 8 years experience within the US The Site Reliability Engineer (SRE) is responsible for improving the reliability, resiliency, observability, and operational excellence of Client's Cash & Money Movement... 
    Suggested
    Work at office

    RIT Solutions

    Malvern, PA
    12 hours ago
  •  ...Job Title: Site Reliability Engineer Locaton: Malvern, PA Duration: Contract Job Description: Ansible Production Elevations support [ this includes right from reviewing bitbucket configs pull request changes, patching Ansible prod job templates & monitoring... 
    Suggested
    Contract work

    Syntricate Technologies

    Malvern, PA
    1 day ago
  •  ...Role Overview We are looking for a Senior Software Engineer for a contract engagement to build resilient backend services and full-stack solutions for strategic initiatives. You will tackle technical discovery, design API contracts, and deliver production-ready code in... 
    Senior
    Contract work

    Drevol

    Malvern, PA
    1 day ago
  •  ...Senior Reliability Engineer hybrid - malvern, pa Job Description As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making... 
    Senior

    RIT Solutions, Inc.

    Malvern, PA
    4 days ago
  • Job ID: 19678221Reference Number: 23-00871Title: Senior DevSecOps EngineerLocation: Malvern, PAPosted Date: 2023-05-12Company: HAN StaffingSenior...  ...infrastructure for deployed servicesWork with peer software engineers to design resilient, scalable, and robust cloud and IoT... 
    Senior

    HAN Staffing

    Malvern, PA
    2 days ago
  •  ...scalable, resilient systems that support the full transaction lifecycle, combining distributed systems engineering, financial domain expertise, and platform reliability. What makes this team exciting is the blend of deep technical ownership and collaboration. Engineers... 
    Senior
    Temporary work

    Envestnet

    Devon, PA
    2 days ago
  •  ...remediation.Develop strategies to secure current and emerging technologies (containers, serverless, API, AI/ML).Provide hands-on engineering support for security incidents, threat events, vulnerability management, and control improvements to strengthen Enterprise application... 
    Senior
    Full time

    Vanguard

    Malvern, PA
    20 hours ago
  •  ...application features. Partner with product managers, designers, engineers, and platform teams to deliver intuitive, high-quality mobile...  ...Continuously improve application performance, responsiveness, reliability, and maintainability. Translate stories into design & code.... 
    Senior
    Full time
    Work experience placement

    Vanguard

    Malvern, PA
    3 days ago
  • Job Title: Senior Scientist / Principal Scientist - Biotransformation/Metabolite Identification.Location: Exton, PennsylvaniaFull-time Job Description: We are seeking a highly motivated and experienced Senior Scientist specializing in Biotransformation and Metabolite Identification... 
    Senior

    Frontage Laboratories

    Lionville, PA
    1 day ago
  •  ...Senior Software Engineer Bentley Systems Location Hybrid – Exton, PA / Philadelphia / Alabama, Huntsville Visa sponsorship is not...  ...modern continuous integration and delivery practices to ship reliable cloud platform capabilities quickly and safely. Translate... 
    Senior
    Worldwide

    Bentley Systems

    Exton, PA
    20 hours ago
  •  ...experience with a scripting language such as PowerShell and/or Python. Also, this is not a developer role but rather an infrastructure role in AWS and on-prem Windows. Required Skills : Cloud Additional Skills : Network Engineer This is a high PRIORITY requisition.... 
    Senior

    Samprasoft

    Malvern, PA
    1 day ago
  • The AI Threat Detection Engineer, Senior Specialist is responsible for developing and implementing AI-driven capabilities that enhance Security Operations Center (SOC) effectiveness. This role focuses on building automation and intelligent solutions to improve threat detection... 
    Senior
    Full time

    Vanguard

    Malvern, PA
    1 day ago
  •  ...automation platforms, and integration frameworks — ensuring they scale reliably across thousands of cloud accounts and multiple business units...  ...-driven pipelines, API contracts, data models) that other engineers build on — establishing the foundational approach for how... 
    Senior
    Full time
    Work experience placement
    Immediate start
    Shift work

    Vanguard

    Malvern, PA
    20 hours ago
  •  ...building an AI Code Modernization team to transform large-scale engineering software into modern cloud-native architectures. You will...  ...Bentley’s products. The role emphasizes production impact, reliability, and scalability, with no travel required. Hybrid work is available... 
    Senior

    Jobleads-US

    Exton, PA
    5 days ago
  • Job Title Responsibilities Operate in a highly collaborative team emphasizing best practices in software development, automation, DevSecOps and SRE. Build cloud-first, consumer-focused and applying lean agile methodologies. Ability to learn and build in ...
    Senior

    Samprasoft

    Malvern, PA
    1 day ago
  • The Senior Manager, AI & Data Engineering leads teams responsible for building and operating reusable, governed data products and AI-ready semantic...  ...outcomes while ensuring data quality, governance, security, reliability, and operational excellence. Success in this role... 
    Senior
    Full time
    For contractors
    Work at office

    Vanguard

    Malvern, PA
    2 days ago
  • Job ID: 19130188Reference Number: 23-00482Title: Senior UI DeveloperLocation: malvern, PA, 07512Posted Date: 2023-03-03Company: HAN Staffing Provides expert level system analysis, design, development, and implementation of external facing web applications. Integrates third... 
    Senior
    Work experience placement

    HAN Staffing

    Malvern, PA
    2 days ago
  • Job Title Focus will be on implementing IAM processes related to our Cloud access to infrastructure, applications and databases, learning the existing process, understanding automation requirements, development, testing, and implementation. Serve as subject matter...
    Senior

    Samprasoft

    Malvern, PA
    1 day ago
  • We are looking for a Sr. Software Engineer to help shape and advance AI-driven software modernization efforts in Exton, Pennsylvania. This role will guide engineering teams on applied AI practices, improve code conversion quality across varied product environments, and... 
    Senior

    Robert Half

    Exton, PA
    4 days ago
  •  ...Senior Technical Engineer Assessment required hybrid malvern, pa Job Description Senior Technical Engineer - Security Automation & Vulnerability Solutions Role Purpose: Lead the design and development of centralized solutions that improve vulnerability... 
    Senior

    RIT Solutions, Inc.

    Malvern, PA
    4 days ago
  •  ...lifecycle. Collaborate with product owners, architects, and legacy system SMEs to translate business and technical requirements into reliable solutions. Develop and execute unit tests, integration tests, and validation scripts to support end-to-end account sync... 
    Senior

    RIT Solutions

    Malvern, PA
    20 hours ago
  • Senior AEM Cloud Developer / ArchitectExperience: 8+ Years AEM as a Cloud Service: 3+ Years...  ...Build enterprise DAM solutions and multi-site architectures.Configure and optimize AEM...  ...understanding of security, scalability, reliability, and production support. Purple Drive
    Senior

    Purple Drive

    Malvern, PA
    5 days ago
  •  ...We believe that if it can be dreamed it can also be measured. And if it can be measured, it can also be realized.The Sr Mechanical Engineer is a key contributor responsible for design and analysis of mechanical systems, equipment and packaging for FARO's metrology... 
    Senior
    Full time
    For contractors
    Work at office
    Immediate start

    Faro

    Lionville, PA
    1 day ago
  •  ...Our solutions improve the techniques and outcomes of surgery so patients can resume their lives as quickly as possible. The Senior Software Engineer will be a part of our rapidly growing surgical navigation division. Here, we develop novel tracking platforms, sensors and... 
    Senior
    Full time
    Work at office

    Globus Medical

    Audubon, PA
    1 day ago
  •  ...solutions improve the techniques and outcomes of surgery so patients can resume their lives as quickly as possible. As a Senior Software Engineer, you will be part of a small pluri-disciplinary development team of 5 to 10 engineers, dedicated to the development of a complete... 
    Senior
    Full time

    Globus Medical

    Audubon, PA
    1 day ago
  •  ...innovative company and help us improve the lives of people and animals everywhere. Apply today!Job DetailsJob Description SummaryDesigns, engineers, and continuously enhances enterprise Identity and Access Management (IAM) platforms with a strong emphasis on automation,... 
    Senior
    Full time
    Local area

    Cencora

    Conshohocken, PA
    4 days ago
  • Shape the future of enterprise identity by engineering, securing, and automating mission-...  ...secure access across the enterprise. As a senior member of the Identity Platform team,...  ...capabilities that improve security, resiliency, reliability, and user experience.This role combines... 
    Senior
    Permanent employment
    Full time

    Vanguard

    Malvern, PA
    1 day ago
  • $110.3k - $204.9k

     ...the world by using our outstanding skills and experiences to create, craft and build solutions to some of the world’s most complex engineering problems. Our culture supports employees to dream big, perform with excellence and create incredible products.Our MissionLockheed... 
    Senior
    Full time
    Temporary work
    Work experience placement
    Casual work
    Flexible hours

    Lockheed Martin

    Valley Forge, Montgomery County, PA
    4 days ago
  • Senior Backend Developer hybird - malvern, pa assessment required Team name: Debit cards compliance and BIN Role: Experienced Java/SpringBoot API Sr Developer Required skillset: Java, SpringBoot, AWS ECS, Databases, Python
    Senior

    Akaasa Technologies

    Malvern, PA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!