Senior Site Reliability Engineer
Castleton Commodities International
Senior Site Reliability Engineer
The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams to design resilient cloud-native architectures, implement Infrastructure as Code (IaC) and CI/CD standards, and drive measurable reliability outcomes. The Senior Site Reliability Engineer will also lead efforts to define and validate recovery objectives (RTO/RPO), design and implement Business Continuity / Disaster Recovery (BCP/DR) plans, and coordinate structured testing to ensure readiness.
Responsibilities:
- Reliability Engineering & Operations: Own and improve service reliability through SLO/SLI definition, error budgets, and operational best practices. Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR and improve proactive detection. Lead incident response practices including on-call improvements, runbooks, post-incident reviews (RCA), and preventative actions. Partner with application teams to improve performance, capacity planning, and resiliency under failure scenarios.
- Infrastructure & Cloud Architecture: Design and operate highly available, fault-tolerant Cloud architectures (multi-AZ and, where required, multi-region). Implement resilient patterns across compute, storage, networking, and managed services (e.g., autoscaling, load balancing, backups, replication). Drive cloud governance best practices (tagging, account/landing zone patterns, least privilege, guardrails) in partnership with security and platform teams.
- Infrastructure as Code (IaC) & DevOps Enablement: Build and maintain IaC modules and standards (e.g., Terraform, CloudFormation, CDK) for repeatable, auditable infrastructure delivery. Develop, standardize, and optimize CI/CD pipelines to enable safe, automated deployments (e.g., GitHub Actions, GitLab CI, Jenkins, AWS CodePipeline). Promote DevOps practices: version-controlled infrastructure, automated testing, immutable deployments, and progressive delivery patterns. Establish environment consistency across dev/test/stage/prod and ensure infrastructure drift detection and remediation.
- BCP/DR, RTO/RPO Definition & Testing: Collaborate with stakeholders to evaluate and define service-level RTO and RPO targets based on business and technical requirements. Design and implement BCP/DR architectures and procedures (backups, restore workflows, replication, failover/failback, data integrity validation). Coordinate and execute structured DR tests (tabletop, simulation, partial failover, full failover) and document outcomes. Maintain DR runbooks, dependency maps, and recovery checklists; drive remediation of gaps identified during testing. Produce metrics and reporting on DR readiness, test results, and continuous improvement actions.
Qualifications:
- 7+ years of experience in SRE, DevOps, Platform Engineering, or Systems Engineering roles supporting production environments.
- Strong proficiency with observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft, etc).
- Strong hands-on AWS experience building and operating production systems.
- Proven expertise with Infrastructure as Code (Terraform and/or CloudFormation/CDK).
- Strong CI/CD and automation background (pipeline design, deployment strategies, testing automation).
- Experience defining and validating RTO/RPO, and implementing BCP/DR plans with structured testing.
- Experience with Kubernetes and auto-scaling container platforms (EKS, ECS, or Kubernetes on-prem).
- Strong Linux fundamentals, networking concepts (DNS, TCP/IP, load balancing), and troubleshooting skills.
- Proficiency in at least one scripting/programming language (Python, Go, Bash, or similar).
- Ability to write clear operational documentation, runbooks, and post-incident reports.
- Ability to work effectively in a fast-paced, dynamic and high-intensity environment including open-floor plan if applicable to the position, with timely responsiveness and the ability to work beyond normal business hours when required.
Preferred Qualifications:
- Familiarity with Azure and/or Oracle Cloud (OCI).
- Familiarity with Service Mesh, API Gateways, and distributed tracing tooling.
- Familiarity with OpenTelemetry, client instrumentations and collector configurations.
- Security and compliance familiarity in cloud environments (IAM design, secrets management, audit logging).
- Experience implementing progressive delivery (blue/green, canary), feature flags, and automated rollback.
- Relevant certifications (AWS Solutions Architect/DevOps Engineer, Kubernetes CKA/CKAD).
- Experience with ArgoCD & Karpenter.
Employee Programs & Benefits: CCI offers competitive benefits and programs to support our employees, their families and local communities. These include: Competitive comprehensive medical, dental, retirement and life insurance benefits Employee assistance & wellness programs Parental and family leave policies CCI in the Community: Each office has a Charity Committee and as a part of this program employees are allocated 2 days annually to volunteer at the selected charities. Charitable contribution match program Tuition assistance & reimbursement Quarterly Innovation & Collaboration Awards Employee discount program, including access to fitness facilities Competitive paid time off Continued learning opportunities
- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...SeniorRemote work
- The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams...SeniorFull timeWork at officeLocal area
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk Technology team, you will solve complex and broad business problems...Suggested
- As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management- Maintain and monitor production systems for availability, latency, and performance.- Lead incident response efforts, including communication, resolution, and postmortem...SuggestedPermanent employment
- ...To Supervisor Analytics Cloud Services.Role OverviewThe Release Engineer is responsible for the deployment release and maintenance of... ...systems architects infrastructure and security teams to deliver reliable and scalable cloud solutions.Key ResponsibilitiesCI/CD Pipeline...Senior
- As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting...Work from home2 days per week
- ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely...Full timeLocal area
- As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate Know Your Customer (KYC) Technology group, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business...
- ...Nscale, a GPU cloud for AI, seeks a senior SRE to raise the reliability bar across the platform. You will own the hardest problems, influence architectural decisions, and mentor others while maintaining an on-call rotation that becomes lighter over time. You’ll drive...
$213.1k - $300k
Lead a team of engineers to maintain service uptime while managing global on-call rotations... ...improve operational practices to drive reliability, maintainability, and stakeholder alignment... ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation...Full timeWork at office- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate Know Your Customer (KYC) team, you will solve complex and...
- Job Title- Senior Engineer - Maintenance & Reliability.Description- Conduct maintenance process audits on Nabors’ rigs globally in line with the Company’s Maintenance Audit Standard. Perform inspections of equipment and verify quality of maintenance activities and integrity...SeniorLocal area
$113k - $141.53k
...with a true leader in global energy. Senior Solutions Engineer - Systems Integration serves as a technical... ...in the field, ensuring safe, reliable, and performant operation across diverse... ...Willingness to travel to factories and project sites (25%).Preferred QualificationsMaster’s...SeniorFull timeFor contractorsLocal areaWorldwideFlexible hours- DescriptionSobre nosotros Worley es una empresa global de expertos en energía, químicos y recursos naturales, con sede en Australia. Trabajamos en asociación con nuestros clientes para desarrollar proyectos y generar valor a lo largo del ciclo de vida de sus activos. Nos...Senior
- Senior Systems Software EngineerThis role has been designed as 'Hybrid' with a requirement... ...that matters alongside experienced engineers across software, firmware, and hardware... ...uncover edge cases, performance limits, and reliability opportunitiesDebug and resolve...SeniorFull timeWork experience placementWork at officeLocal areaImmediate start2 days per week
- ...Description Job Description Job Title: System Software Programmer / Senior Software Developer Location: Columbia, SC Job Type:... ...imaging solutions by County. Collaborate with software engineers and technical leaders to understand subsystem and component specifications...SeniorContract workWork experience placement2 days per week
- Senior Principal Presales Systems Engineer, Cloud and AI NetworkingThis role has been designated as ‘Remote/Teleworker’, which means you will primarily work from home.Who We Are:Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live...SeniorFull timeWork experience placementLocal areaImmediate startRemote workWork from home
- ...become part of a global team of over 50,000 planners, designers, engineers, scientists, digital innovators, program and construction... ...deliver a better world. Join us.Job DescriptionAECOM is seeking a Senior Principal Engineer / Lead Engineer to serve as a senior technical...SeniorWork at officeLocal areaWorldwideRelocationFlexible hours
- We are currently recruiting a Senior Software Engineer for a Financial client in Houston. You will be supporting their trading platform and consolidation of their legacy P&L application into the platform.Required experience:-Strong java experience-Python, Kubernetes, cloud...Senior
- ...native services, and infrastructure automation to improve reliability and speed to delivery. Our global team operates in a... ...improvement and operational excellence.About the Role:As a Senior Network Platform Engineer, you will be a key technical contributor responsible for...SeniorFull timeWork at officeFlexible hours
- ...rewarded, and diversity is prized. Come find out how our people power modern life. About this Role We are seeking a Senior ITSM/ITAM Platform Engineer to lead the initial development and the ongoing administration of our enterprise IT Service Management and IT Asset...SeniorFull timeWork at officeLocal areaWorldwide
$61k - $101k
...formal training or certification in software engineering concepts, along with 5+ years of applied... .... We need deep expertise in reliability, scalability, performance, security, enterprise... ...architecture, toil reduction, and other site reliability practices, with the ability...Full time- We are looking for a solid, hands-on Back-End Engineer to join our engineering team and lead application development for our next-generation operational platform.You will be building an AI-enabled platform designed to improve production operations efficiency and reduce...Senior
- ...Responsibilities: RCM Analysis Leadership: Independently lead end-to-end Reliability Centered Maintenance (RCM) studies on rotating, static,... ...client reliability/maintenance managers, operations, and engineering disciplines to validate assumptions, findings and...Senior
$120k - $156k
Houston, TXOnsiteDirect Hire$120k - $156k Engineering Lead (Platform & Architecture) This full-time, onsite opportunity west of Boston is... ...the engineering organization. You'll partner closely with senior technical leadership to evaluate emerging technologies, modernize...SeniorFull timeFlexible hours- JOB SUMMARY As a Senior Software Engineer at Baylor Genetics, you will play a pivotal role in designing, developing, and maintaining high-quality software solutions. This position is ONSITE and requires extensive technical expertise, strong problem-solving skills, and...SeniorFull timeLocal area
- ...company perks await!Job Description:Microchip Technology Inc. has a Senior Engineer II-Applications (Software/Firmware) opening based in Houston,... ...America and internationally, as needed, to work at customer sites to debug firmware and softwareRequirements/Qualifications:...SeniorFull time
$200k
...Senior Mechanical Reliability Engineer Our client is seeking a full-time Senior Mechanical Reliability Engineer. Hybrid. This individual may reside... ...opportunities for operational improvement throughout the site and organization. Provide technical support and expertise...SeniorFull time- ...critically about backlog prioritizationPractical knowledge of software engineering best practices, including Agile development methodologies,... ...well-being for you and your family.WHAT YOU'LL DOAs a Senior Software Engineer with the Life Sciences Data and Analytics Center...SeniorApprenticeshipEasy work
$101.6k - $152.4k
We are seeking a talented Senior Engineer I, Digital Solutions to join our team and take charge of designing, developing, and deploying cutting... ...alarms, and reports.Travel: Willingness to travel to customer sites as required. Travel is roughly expected to be around 25% but...SeniorFull timeTemporary workImmediate startRemote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Houston, TX
- site reliability engineer Houston, TX
- senior technical service engineer Houston, TX
- senior technical consultant Houston, TX
- senior director product management Houston, TX
- senior vice president human resources Houston, TX
- senior automation controls engineer Houston, TX
- senior grant accountant Houston, TX
- senior technical recruiter Houston, TX
- senior compliance officer Houston, TX



