Senior Site Reliability Engineer
Castleton Commodities International
The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams to design resilient cloud-native architectures, implement Infrastructure as Code (IaC) and CI/CD standards, and drive measurable reliability outcomes. The Senior Site Reliability Engineer will also lead efforts to define and validate recovery objectives (RTO/RPO), design and implement Business Continuity / Disaster Recovery (BCP/DR) plans, and coordinate structured testing to ensure readiness.Responsibilities:Reliability Engineering & OperationsOwn and improve service reliability through SLO/SLI definition, error budgets, and operational best practices.Design, implement, andmaintainobservability (monitoring, logging, tracing, alerting) to reduce MTTR and improve proactive detection.Lead incident response practices including on-call improvements, runbooks, post-incident reviews (RCA), and preventative actions.Partner with application teams to improve performance, capacity planning, and resiliency under failure scenarios.Infrastructure & Cloud ArchitectureDesign andoperatehighly available, fault-tolerantCloudarchitectures (multi-AZ and, whererequired, multi-region).Implementresilient patterns across compute, storage, networking, and managed services (e.g., autoscaling, load balancing, backups, replication).Drive cloud governance best practices (tagging, account/landing zone patterns, least privilege, guardrails) in partnership with security and platform teams.Infrastructure as Code (IaC) & DevOps EnablementBuild and maintainIaCmodules and standards (e.g., Terraform, CloudFormation, CDK) for repeatable, auditable infrastructure delivery.Develop, standardize, andoptimizeCI/CD pipelines to enable safe, automated deployments (e.g., GitHub Actions, GitLab CI, Jenkins, AWS CodePipeline).Promote DevOps practices: version-controlled infrastructure, automated testing, immutable deployments, and progressive delivery patterns.Establish environment consistency across dev/test/stage/prod and ensure infrastructure drift detection and remediation.BCP/DR, RTO/RPO Definition & TestingCollaborate with stakeholders to evaluate and define service-level RTO and RPO targets based on business and technical requirements.Design and implement BCP/DR architectures and procedures (backups, restore workflows, replication, failover/failback, data integrity validation).Coordinate and execute structured DR tests (tabletop, simulation, partial failover, full failover) and document outcomes.Maintain DR runbooks, dependency maps, and recovery checklists; drive remediation of gapsidentifiedduring testing.Produce metrics and reporting on DR readiness, test results, and continuous improvement actions.Qualifications:7+ years of experience in SRE, DevOps, Platform Engineering, or Systems Engineering roles supporting production environments.Strongproficiencywith observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft,etc).Strong hands-on AWS experience building and operating production systems.Provenexpertisewith Infrastructure as Code (Terraform and/or CloudFormation/CDK).Strong CI/CD and automation background (pipeline design, deployment strategies, testing automation).Experience defining and validating RTO/RPO, andimplementing BCP/DR plans with structured testing.Experience with Kubernetes andauto-scalingcontainer platforms (EKS, ECS, or Kubernetes on-prem).Strong Linux fundamentals, networking concepts (DNS, TCP/IP, load balancing), and troubleshooting skills.Proficiencyin at least one scripting/programming language (Python, Go, Bash, or similar).Ability to write clear operational documentation, runbooks, and post-incident reports.Ability to work effectively in a fast-paced, dynamic and high-intensity environment including open-floor plan if applicable to the position, with timely responsiveness and the ability to work beyond normal business hours when required.Preferred Qualifications:Familiarity with Azure and/or Oracle Cloud (OCI).Familiarity with Service Mesh, API Gateways, and distributed tracing tooling.Familiarity withOpenTelemetry, client instrumentations and collector configurations.Security and compliance familiarity in cloud environments (IAM design,secretsmanagement, audit logging).Experience implementing progressive delivery (blue/green, canary), feature flags, and automated rollback.Relevant certifications (AWS Solutions Architect/DevOps Engineer, Kubernetes CKA/CKAD).Experience with ArgoCD & Karpenter.Employee Programs & Benefits:CCI offers competitive benefits and programs to support our employees, their families and local communities. These include:Competitive comprehensive medical, dental, retirement and life insurance benefitsEmployee assistance & wellness programsParental and family leave policiesCCI in the Community: Each office has a Charity Committee and as a part of this program employees are allocated 2 days annually to volunteer at the selected charities.Charitable contribution match programTuition assistance & reimbursementQuarterly Innovation & Collaboration AwardsEmployee discount program, including access to fitness facilitiesCompetitive paid time offContinued learning opportunitiesVisit to learn more!#LI-CD1SummaryLocation: Stamford, CT; Houston, TXType: Full time
- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...SeniorRemote work
- ...Nscale, a GPU cloud for AI, seeks a senior SRE to raise the reliability bar across the platform. You will own the hardest problems, influence architectural decisions, and mentor others while maintaining an on-call rotation that becomes lighter over time. You’ll drive...Suggested
- ...accelerate autonomy development. We are seeking a software engineer with strong C++ expertise and a passion for building scalable simulation... ...in architecture and technical design discussions Build reliable, maintainable, and well-tested systems Contribute to code...SeniorFull time
- ...Ventures, and many more. About you: You are a talented engineer who thrives in a collaborative and fast-paced environment. You... ...Science or relevant experience. About the role: As a Senior Engineer, you will play a pivotal role in designing and implementing...SeniorFull time
- ...they're harder. Knows when to cut corners and when to build for the long haul. Bridges the gap between business goals and engineering reality without losing sight of either. Guides and grows other developers—teaching them how to think, not just how to code....Senior
- ...JOB DESCRIPTION As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management - Maintain and monitor production systems for availability, latency, and performance. - Lead incident response efforts, including communication...Permanent employmentFull time
- ...A Senior Applications Engineer II is responsible for assessing the costs and resources required for a project or service. They analyze project specifications... ...single-lines, three-lines, elevations, control schematics, site plans and electrical enclosure layouts Generate accurate...SeniorFull timeFor subcontractorWork at office
$140k - $170k
...financial industry? Our client is seeking a highly motivated Senior Software Developer to design and build scalable, secure, and high... ...to trade and risk management systems, and mentor junior engineers while aligning technical delivery with business goals. About...SeniorFull timeLocal areaWorldwideVisa sponsorshipWork visa- ...To Supervisor Analytics Cloud Services.Role OverviewThe Release Engineer is responsible for the deployment release and maintenance of... ...systems architects infrastructure and security teams to deliver reliable and scalable cloud solutions.Key ResponsibilitiesCI/CD Pipeline...Senior
- ...RabbitMQ Security (OAuth2, OIDC, JWT, IAM) Performance Engineering & Production Troubleshooting AI Requirements... ...engineering teams Solves complex production issues Drives scalability, reliability, security, and AI-enabled solutions...Senior
- ...Senior Software Developer Benchmark Insurance Group Location : 827 N Loop W Suite B, Houston, TX 77008 Job Type: Full-time... ...growth and leadership opportunities Work Location On-site: 827 N Loop W Suite B, Houston, TX 77008 Schedule ~ Monday...SeniorFull timeRemote workMonday to FridayFlexible hours
$157k - $298k
...educational, social, and developmental outcomes in ways that create meaningful and lasting impact. ABOUT THE ROLE As a Senior AI Applications Engineer at Bezos Academy, you will help build AI-powered software applications that support teachers’ abilities to deliver...SeniorFull timeRelocationRelocation packageFlexible hours- ...ENGINEERLocation: HOUSTON, TXFLSA Class: EXEMPTResponsible to: Directo of Software EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely...Full timeLocal area
$61k - $101k
...formal training or certification in software engineering concepts, along with 5+ years of applied... .... We need deep expertise in reliability, scalability, performance, security, enterprise... ...architecture, toil reduction, and other site reliability practices, with the ability...Full time- As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing products. Responsibilities will include collecting system requirements; improving existing tools and processes through scripting...Work from home2 days per week
$140k - $200k
...experiencing exponential growth. Overview We're looking for a Senior Software Engineer to join our Core Experiences Team. This team builds and... ...strategically, and is passionate about designing clear, reliable APIs and simple systems that directly enhance the user...SeniorFull timeRemote work- ...Senior Systems Software Engineer This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from... ...scenarios to uncover edge cases, performance limits, and reliability opportunities Debug and resolve challenging issues...SeniorWork experience placementWork at officeLocal areaImmediate start2 days per week
- ...Engineering Lead (Platform & Architecture) This full-time, onsite opportunity west of Boston is with a growing technology organization... ...real-world systems and hardware. You'll partner closely with senior technical leadership to modernize legacy applications, evaluate...SeniorFull time
- ...JOB SUMMARY As a Senior Software Engineer at Baylor Genetics, you will play a pivotal role in designing, developing, and maintaining high-quality software solutions. This position is ONSITE and requires extensive technical expertise, strong problem-solving...SeniorLocal area
- ...Role Overview PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s... ...people leader who can coach managers and represent SRE to senior engineering and business stakeholders. Key Responsibilities...
- ...Software Engineer Duration: 7+ Months Location: Houston, Texas (Hybrid) Job Description Guide and Support other team members to create excellence in delivery Align team targets to business goals and focus on team efforts (Bottom Line) Perform Real...SeniorLocal area
$120 per hour
...Senior Python Software Engineer (Commodities) Role Role: Senior Python Software Engineer (Commodities) Client: Leading Commodities Trading... ...that power real-time trading decisions, ensuring reliability and performance in a front-office environment. Engineers...SeniorHourly payContract work$160k - $220k
...Senior Software Engineer, Applied AI Houston, Texas (preferred), or the San Francisco Bay Area Company Introduction At Bot Auto, we are... ...directly with users to convert high-value workflows into reliable products. Required Qualifications ~6+ years of relevant...Senior$213.1k - $300k
...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering Google Houston, TX, USA Advanced Experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders...Full timeWork at office- The Cooper Companies Inc. is looking for a Staff Bioinformatics Software Engineer to join the Cooper Surgical R&D Software team in Houston, Texas. In this role, you will design, develop, and maintain advanced software solutions for genetic testing products, focusing on...Senior
- ...Senior Software Engineer As a Senior Software Engineer, you will design, build, and enhance enterprise software applications, AI-driven solutions... ..., and services while improving system performance, reliability, and security. Design and document application and system...Senior
- ...Senior Software Engineer Location: Houston, Texas (Hybrid, potential for Remote) Role Overview Our organization is seeking a Senior Software Engineer for a newly formed team focused on modernizing complex legacy applications. This role is pivotal in transforming...SeniorRemote work
$135k - $200k
...Posting/External Job Title Software Engineer (Senior-Level) Location Houston, TX 77058 US (Primary) Remote/Onsite/Hybrid 100... ...implementation and deployment. Develop robust, well-architected, and reliable software, typically using C/C++ and Python on Linux....SeniorPermanent employmentFull timeContract workLocal areaRemote workWeekend workAfternoon shift$124.2k - $207k
...relevant educational attainment and/or training. Job Title: Senior Software Engineer Location: Lake Forest, CA/ Houston, TX/ Philadelphia, PA... ...at a local AVEVA office, with an expectation of being on-site 50% of your working hours to support collaboration and...SeniorFull timeWork at officeLocal areaRemote workFlexible hours- Job Title- Senior Engineer - Maintenance & Reliability.Description- Conduct maintenance process audits on Nabors’ rigs globally in line with the Company’s Maintenance Audit Standard. Perform inspections of equipment and verify quality of maintenance activities and integrity...SeniorLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Houston, TX
- site reliability engineer Houston, TX
- senior technical service engineer Houston, TX
- senior technical consultant Houston, TX
- senior director product management Houston, TX
- senior vice president human resources Houston, TX
- senior automation controls engineer Houston, TX
- senior grant accountant Houston, TX
- senior technical recruiter Houston, TX
- senior compliance officer Houston, TX




