Senior Site Reliability Engineer
Radley James
A leading High-Frequency Trading firm is seeking an experienced Senior Site Reliability Engineer.
This is a critical role where reliability, performance, automation and operational excellence are essential. You will work closely with software engineers, infrastructure teams and traders to build and maintain highly available, resilient and scalable platforms supporting time-sensitive trading systems.
The ideal candidate will have a strong background in Site Reliability Engineering, DevOps and cloud infrastructure, with significant hands-on experience working with AWS.
The Role
As a Senior SRE, you will be responsible for ensuring the reliability, scalability and performance of the technology platforms.
You will:
- Design, build and maintain highly available and resilient infrastructure on AWS.
- Develop and improve automation across infrastructure, deployment and operational processes.
- Establish and maintain monitoring, observability, alerting and incident-management practices.
- Work closely with development teams to improve application reliability and performance.
- Participate in the design and implementation of highly resilient systems supporting trading and business-critical workloads.
- Lead incident response, troubleshooting and root-cause analysis for complex production issues.
- Identify and eliminate recurring operational problems through automation and engineering.
- Contribute to capacity planning, performance optimisation and disaster-recovery strategies.
- Improve CI/CD pipelines and deployment processes.
- Champion SRE and DevOps best practices across the engineering organisation.
- Mentor engineers and provide technical leadership on reliability and infrastructure matters.
Essential Experience
We are looking for candidates with:
- 5+ years' experience in SRE, DevOps, Platform Engineering or a similar infrastructure-focused role.
- Strong, hands-on AWS experience in production.
- Excellent understanding of AWS services such as EC2, EKS/ECS, S3, IAM, VPC, CloudWatch, RDS and Route 53.
- Strong Linux/Unix administration and troubleshooting skills.
- Experience with Infrastructure as Code, ideally Terraform.
- Strong scripting/programming ability with Python.
- Experience with Kubernetes and containerised environments.
- Strong understanding of CI/CD principles and tooling.
- A strong understanding of networking, security and distributed systems.
- Proven experience managing production incidents and conducting root-cause analysis.
- Experience building systems with high availability, resilience and fault tolerance in mind.
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre New York, NY
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- senior service associate New York, NY
- senior safety specialist New York, NY
- senior vice president of business development New York, NY
- senior service designer New York, NY
- senior sales recruiter New York, NY
- senior mulesoft developer New York, NY
- senior media manager New York, NY
