SRE
The Cypress Group
Job Description
Job Description
Location: New York City (Nolita)
Work Arrangement: Full-time, onsite/hybrid (based on team expectations)
About the Role
We're looking for a Senior Site Reliability Engineer to help build and scale the reliability foundation of our engineering organization. You'll lead initiatives that improve system availability, observability, deployment reliability, and operational excellence while enabling engineering teams to ship quickly and confidently.
This is a highly impactful role with significant ownership, where you'll influence architecture, establish SRE best practices, and help scale infrastructure supporting millions of users.
What You'll Do
- Lead production incident response and continuously improve operational readiness.
- Build and evolve sustainable on-call practices, including runbooks, automation, and blameless postmortems.
- Design and maintain modern observability solutions that provide actionable insights into application health and performance.
- Develop and manage cloud infrastructure using Infrastructure as Code.
- Partner with engineering teams to build resilient, scalable systems across databases, networking, asynchronous processing, and data infrastructure.
- Improve CI/CD pipelines and deployment strategies to enable reliable, low-risk software releases.
- Advocate for reliability, scalability, and operational excellence throughout the software development lifecycle.
- Help define engineering standards and best practices as the organization continues to grow.
What We're Looking For
Required Qualifications
- 5+ years of experience in Site Reliability Engineering or DevOps, or 7+ years as a Software Engineer with significant infrastructure ownership.
- Demonstrated experience supporting highly available, production-scale systems.
- Strong expertise leading incident response, conducting root cause analysis, and driving continuous operational improvements.
- Experience designing resilient cloud infrastructure with automated scaling, health monitoring, and fault-tolerant architectures.
- Deep knowledge of AWS services and Infrastructure as Code using tools such as Terraform.
- Experience implementing monitoring and observability platforms such as Datadog, Prometheus, Grafana, ELK, or similar.
- Experience designing and maintaining modern CI/CD pipelines and deployment automation.
Preferred Experience
- Experience supporting high-growth SaaS or fintech platforms.
- Familiarity with container orchestration and cloud-native architectures.
- Strong scripting or programming skills for infrastructure automation.
- Excellent written and verbal communication skills with experience documenting operational processes.
- Passion for mentoring engineers and fostering a culture of operational excellence.
Who You Are
You'll thrive in this role if you:
- Take ownership of complex infrastructure challenges.
- Enjoy solving reliability and scalability problems before they become customer issues.
- Balance engineering rigor with pragmatic decision-making.
- Collaborate effectively across engineering teams to improve platform health.
- Enjoy building systems, processes, and tooling that help the entire organization move faster.
- ...We are growing and have been recognized by G2 as one of 2024's Top 50 Software Companies in the world. Glassbox is looking for an SRE to join our global Cloud team. We are Glassbox, and our mission is to reveal the insights that empower organizations to deliver exceptional...SuggestedRemote jobNight shift
$205k - $225k
...access to credit. We’re restoring dignity and paving the way for millions to achieve financial freedom. About the Role As a Senior SRE you will ensure the reliability, scalability, and performance of January’s production and internal systems as we scale from thousands...SuggestedCurrently hiringWork at office$160k - $180k
Overview Socure is building the identity trust infrastructure for the digital economy — verifying 100% of good identities in real time and stopping fraud before it starts. The mission is big, the problems are complex, and the impact is felt by businesses, governments, and...Suggested$70k
...discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity...SuggestedPermanent employmentFull timeRelocationVisa sponsorshipWork visaFlexible hours$111.5k - $159k
...& Vulnerability Remediation Support Support AWS remediation by validating application impact, workload criticality, reboot windows, SRE coordination, dependencies, and post-patch sign-off. Support EOL OS remediation through code/configuration changes, dependency remediation...SuggestedShift work- ...assured they are accessing the loan market’s most credible source of deal information. About You: Versana is seeking a motivated SRE/DevOps Engineer with strong observability experience to join our growing Platform Engineering squad. The squad’s goal is to manage...Work experience placementLocal area
- ...leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You will own the...Full timeWork at office
- ...meet audit and governance expectations.What You'll Bring:5+ years in infrastructure operations, platform engineering, cloud operations, SRE, or enterprise patching.Experience with Windows, Linux, middleware, database, virtualization, cloud infrastructure, or AWS EC2...Local area
- ...candidate for the Head of Infrastructure position at the company will have a background of at least 5 years in infrastructure, DevOps, or SRE roles, with a strong emphasis on engineering leadership. Proficiency in Kubernetes is essential, and experience with machine learning...
$165k - $190k
...of the cloud infrastructure for our products. This position involves close collaboration with Application Developers, Data Engineers, SRE, Security, and IT professionals to automate and streamline operations and processes. Your contributions will be essential in ensuring...- ...reusable, and efficient code following engineering best practices and internal guidelines. Collaborate with product managers, architects, SRE teams, and other developers to translate business requirements into engineering solutions. Integrate workflows with external systems...Full timeWeekday work
$160k - $200k
...60,000.00/yr - $200,000.00/yr Direct message the job poster from AppCard, Inc. We are seeking an experienced Infrastructure Engineer/SRE to design, build, and maintain our hybrid infrastructure, which spans both cloud and on‑premises environments . This role is hands‑on...Full time$80k - $100k
...documentation and internal enablement materials for platforms, pipelines, and standards. Required Qualifications 4-7+ years in DevOps/SRE/Platform Engineering roles supporting enterprise cloud applications. Hands‑on with IaC (Terraform or CloudFormation/ARM/Bicep) and...Full timeLocal areaRemote work$150k - $300k
About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in the world to troubleshoot, remediate, and even prevent the most complex production incidents. Our mission is to free engineers from...Full timeWork at officeFlexible hours- ...Enforce architectural patterns and standards Develop infrastructure design solutions in collaboration with CPS core engg team Leverage SRE organization embedded in application teams to scale for the PEArch charter Keep rest of the CPS org apprised on what is coming soon...Contract workFlexible hours
$30 - $33 per hour
...to a wide range of technologies and application types Excellent foundation for careers in security engineering, platform operations, SRE, or infrastructure roles Structured learning path with hands‑on experience in a critical domain Opportunity to become a trusted...Contract workRemote work$135k - $200k
...familiarity with containers (Docker) and orchestration (Kubernetes). Experience working with a cloud provider (AWS / Azure / GCE) or sysadmin/SRE experience in data centers. Up to date with modern industry practices and open‑source advancements. Solid understanding of...Work at officeRemote workWork from homeRelocation packageFlexible hours2 days per week1 day per week- ...of Engineering. What you will do Technical Breadth across disciplines. Focus deeply on one problem at a time while moving between SRE, DevOps, Infrastructure, and Platform work as priorities shift. Most of the time you will be heads‑down on a substantial initiative such...Full timeWork at officeLocal areaFlexible hoursShift work
- ...policy‑as‑code (OPA, Sentinel) and IaC security scanning. Experience with Terraform Cloud/Enterprise, Terragrunt, or similar. Prior SRE or platform engineering experience building developer self‑service portals. Benefits Opportunity to travel (if applicable)...Work at officeLocal areaFlexible hours
- ...reliable infrastructure for data processing and content classification workloads Who you are: You have 5+ years of experience in DevOps, SRE, or infrastructure engineering, with a strong focus on cloud‑native environments (preferably AWS) You have managed cloud...Work at officeWorldwideMonday to Friday3 days per week
- ...how engineering teams detect, investigate and remediate cloud native deployments at scale by creating the industry’s first agentic AI SRE platform. Kubernetes promises agility, elasticity, reliability and high availability, but it also introduces complexity, high...Remote work
- ...Participate in all phases of solution development, from initial planning and design to implementation and go-live, ensuring alignment with SOE/SRE standards Design and develop reusable, scalable data models that effectively meet business needs Foster a culture of technical...Full timeContract workWork experience placementRemote work
$60 per hour
...and New Jersey. The Wireless Engineer will architect and support advanced netwo Salary: $60.00 hourly Employment Type: Full-time SRE Engineer (W2) Jersey City, NJ (Onsite) - 6 Months Contract to Hire Proficient in application development skills for more than one technology...Hourly payFull timeContract workLocal area$178.5k - $241.5k
...Unicorns Tech Leads across the company to ensure the team delivers mission value across platform engineering, DevSecOps, cloud, and SRE disciplines. First-Year Outcomes Success in the first year will be measured by three primary outcomes: A smooth transition from the...Contract workFor contractorsFor subcontractorLocal areaRemote work- ...reliability improvements, and collaborate closely with product and backend teams. Common backgrounds for this role include DevOps Engineer, SRE, Platform Engineer, or Backend Engineer (Infrastructure). What You’ll Do Design, build, and operate AWS-based infrastructure....Full timeWork at officeFlexible hours
- ...Design and package applications with latest containerization and orchestration. Configure logging, metrics, tracing, and alerting - apply SRE practices. Produce cost-and-capacity reports with FinOps tooling. Produce well-structured design documents, diagrams, and hand-over...Full time
$180k - $230k
...activities. Up-to-date knowledge with a range of foundational internet technologies, eg TCP/IP, TLS, DNS & BGP. Effective deployment of SRE to achieve consistent, high‑availability technology operations. Hybrid This position currently has a hybrid schedule, which requires...Work at office$60k - $140k
...including security, scalability, observability (logging/metrics/tracing), resilience, and disaster recovery. Collaborate with DevOps/SRE to support CI/CD pipelines, deployment architecture, and operational readiness. Lead architecture and design reviews, mentor...Work at office- ...MCP, OAuth, API gateways, or multi-tenant integration platforms. Experience with modern observability (OTEL, distributed tracing) and SRE best practices. Experience building internal admin or developer tooling consumed by both humans and agents. Experience working on...Work at officeFlexible hoursNight shift
$250k
...Unlimited PTO | Equinox Membership | $5,000 Lifestyle Stipend | $5,000 Relocation Support Experience 4+ years in Platform, Infrastructure, SRE, or DevOps Engineering Scaled production cloud infrastructure (AWS, GCP, and/or Azure) Strong K8s, Docker, containerisation...Relocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!

