Lead Site Reliability Engineer
$179k - $226kAlloy
Alloy is where you belong! Alloy helps solve the identity risk problem for companies that offer financial products by enabling them to outpace fraud and confidently serve more people around the world. Over 800 of the world's largest financial institutions and fintechs turn to Alloy to take control of fraud, credit, and compliance risk, and grow with the clearest picture of their customers. Through our values: Be Bold, Get Scrappy, Collaborate, and Celebrate Our Differences, we are creating a workplace where you can grow, thrive, and belong. See how we've been continuously recognized and named one of Inc. Magazine's Best Workplaces, Forbes America's Best Startup Employers, Best Fintech to Work for by American Banker, year after year. Check out our investors and read more about us here.
About the team Alloy's Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint: 15+ Kubernetes clusters, 100+ databases, dozens of services, and complex data organization. Our challenge isn't just scale-it's making that scale reliable, secure, and operable with less manual work. We're looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees.
What you'll be doing Reporting to the Engineering Manager of Infrastructure, you'll:
Benefits and Perks
About the team Alloy's Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint: 15+ Kubernetes clusters, 100+ databases, dozens of services, and complex data organization. Our challenge isn't just scale-it's making that scale reliable, secure, and operable with less manual work. We're looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees.
What you'll be doing Reporting to the Engineering Manager of Infrastructure, you'll:
- Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)
- Reduce operational toil by turning manual processes into reliable, repeatable workflows
- Build internal tooling and platforms that enable safe self-service changes for other engineers
- Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)
- Implement and evolve systems for deploying and running applications in Kubernetes
- Contribute to architecture decisions across infrastructure, reliability, and security
- Write and review production-quality code
- Participate in on-call rotations-but focus on building systems that prevent incidents, not just respond to them
- 10+ years of experience in infrastructure, SRE, or software engineering roles
- Strong software engineering skills-you build systems, not just scripts
- Experience managing production infrastructure at scale (cloud + containerized systems)
- Experience with Infrastructure as Code (e.g., Terraform)
- Experience running and troubleshooting distributed systems (Docker/Kubernetes)
- Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.)
- Proficiency in at least one programming language (Python, Go, JavaScript, etc.)
- Experience participating in on-call rotations and improving systems based on incidents
- Strong communication and collaboration skills
- Default to automation over manual processes
- See repetitive work and immediately want to eliminate it
- Think in terms of systems, failure modes, and long-term scalability
- Care about building infrastructure that other engineers can use safely and confidently
- Enjoy working in a small team with high ownership and impact
- Experience running Kubernetes in production at scale
- Deep familiarity with AWS
- Experience building internal platforms or developer tooling
- Background in distributed systems or large-scale data systems
Benefits and Perks
- Unlimited PTO and flexible work policy
- Employee stock options
- Medical, dental, vision plans with HSA (monthly employer contribution) and FSA options
- 401k with 100% match up to 4% of annual employee compensation
- Eligible new parents receive 16 weeks of paid parental leave
- Home office stipend for new employees
- Annual Learning & Development annual stipend
- Well-being benefits include access to ClassPass, OneMedical, UrbanSitter, and Spring Health
- Hybrid work environment: employees are expected to work Tuesdays through Thursdays from our HQ in Union Square, Manhattan. Tasty lunches catered from a variety of local restaurants and frequent employee-organized cultural events contribute to our positive office energy. On Monday/Friday most employees Zoom into work from home while some take advantage of the quieter office.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in New York, NY vacancy
$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly... ...issues to resolution with very infrequent after‑hours support. Lead blameless postmortems and implement long‑term improvements...Suggested$100k - $250k
...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial... ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics...SuggestedLocal area- ...Applications Deployment Responsible for reliability and support of Container Platform on-... ...Perform blameless RCA, partner with engineering and operation teams across the... ...Additional Skills : Automation Process Engineer,Site Reliability Engineer,Full Stack DeveloperThis...Suggested
- ...and Antler, we empower CISOs to proactively manage human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our...SuggestedFull timeWork at office
- ...paced, regulated financial environment. • Excellent communication and collaboration skills. Role Overview The System Engineer will be responsible for designing, implementing, and maintaining enterprise-level infrastructure solutions across Linux...Suggested
- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....Local area
- ...DevOps Engineer DevOps teams in our Infrastructure Engineering group enable Company to continually disrupt the Insure tech space. Our... ...that enables Company Life product teams to ship industry leading and innovative systems that make everyone feel good about life...
$100 per hour
...and deeply committed team from leading tech and accounting companies... ...impact Improve reliability of our systems Build & maintain... ...frameworks and solutions to engineering problems Fast-moving: you... ...benefits ~401k benefits ~ On-site team culture - high...Immediate startWeekend work- ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your...Full timeWork visa
$123k - $165k
...Department/Group Overview Our engineering fleet is a horizontal set of teams providing engineering... .... Our specific team provides reliability engineering and operational support to backend... ...and brands. We are seeking a Site Reliability Engineer who will contribute...$125k - $150k
...Site Reliability Engineer Virtu is a leading financial firm that leverages cutting edge technology to deliver liquidity to the global markets and innovative, transparent trading solutions to our clients. As a market maker, Virtu provides deep liquidity that helps to...Worldwide$127k - $249k
THE TEAM Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ..., alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours- ...Asia and the Middle East. We are creative, low-ego and team-spirited. The Role We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work...Relocation package
- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...CD, ArgoCD). Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis....Flexible hours
$189k - $283.6k
...proactively and reactively improve the reliability of Block's platform and critical infrastructure... ...0) services. In this role, you will lead incident command, coordinate mitigation,... ...strong desire to perform and grow as an engineer ~5+ years of software development...Full timeLocal areaRemote workRelocation packageFlexible hoursShift work- ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications...Flexible hours
$176.75k - $209.1k
...Site Reliability Engineer At Peloton, we view Platform as a Product. A phenomenal platform unlocks speed of development and learning. It allows us to scale easily, enabling our engineers to maximize attention on new features and capabilities. A key to crafting a phenomenal...Temporary work$150k - $175k
...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we're guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed...Remote work- ...self-healing, deployment/rollback automation). Establish reliability standards: SLOs/SLIs, error budgets, production readiness reviews... ..., and release risk controls. Performance and reliability engineering: capacity planning, load/performance analysis, resilience...
- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient...
$225k - $325k
...What you'll do day-to-day Ensure the scalability, reliability, and observability of our systems to maintain and improve the firm's core infrastructure environment. Lead a range of engineering projects, from developing proprietary platforms for configuration...Hourly pay$140k - $170k
...optimize business interactions. Role Description: As a Site Reliability Engineer, you will work with Agile engineering teams to provide... ...work both independently and collaboratively. Ability to lead and work on projects. Ability to multitask and adapt quickly...Local area$152.5k - $219.2k
...global cloud platform. As a team of six engineers distributed across the US, Canada, and the... ...with a strong focus on automation, reliability, and operational excellence. We are one... ...Qualifications ~2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$115k - $125k
...Site Reliability Engineer New York City, NY Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions. As financial industry experts at the center of markets and technology, we help...Work experience placementWork at officeWork from homeMonday to FridayFlexible hoursShift workWeekend workAfternoon shiftEarly shift$194k - $267k
...something more than once, automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$175k - $230k
...those residing in senior living facilities. Falls are the leading cause of injury-related death among adults over 65. And yet... ...be a 24x7, highly available platform for elder care. As a Site Reliability Engineer, you'll partner with engineering teams across the organization...ApprenticeshipWork at officeLocal areaRemote work2 days per week$205k - $305k
...Director Of Site Reliability Engineering Interested in working on cutting-edge blockchain technology and creating equitable access to the global... ...looking for a Director of Site Reliability Engineering to lead a small, high-leverage SRE team and help shape how engineering...Temporary workWork at officeLocal areaWorldwideFlexible hours$131k - $164k
...Position Overview We are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and... ...they need to drive greater impact and accountability - to lead with purpose. Our employees are passionate, smart, and creative...Work at officeLocal areaVisa sponsorshipFlexible hours- ...strategy sessions with other Optum Teams Require 2+ years of experience with Terraform Require 2+ years of experience with DevOps Solution Architect, DevOps, or System Engineer certification in one or more public cloud providers Terraform certification #J-18808-Ljbffr...
- ...coasts. If you're driven by impact, pace, and raising the bar. This is the place. The Role As a Staff Site Reliability Engineer you'll play a lead role on the founding SRE team at our new NYC engineering hub. You'll own multi-team reliability and...Work at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
Related searches
- lead web developer New York, NY
- lead algorithm engineer New York, NY
- lead network engineer New York, NY
- lead infrastructure engineer New York, NY
- lead engineer New York, NY
- lead operating engineer New York, NY
- lead system engineer New York, NY
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- site reliability engineer sre New York, NY

