Lead Site Reliability Engineer
$184.2k - $240kMovable Ink
Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility. Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.
As one of our Lead Site Reliability Engineers, you will combine hands‑on technical expertise with strategic technical leadership across infrastructure and software development. You will own the design and evolution of major systems within our multi‑cloud, multi‑region, active‑active content serving platform that serves upwards of 25 Billion requests daily. Through a combination of architectural vision, cross‑team collaboration and mentorship, you will help drive the reliability initiatives and define the technical strategy that scales our platform to 50 Billion requests per day and beyond.
Responsibilities
- Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents
- Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long‑term business objectives
- Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization
- Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios
- Lead cross‑functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery
- Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.
Qualifications
- Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long‑term reliability strategy
- Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges
- Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution
- Deep, hands‑on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi‑cloud architecture and strategy (AWS and GCP).
- Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo
- Experience leading on‑call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on‑call rotation
- Expert‑level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef
- Advanced Kubernetes expertise, including cluster architecture design, multi‑tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE
- Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting
- Advanced Linux systems expertise, with the ability to diagnose complex system‑level issues and mentor others on performance tuning and troubleshooting
The base pay range for this position is $184,200 - 240,000 /year, which can include additional bonus depending on the position ultimately offered, in addition to a full range of medical, financial, and/or other benefits. The base pay offered may vary depending on job‑related knowledge, skills, and experience.
Studies have shown that women, communities of color, and historically underrepresented people are less likely to apply to jobs unless they meet every single qualification. We are committed to building a diverse and inclusive culture where all Inkers can thrive. If’re excited about the role but don’t meet all of the abovementioned qualifications, we encourage you to apply. Our differences bring a breadth of knowledge and perspectives that makes us collectively stronger.
We welcome and employ people regardless of race, color, gender identity or expression, religion, genetic information, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, ethnicity, family or marital status, physical and mental ability, political affiliation, disability, Veteran status, or other protected characteristics. We are proud to be an equal opportunity employer.
#J-18808-Ljbffr- ...The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve... ...part of a cutting-edge technology company building industry leading products, please apply.Shutterstock is an Equal Opportunity...SuggestedFull timeWork experience placementRemote work
- ...DEPARTMENT: Product Engineering / Operational Readiness REPORTING... ...team is responsible for the reliability, monitoring, automation, and... ...connectivity solutions for the world's leading financial institutions.... ...services. Role Overview: The Site Reliability Engineer II (SRE)...SuggestedFull timeTemporary workWork at officeRemote workFlexible hours
- ...Ireland Full time Are you passionate about building reliable, scalable systems that power critical business solutions?... ...can learn more about LexisNexis Risk at the Role As a Site Reliability Engineer (SRE), you will bridge software development and IT operations...SuggestedFull timeWork from home
$100k - $135k
...Model-Based Manufacturing, where context-aware production planning informs design in real-time—eliminating the disconnect between engineering and manufacturing. Our platform automates what can be automated and captures tribal knowledge where automation falls short,...SuggestedWork at officeFlexible hours$104k - $178k
## Sr. Site Reliability Engineer IApply: Hybrid: NYC Global HQ: Full time: Posted 12 Days Ago: JR00000779# ****Who We Are****DV is the leader in... ...incident reviews to minimize downtime and prevent recurrence.* Lead technical projects from planning through deployment,...SuggestedFull time$100k - $250k
...systems) can be yours. What you’ll do Improve observability, reliability and availability by defining and measuring key metrics.... ...function improvements. Educate, mentor and hold accountable the engineering team to improve the reliability of our systems and make...Local area$171.6k - $223k
...development and learning. It allows us to scale easily, enabling our engineers to enhance attention on new features and capabilities. A key... ...of members all over the world. Peloton is looking for a Site Reliability Engineer to create the tooling and services which simplify...Temporary workLocal area- ...Improve the reliability of mission critical solutions, applications, and platforms Software... ...issues that affect applications Lead efforts to troubleshoot and/or debug issues... ...Windows and Linux Years of Experience: 5 Years of Software Engineering #J-18808-LjbffrWork experience placement
$207k - $300k
...pushing for changes that improve reliability and velocity.Practice... ...languages.3 years of experience leading projects.3 years of experience... ...degree in Computer Science or Engineering.Experience mentoring... ...across cross-functional teams.Site Reliability Engineering (SRE)...$208.5k - $216.5k
...including its signal quality and its cost.Drive reliability improvements using SLOs and telemetry... ...access paths.Shape incident management: lead incident response, postmortems, failure... ...years in SRE, DevOps or infrastructure engineering, though scope and impact weigh more than...Full timeTemporary workFlexible hours- ...launch. Our team brings former experience at the world's leading quantitative firms, including Two Sigma, Flow Traders,... ...for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability...Full timeLocal areaRemote work
$120k - $160k
...beautiful modern office, daily catered lunches, and more. As a Site Reliability Engineer (SRE), you will work at the intersection of production... ...and trading systems Diagnose and fix bugs in code Lead complex deployments Automate manual workflows Track and...Work at officeLocal area- ...millions of global users and a reputation for rigorous engineering. The Role As Senior Site Reliability Engineer , you will own the infrastructure... ...the internal authority on container orchestration Leading application migration efforts onto Kubernetes in...
$120k - $150k
...our teams and communities. We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.... ...with cloud security posture or compliance concepts. As a leading investment bank, we enable growth and success for our clients...- ...Responsibilities Support and enhance the reliability, availability, and performance of... ...service improvements. Work closely with engineering teams to implement, test, and deploy... ...production environments, platform engineering, site reliability, DevOps, or infrastructure...
$150k - $170k
...Senior Site Reliability Engineer – Zip CoJoin to apply for the Senior Site Reliability Engineer role at Zip CoAt Zip, we build cloud-native software... ...Union Square office with a casual dress codeIndustry-leading, employer-sponsored insurance for you and your dependents,...Casual workWork at officeRemote work- ...About Ontrac Solutions Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that... ...Overview We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring...
$150k - $200k
...Defence and Government. They’re continuing their expansion of their New York (and Washington DC) engineering teams and looking for an Infrastructure Engineer / Site Reliability Engineer / Forward Deployed Infrastructure Engineer with a strong software engineering...- ...Job Description Chariot’s engineering hire will be responsible for taking the Chariot platform to the next level. You will lead the evolution of our banking product, DAF product and technology strategy, along with building a growing team of software engineers. You will...Work experience placementWork at officeWork from homeMonday to FridayMonday to Thursday
$200k - $240k
...better patient care. Backed by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health,... ...to meet you! About the role We're looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team and get their...Work at office3 days per week- ...infrastructure, behind their own controls, with the reliability and operational clarity they would... ...TAMs to trust. Partner with product engineers on infrastructure requirements for new... ...when they introduce new dependencies Lead through ambiguity, make careful risk...
- ...part of our global expansion, we’re looking for a hands‑on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our... ...reliability and resiliency are built in from day one. Lead incident response: Drive on‑call processes, conduct root‑cause...Remote work
- ...Plaid Inc is looking for a Staff Site Reliability Engineer to lead the reliability practices across product engineering. You will architect SLO and error-budget programs, ensuring new products are production-ready while promoting safety gates for high velocity. The...
- ...Site Reliability Engineer Our Client, a multinational telecommunications technology company is seeking a Site Reliability Engineer (SRE I) to join our Video Platform Engineering Team. As a Level 1 SRE, you will work closely with senior engineers to respond to incidents...Temporary work
- ...Job Description Job Description Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well)...
$191k - $226k
...anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the... .... You will run the machine: defining and upholding SLOs, leading incident response, and driving the automation and standards...Remote workWork visaFlexible hours$182.8k - $247.3k
...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure the company’s sophisticated distributed...Work experience placement- ...right care, at the right time, in the right setting. Role Description You'll join Tennr's Infrastructure team as a Site Reliability Engineer, focused on the systems that keep us reliable, observable, and secure. This is a hands-on role with real room to grow: you...Work at office
$120k - $180k
...people, and works with high-profile manufacturers including leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield opportunity to architect the path from AWS to on-premises...Permanent employmentFull timeRelocation package- ...We are a mission-driven team of brilliant minds trusted by leading organizations including Intermountain Health, OSF HealthCare... ...| Careers About The Role We are looking for a Site Reliability Engineer to help us evolve and safeguard the infrastructure powering...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead operating engineer New York, NY
- lead algorithm engineer New York, NY
- lead security engineer New York, NY
- lead infrastructure engineer New York, NY
- lead web developer New York, NY
- lead system engineer New York, NY
- lead app. developer New York, NY
- lead network engineer New York, NY
- lead engineer New York, NY
- site reliability engineer New York, NY



