Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$105.79k - $141.05k
Full-time

Lumen

Lumen is the trusted network for the AI‑powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high‑performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities.

At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where performance, security, and resilience matter.

This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you’re ready to take ownership, deliver meaningful impact, and help shape the future of AI‑ready connectivity, join us today.

The Role

We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is critical to ensuring the reliability, scalability, and efficiency of our systems, with a strong emphasis on AWS infrastructure, observability, automation, and AI-assisted engineering practices.

The Lead SRE requires an AI-native mindset, understands the software development lifecycle (from coding to support) and applies modern AI tools to enhance productivity, quality, and operational excellence. This role will shape how Lumen combines the latest technologies, including AI-driven automation, to modernize software delivery and application lifecycle management.

This role will collaborate with key stakeholders across the engineering organization — including product owners, developers, and testers — to design, optimize, and automate business and technical processes, while effectively navigating multiple teams within a large and complex organization.

Location

This role is designated as a fully remote position within the United States.

The Main Responsibilities

Production Support & Incident Management

  • Implement AI systems and automations to assist during ongoing outages and triage potential ones. You will work with the development teams to ensure that they have all the data normally needed during an outage at their fingertips including preliminary analysis by AI.
  • Provide Tier 3 support for issues across portal services by troubleshooting and resolving technical issues in test and production environments.
  • Lead root cause analysis and post-mortem processes, incorporating AI-assisted analysis and pattern detection to ensure continuous improvement.

Performance Optimization

  • Monitor system performance and proactively identify bottlenecks or degradation using AI-driven observability and anomaly detection tools.
  • Implement tuning strategies across application layers, databases, and infrastructure.
  • Drive initiatives to improve latency, throughput, and resource utilization.

Monitoring & Observability

  • Deploy improved alerting for Lumen Connect in depth, focusing on outside in but including early indicators for fulfilment and other areas. Combining traditional monitoring with AI-based anomaly detection and noise reduction.
  • Proactively monitor the errors and performance on Lumen Connect. Implement rules to detect deviations, implement improvements together with the teams.
  • Design and maintain dashboards, alerts, and metrics using tools like Datadog, AppInsights, CloudWatch, or similar.

Automation & Infrastructure as Code

  • Develop and maintain automation scripts and tools for deployment, scaling, and recovery.
  • Develop and maintain automation scripts and tools for deployment, scaling, and recovery, leveraging AI-assisted code generation and validation tools
  • Use Terraform, or similar IaC tools to manage AWS resources.

Reliability Engineering

  • Perform an in-depth analysis of the overall system and its dependencies, implementing techniques to increase the global availability, reduce the reliance on unstable dependencies and guide ecosystem improvements.
  • Champion SRE principles such as SLIs, SLOs, and error budgets.
  • Advocate for resilient architecture and fault-tolerant design patterns, incorporating AI-assisted design reviews and architecture evaluation.

Collaboration & Communication

  • Work closely with software engineers, DevOps, and product teams to align reliability goals.
  • Document processes, runbooks, and best practices for knowledge sharing.
  • Provide mentorship and guidance on reliability and operational excellence.

What We Look For in a Candidate

Required Qualifications:

  • 5 years overall professional experience in SRE, DevOps, or infrastructure engineering roles.
  • Experience with Terraform, or similar IaC tools to manage Cloud resources.
  • Proficiency in scripting languages (Python, Bash, etc.) and automation frameworks.
  • Experience with CI/CD pipelines and tools like GitHub Actions, Jenkins or GitLab CI.
  • Solid understanding of monitoring and logging tools (e.g., CloudWatch, ELK, Datadog).
  • Familiarity with containerization and orchestration (Docker, Kubernetes).
  • Excellent AI and problem-solving skills, and a proactive mindset.

Preferred Qualifications:

  • Experience in AWS services (EC2, CloudFront, EKS, RDS, S3, etc.).
  • Certifications in AWS or related technologies are a plus.
  • Experience of application development using Java Microservices and Spring Boot framework
  • Experience with Agile/SCRUM Methodologies and development practices

Compensation

This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.

Location Based Pay Ranges

$105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY $111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI $116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA

Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process.

Learn more about Lumen's:

LI-Remote

LI-VK1

Requisition #: 342698

Life at Lumen

Life at Lumen is human and connected, even in a fast moving, AI‑focused organization. We set clear expectations and trust people to meet them. With real support and shared accountability, teams collaborate better, move faster, and deliver meaningful outcomes.

Our Lumen 8 behaviors guide how we interact, make decisions, and work together, shaping a culture built to perform and win.

To learn more about Life at Lumen and how we live the Lumen 8, please visit:

Background Screening

If you are selected for a position, there will be a background screen, which may include checks for criminal records and/or motor vehicle reports and/or drug screening, depending on the position requirements. For more information on these checks, please refer to the Post Offer section of our FAQ page . Job-related concerns identified during the background screening may disqualify you from the new position or your current role. Background results will be evaluated on a case-by-case basis.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Equal Employment Opportunities

We are committed to providing equal employment opportunities to all persons regardless of race, color, ancestry, citizenship, national origin, religion, veteran status, disability, genetic characteristic or information, age, gender, sexual orientation, gender identity, gender expression, marital status, family status, pregnancy, or other legally protected status (collectively, “protected statuses”). We do not tolerate unlawful discrimination in any employment decisions, including recruiting, hiring, compensation, promotion, benefits, discipline, termination, job assignments or training.

Privacy Notice

Lumen is committed to protecting the privacy and security of personal information collected during the recruitment and hiring process. Our Applicant Privacy Notice explains how we collect, use, disclose, and protect applicant information, as well as how individuals may request access to or deletion of their personal data.

To review Lumen’s Global Employment Applicant and Talent Community Privacy Notice, please visit:

Disclaimer

The job responsibilities described above indicate the general nature and level of work performed by employees within this classification. It is not intended to include a comprehensive inventory of all duties and responsibilities for this job. Job duties and responsibilities are subject to change based on evolving business needs and conditions.

In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.

Please be advised that Lumen does not require any form of payment from job applicants during the recruitment process. All legitimate job openings will be posted on our official website or communicated through official company email addresses. If you encounter any job offers that request payment in exchange for employment at Lumen, they are not for employment with us, but may relate to another company with a similar name.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in New York, NY vacancy
  • $158.5k - $172k

     ...velocity energy of a powerhouse startup. As a leading U.S. ordering and delivery marketplace,...  .... About The Opportunity As a Senior Engineer on the Runtime Automation team, you will...  ...high-impact position driving continuous reliability, deep system optimization, and... 
    Suggested
    Full time
    Work at office
    3 days per week

    Wonder

    New York, NY
    21 hours ago
  • $111k - $218k

     ...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency... 
    Suggested
    Full time
    Local area
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    4 days ago
  • $130k - $180k

    About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI...  ...house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration...  ...THE ROLE Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure... 
    Suggested
    Temporary work
    Work at office
    Immediate start
    Remote work

    Nebius

    New York, NY
    1 day ago
  •  ...A dynamic fintech company in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability of its fintech products. The role focuses on end-to-end monitoring of customer journeys, incident management, and collaboration with various teams. Candidates... 
    Suggested

    ClarityPay

    New York, NY
    11 hours ago
  •  ...Karsun Solutions in the DMV area is seeking a Site Reliability Manager to ensure reliability, scalability, and performance of our systems. You will lead a team focusing on Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate has 1... 
    Suggested

    Karsun Solutions

    New York, NY
    11 hours ago
  •  ...Karsun Solutions, LLC is seeking a Site Reliability Manager to lead a multi-disciplinary team responsible for reliability, security, and platform lifecycle across AWS-based services. The role emphasizes collaboration, observability, and continuous improvement in a client... 

    Karsun Solutions

    New York, NY
    11 hours ago
  •  ...researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At...  ...new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform and customer... 

    Modal Labs

    New York, NY
    12 hours ago
  •  ...The Consulting Solutions is seeking an experienced Senior / Staff Engineer for our SRE, InfraSec team in Seattle. The role involves leading the security of cloud-based infrastructure, mentoring a team of SREs, and collaborating with other engineering teams to ensure high... 
    Remote work

    The Consulting Solutions

    New York, NY
    11 hours ago
  •  ...Komodor, a remote-first company, is seeking a Solutions Engineer to connect customer business initiatives to the Komodor platform, understanding developers, DevOps and Incident response teams working with Kubernetes. You will identify customer pain points and communicate... 
    Remote work

    Komodor

    New York, NY
    11 hours ago
  •  ...A financial technology company based in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability and health of their fintech products. The role involves monitoring customer journeys, APIs, and incident management, requiring 5+ years of experience... 

    ClarityPay Program Services, LLC

    New York, NY
    12 hours ago
  •  ...optimize production infrastructure across CI/CD, cloud deployments, and security. You will collaborate with our internal product and engineering teams to keep services scalable, secure, and highly available. The role emphasizes GitHub Actions, Terraform, Vercel, AWS core... 
    Remote work

    Remote Leverage

    New York, NY
    12 hours ago
  •  ...and Antler, we empower CISOs to proactively manage human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our... 
    Full time
    Work at office

    Dune Security

    New York, NY
    4 days ago
  • $70k - $90k

     ...TD SYNNEX is seeking a Business Systems Engineer for its Global eCommerce and ERP team. This position involves translating business requirements into technical specifications for their systems. The ideal candidate has 1-3 years of experience, is knowledgeable in Java,... 
    Remote work

    TD SYNNEX

    New York, NY
    11 hours ago
  • Fun is seeking a Business Development professional to own enterprise deals end-to-end and accelerate commercial growth at the frontier of on-chain payments. This role is primarily in-person at our Midtown, NYC headquarters with a Monday–Thursday collaboration rhythm and...
    Work from home

    Fun Inc

    New York, NY
    11 hours ago
  • $93k - $160k

     ...Palantir Technologies is seeking a Site Reliability Operations Analyst in New York, NY. In this role, you will streamline workflows and reduce friction in deployments. Your responsibilities include supporting deployments, removing roadblocks, and managing multiple challenges... 

    Dormont Manufacturing Company

    New York, NY
    12 hours ago
  • $130k - $170k

     ...NBC Universal is looking for a Staff Software Engineer (SRE Lead) in New York, NY. This role involves overseeing day-to-day operations of SAP BTP CPI applications, managing incidents, leading offshore support teams, and ensuring high system performance. Candidates should... 
    Remote work

    NBCUniversal

    New York, NY
    11 hours ago
  • $160k - $180k

     ...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience... 

    Socure Inc

    New York, NY
    11 hours ago
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly...  ...issues to resolution with very infrequent after‑hours support. Lead blameless postmortems and implement long‑term improvements... 

    Ridge Line Services

    New York, NY
    5 days ago
  • $123k - $165k

     ...Site Reliability Engineer II Our engineering fleet is a horizontal set of teams providing engineering services across the organization. Our specific team provides reliability engineering and operational support to backend service development teams. Technology is... 

    Disney

    New York, NY
    2 days ago
  • $176.75k - $209.1k

     ...Site Reliability Engineer At Peloton, we view Platform as a Product. A phenomenal platform unlocks speed of development and learning. It allows us to scale easily, enabling our engineers to maximize attention on new features and capabilities. A key to crafting a phenomenal... 
    Temporary work

    Peloton

    New York, NY
    2 days ago
  • $150k - $175k

     ...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we're guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed... 
    Remote work

    ASAPP

    New York, NY
    1 day ago
  •  ...self-healing, deployment/rollback automation). Establish reliability standards: SLOs/SLIs, error budgets, production readiness reviews...  ..., and release risk controls. Performance and reliability engineering: capacity planning, load/performance analysis, resilience... 

    Bahwan CyberTek

    New York, NY
    3 days ago
  • $100k - $250k

     ...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial...  ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics... 
    Local area

    Kalshi

    New York, NY
    5 days ago
  •  ...Applications Deployment Responsible for reliability and support of Container Platform on-...  ...Perform blameless RCA, partner with engineering and operation teams across the...  ...Additional Skills : Automation Process Engineer,Site Reliability Engineer,Full Stack DeveloperThis... 

    Kaav Inc.

    New York, NY
    4 days ago
  •  ...exceptional professionals for this role. JOB DESCRIPTION As a Site Reliability Engineering at JPMorgan Chase within the Enterprise technology,...  ...situations with composure and tact. J ob responsibilitie s Lead SRE practices that balance delivery speed, efficiency, and... 

    J.P. Morgan

    New York, NY
    3 days ago
  •  ...strategy sessions with other Optum Teams Require 2+ years of experience with Terraform Require 2+ years of experience with DevOps Solution Architect, DevOps, or System Engineer certification in one or more public cloud providers Terraform certification #J-18808-Ljbffr... 

    TechDigital Group

    New York, NY
    2 days ago
  • $175k - $230k

     ...those residing in senior living facilities. Falls are the leading cause of injury-related death among adults over 65. And yet...  ...be a 24x7, highly available platform for elder care. As a Site Reliability Engineer, you'll partner with engineering teams across the organization... 
    Apprenticeship
    Work at office
    Local area
    Remote work
    2 days per week

    Sage Group plc

    New York, NY
    4 days ago
  •  ...coasts. If you're driven by impact, pace, and raising the bar. This is the place. The Role As a Staff Site Reliability Engineer you'll play a lead role on the founding SRE team at our new NYC engineering hub. You'll own multi-team reliability and... 
    Work at office

    Legora

    New York, NY
    5 days ago
  • $120k - $165k

     ...and shape the future of our communities. This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the...  ...Overview The Wealth Management Production Management Site Reliability Engineer position is a highly visible/... 
    Temporary work
    Work at office

    PowerToFly

    New York, NY
    3 days ago
  • $132.23k - $176.31k

     ...the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This... 
    Full time
    Temporary work
    Remote work

    Lumen

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!