Sr Lead Site Reliability Engineer
$132.23k - $176.31kLumen Inc
Lumen is the trusted network for the AI‑powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high‑performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities.
At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where performance, security, and resilience matter.
This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you’re ready to take ownership, deliver meaningful impact, and help shape the future of AI‑ready connectivity, join us today.
The Role
We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is critical to ensuring the reliability, scalability, and efficiency of our systems, with a strong emphasis on AWS infrastructure, observability, automation and AI-assisted engineering practices.
The Senior Lead SRE requires an AI-native mindset, understands the software development lifecycle (from coding to support) and applies modern AI tools to enhance productivity, quality, and operational excellence. This role will shape how Lumen combines the latest technologies, including AI-driven automation, to modernize software delivery and application lifecycle management.
This role will collaborate with key stakeholders across the engineering organization — including product owners, developers, and testers — to design, optimize, and automate business and technical processes, while effectively navigating multiple teams within a large and complex organization.
Location
This role is designated as a fully remote position within the United States.
The Main Responsibilities
Production Support & Incident Management
Help define and improve the processes and best practices for incident management within the customer facing applications.
Implement AI systems and automations to assist during ongoing outages and triage potential ones. You will work with the development teams to ensure that they have all the data normally needed during an outage at their fingertips including preliminary analysis by AI.
Design and implement improved SRE processes to handle incidents. From proactive analysis, faster response, automatic remediation, AI guided analysis, partially automatic root cause analysis.
Performance Optimization
Monitor system performance and proactively identify bottlenecks or degradation using AI-driven observability and anomaly detection tools.
Implement tuning strategies across application layers, databases, and infrastructure.
Drive initiatives to improve latency, throughput, and resource utilization.
Monitoring & Observability
Deploy improved alerting for Lumen Connect in depth, focusing on outside in but including early indicators for fulfilment and other areas. Combining traditional monitoring with AI-based anomaly detection and noise reduction.
Proactively monitor the errors and performance on Lumen Connect. Implement rules to detect deviations, implement improvements together with the teams.
Design and maintain dashboards, alerts, and metrics using tools like Datadog, AppInsights, CloudWatch, or similar.
Automation & Infrastructure as Code
Develop and maintain automation scripts and tools for deployment, scaling, and recovery, leveraging AI-assisted code generation and validation tools
Use Terraform, or similar IaC tools to manage AWS resources.
Reliability Engineering
Perform an in-depth analysis of the overall system and its dependencies, implementing techniques to increase the global availability, reduce the reliance on unstable dependencies and guide ecosystem improvements.
Champion SRE principles such as SLIs, SLOs, and error budgets.
Advocate for resilient architecture and fault-tolerant design patterns, incorporating AI-assisted design reviews and architecture evaluation.
Help define and improve better SRE processes and lead significant improvements in reliability for the Lumen Connect platform.
Collaboration & Communication
Work closely with software engineers, DevOps, and product teams to align reliability goals.
Document processes, runbooks, and best practices for knowledge sharing.
Provide mentorship and guidance on reliability and operational excellence.
What We Look For in a Candidate
**Required Qualifications: **
10+ years overall professional experience in SRE, DevOps, or infrastructure engineering roles.
Experience with Terraform, or similar IaC tools to manage Cloud resources.
Proficiency in scripting languages (Python, Bash, etc.) and automation frameworks.
Experience with CI/CD pipelines and tools like GitHub Actions, Jenkins or GitLab CI.
Solid understanding of monitoring and logging tools (e.g., CloudWatch, ELK, Datadog).
Familiarity with containerization and orchestration (Docker, Kubernetes).
Excellent AI and problem-solving skills, and a proactive mindset.
Preferred Qualifications:
Experience in AWS services (EC2, CloudFront, EKS, RDS, S3, etc.).
Certifications in AWS or related technologies are a plus.
Experience of application development using Java Microservices and Spring Boot framework
Experience with Agile/SCRUM Methodologies and development practices
Compensation
This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.
Location Based Pay Ranges
$132,232 - $176,310 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY
$138,844 - $185,124 in these states: CO HI MI MN NC NH NV OR RI
$145,456 - $193,940 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA
Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process.
Learn more about Lumen's:
Benefits (
Bonus Structure
#LI-Remote
#LI-VK1
Requisition #: 342707
Life at Lumen
Life at Lumen is human and connected, even in a fast moving, AI‑focused organization. We set clear expectations and trust people to meet them. With real support and shared accountability, teams collaborate better, move faster, and deliver meaningful outcomes.
Our Lumen 8 behaviors guide how we interact, make decisions, and work together, shaping a culture built to perform and win.
To learn more about Life at Lumen and how we live the Lumen 8, please visit:
Background Screening
If you are selected for a position, there will be a background screen, which may include checks for criminal records and/or motor vehicle reports and/or drug screening, depending on the position requirements. For more information on these checks, please refer to the Post Offer section of our FAQ page ( . Job-related concerns identified during the background screening may disqualify you from the new position or your current role. Background results will be evaluated on a case-by-case basis.
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
Equal Employment Opportunities
We are committed to providing equal employment opportunities to all persons regardless of race, color, ancestry, citizenship, national origin, religion, veteran status, disability, genetic characteristic or information, age, gender, sexual orientation, gender identity, gender expression, marital status, family status, pregnancy, or other legally protected status (collectively, “protected statuses”). We do not tolerate unlawful discrimination in any employment decisions, including recruiting, hiring, compensation, promotion, benefits, discipline, termination, job assignments or training.
Privacy Notice
Lumen is committed to protecting the privacy and security of personal information collected during the recruitment and hiring process. Our Applicant Privacy Notice explains how we collect, use, disclose, and protect applicant information, as well as how individuals may request access to or deletion of their personal data.
To review Lumen’s Global Employment Applicant and Talent Community Privacy Notice, please visit:
Disclaimer
The job responsibilities described above indicate the general nature and level of work performed by employees within this classification. It is not intended to include a comprehensive inventory of all duties and responsibilities for this job. Job duties and responsibilities are subject to change based on evolving business needs and conditions.
In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.
Please be advised that Lumen does not require any form of payment from job applicants during the recruitment process. All legitimate job openings will be posted on our official website or communicated through official company email addresses. If you encounter any job offers that request payment in exchange for employment at Lumen, they are not for employment with us, but may relate to another company with a similar name.
$185k - $230k
As a Sr. Site Reliability Engineer (SRE) III, you’ll work as part of a collaborative and high-performing team providing your expertise to deliver... ...solutions to meet defined Service Level Objectives.Experience leading or participating in incident response, root cause analysis...SeniorFull timeLocal areaImmediate start$165k - $230k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts...SeniorPermanent employmentTemporary workImmediate startWeekend work- ...automation to reduce manual operational effort. Support capacity planning, performance optimization, reliability, and security initiatives. Collaborate with engineering and infrastructure teams on incident response and migration readiness . Participate in on-...Senior
- ...Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the... ...Management: Act as a primary responder in on-call rotations, leading the technical resolution of production outages....SeniorLocal area
- ...Description: Onsite in Washington, DC our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical... ...alerting to support defined service level objectives. Lead or participate in incident response, root cause analysis,...SeniorHourly payPermanent employmentFull timeLocal areaImmediate start
$125k - $230k
...lasting value. ResponsibilitiesLMI is seeking a skilled Sr. Systems Engineer / Endpoint Lead to manage systems engineering and endpoint device support... ...endpoint support staff across geographically dispersed sites. Support systems integration efforts with network, security...SeniorFull timeContract work$146k - $234k
Responsibilities As Senior Intelligence Team Lead, the candidate will be responsible for the day-to-day management of the contract including staffing, financial management, as well as technical and programmatic reporting. They will be responsible for overseeing employees...SeniorContract workFor subcontractorShift work$139.5k - $188.7k
...ensure the accuracy, completeness, and reliability of Amazon's mandatory sustainability disclosures... ...strategy experience- 5+ years of leading cross-functional initiatives that drive... ...and 7+ years of quantitative role (engineering, process re-engineering, quality assurance...SeniorWorldwideFlexible hours$124.5k - $216.5k
Who We AreFTI Consulting is the leading global expert firm for organizations facing crisis and transformation. We work with many of the... ...The ideal candidate will come with a proven track record of multi-site management and be ready for their next career progression to a...SeniorFull timeWork experience placementWork at officeVisa sponsorshipWork visa$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...SeniorCurrently hiringRemote work$150k - $180k
...are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale... ...organization.This position is based on-site in either our Arlington, VA office,... ...solutions that enhance the team's capabilities.Lead by example in fostering a culture of...SeniorPermanent employmentFull timeWork at officeLocal areaRemote workWorldwide$166k - $220k
...expectations. Our systems integration engineers internalize the nuances of each deployment... ...ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly... ...design and code. They are comfortable leading large, focused projects. They lead in...SeniorFull timeWork experience placementImmediate start$207k - $284.9k
...you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from... ...too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports... ...health of Okta's federal environments, lead a team of engineers working within...SeniorPermanent employmentLocal areaWorldwideFlexible hoursDay shift$106.3k - $221.1k
...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key...SeniorLive inWork at officeLocal area$121.4k - $218.6k
...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure... ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing... ...building technical runbooks, leading complex incident response bridges, and...SeniorWork experience placementWork at office- ...Participate in on-call rotations, respond to critical incidents, lead root cause analysis, and implement preventative measures.... ...AI/ML experience or a strong interest in applying AI/ML to reliability, security, or operational efficiency is a plus. Benefits...SeniorFull timeWork at officeFlexible hours
- ...Position: Senior Site Reliability Engineer (SRE) Location: Redmond WA (Onsite) Duration: Fulltime Job Description We... ...enterprise platform engineering organizations. Experience leading technical execution across multiple engineering teams....SeniorFull time
$147k - $202.4k
...let's talk. Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in... ...rotations and be a primary responder for critical incidents, leading root cause analysis and implementing preventative measures...SeniorWork at officeLocal areaWorldwideFlexible hoursShift work$112k - $218.4k
...: Our team is looking for a Senior Active Directory Site Reliability Engineer. Our mission is to improve the availability, latency, performance and security of the Identity systems behind Microsoft's cloud. Like traditional operations, we keep important revenue-critical...SeniorFull timeLocal area- ...Apogee Global RMS is seeking a Senior Cybersecurity Engineer / Offensive Security Lead to support high‑visibility federal and IC programs. This role is designed for operators who bring hands‑on offensive tradecraft, current certifications, and recent red‑team experience...SeniorFull time
$160k - $210k
...change and achieving remarkable growth in a rapidly evolving industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale...SeniorWork at officeImmediate startRemote workWork from home- ...RMF and SAA requirements, working with system owners, ISSOs, ISSEs, assessors, and Authorizing Officials. Responsibilities include leading the SAA/ATO process, guiding on baseline controls, and coordinating assessment activities to deliver complete SAA/ATO packages and...Senior
- ...celebrating 26 years ofexcellence—is seeking Senior LPD Financial Team Lead toimmediately support our PMS 377 Team at the Naval Sea Systems... .... Our diverse team of acquisition experts, financial analysts, engineers, logisticians, IT professionals, and technical specialists is...SeniorFor contractors
$117.2k - $313.7k
...up your career at the company leading workforce transformation in... ...Distributed Systems Software Engineer - Public Cloud (Senior/Lead/Principal... ...on our platform to be highly reliable, lightning fast, supremely... ...experience balancing live-site management, feature delivery,...SeniorFull time$50.23k
...Environmental Service has multiple openings for Inspectors at the Senior to Lead level within our Technical and Environmental Services Group (TES)... ...to ensure compliance. Key responsibilities include completion of site visits or progress meetings, review work for compliance with...SeniorFull timeFor contractorsWork at officeNight shiftWeekend work$80k - $161k
Washington, DCAdministrative and Logistics Support /Full Time On-Site /On-siteLegislative/Legal Research Lead (LRL) - Senior Work Location: Washington, DC Employment Type: Full-Time, Senior-Level Department: Administrative and Logistics Support CGS is seeking a skilled...SeniorFull time$112k - $179k
...About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in... ...extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise...Contract workWorldwideShift work$125k - $185k
Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions and operations. By bringing the... ..., and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate...Full timeWork experience placementWork at officeRemote workWork from homeRelocation package- OB SUMMARYThe Systems Engineer - Site Reliability Engineering (SRE) is responsible for the reliability, scalability, and performance of mission-critical... ...improvement and service restoration Monitors, manages and leads Service Provider outcomes required to ensure operational...Full timeFor contractorsWork at officeRemote workFlexible hoursShift work
$115.5k - $164.8k
...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant... ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations...Work experience placementWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr Lead Site Reliability Engineer. Be the first to apply!
- lead network engineer Washington DC
- lead infrastructure engineer Washington DC
- lead operating engineer Washington DC
- lead engineer Washington DC
- site reliability engineer Washington DC
- site reliability engineer remote Washington DC
- site reliability engineer sre Washington DC
- senior technical analyst Washington DC
- senior associate attorney Washington DC
- senior developer Washington DC



