Lead Site Reliability Engineer
$105.79k - $141.05kLumen
Lumen is the trusted network for the AI‑powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high‑performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities.
At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where performance, security, and resilience matter.
This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you’re ready to take ownership, deliver meaningful impact, and help shape the future of AI‑ready connectivity, join us today.
The Role
We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is critical to ensuring the reliability, scalability, and efficiency of our systems, with a strong emphasis on AWS infrastructure, observability, automation, and AI-assisted engineering practices.
The Lead SRE requires an AI-native mindset, understands the software development lifecycle (from coding to support) and applies modern AI tools to enhance productivity, quality, and operational excellence. This role will shape how Lumen combines the latest technologies, including AI-driven automation, to modernize software delivery and application lifecycle management.
This role will collaborate with key stakeholders across the engineering organization — including product owners, developers, and testers — to design, optimize, and automate business and technical processes, while effectively navigating multiple teams within a large and complex organization.
Location
This role is designated as a fully remote position within the United States.
The Main Responsibilities
Production Support & Incident Management
- Implement AI systems and automations to assist during ongoing outages and triage potential ones. You will work with the development teams to ensure that they have all the data normally needed during an outage at their fingertips including preliminary analysis by AI.
- Provide Tier 3 support for issues across portal services by troubleshooting and resolving technical issues in test and production environments.
- Lead root cause analysis and post-mortem processes, incorporating AI-assisted analysis and pattern detection to ensure continuous improvement.
Performance Optimization
- Monitor system performance and proactively identify bottlenecks or degradation using AI-driven observability and anomaly detection tools.
- Implement tuning strategies across application layers, databases, and infrastructure.
- Drive initiatives to improve latency, throughput, and resource utilization.
Monitoring & Observability
- Deploy improved alerting for Lumen Connect in depth, focusing on outside in but including early indicators for fulfilment and other areas. Combining traditional monitoring with AI-based anomaly detection and noise reduction.
- Proactively monitor the errors and performance on Lumen Connect. Implement rules to detect deviations, implement improvements together with the teams.
- Design and maintain dashboards, alerts, and metrics using tools like Datadog, AppInsights, CloudWatch, or similar.
Automation & Infrastructure as Code
- Develop and maintain automation scripts and tools for deployment, scaling, and recovery.
- Develop and maintain automation scripts and tools for deployment, scaling, and recovery, leveraging AI-assisted code generation and validation tools
- Use Terraform, or similar IaC tools to manage AWS resources.
Reliability Engineering
- Perform an in-depth analysis of the overall system and its dependencies, implementing techniques to increase the global availability, reduce the reliance on unstable dependencies and guide ecosystem improvements.
- Champion SRE principles such as SLIs, SLOs, and error budgets.
- Advocate for resilient architecture and fault-tolerant design patterns, incorporating AI-assisted design reviews and architecture evaluation.
Collaboration & Communication
- Work closely with software engineers, DevOps, and product teams to align reliability goals.
- Document processes, runbooks, and best practices for knowledge sharing.
- Provide mentorship and guidance on reliability and operational excellence.
What We Look For in a Candidate
Required Qualifications:
- 5 years overall professional experience in SRE, DevOps, or infrastructure engineering roles.
- Experience with Terraform, or similar IaC tools to manage Cloud resources.
- Proficiency in scripting languages (Python, Bash, etc.) and automation frameworks.
- Experience with CI/CD pipelines and tools like GitHub Actions, Jenkins or GitLab CI.
- Solid understanding of monitoring and logging tools (e.g., CloudWatch, ELK, Datadog).
- Familiarity with containerization and orchestration (Docker, Kubernetes).
- Excellent AI and problem-solving skills, and a proactive mindset.
Preferred Qualifications:
- Experience in AWS services (EC2, CloudFront, EKS, RDS, S3, etc.).
- Certifications in AWS or related technologies are a plus.
- Experience of application development using Java Microservices and Spring Boot framework
- Experience with Agile/SCRUM Methodologies and development practices
Compensation
This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.
Location Based Pay Ranges
$105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY $111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI $116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA
Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process.
Learn more about Lumen's:
LI-Remote
LI-VK1
Requisition #: 342698
Life at Lumen
Life at Lumen is human and connected, even in a fast moving, AI‑focused organization. We set clear expectations and trust people to meet them. With real support and shared accountability, teams collaborate better, move faster, and deliver meaningful outcomes.
Our Lumen 8 behaviors guide how we interact, make decisions, and work together, shaping a culture built to perform and win.
To learn more about Life at Lumen and how we live the Lumen 8, please visit:
Background Screening
If you are selected for a position, there will be a background screen, which may include checks for criminal records and/or motor vehicle reports and/or drug screening, depending on the position requirements. For more information on these checks, please refer to the Post Offer section of our FAQ page . Job-related concerns identified during the background screening may disqualify you from the new position or your current role. Background results will be evaluated on a case-by-case basis.
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
Equal Employment Opportunities
We are committed to providing equal employment opportunities to all persons regardless of race, color, ancestry, citizenship, national origin, religion, veteran status, disability, genetic characteristic or information, age, gender, sexual orientation, gender identity, gender expression, marital status, family status, pregnancy, or other legally protected status (collectively, “protected statuses”). We do not tolerate unlawful discrimination in any employment decisions, including recruiting, hiring, compensation, promotion, benefits, discipline, termination, job assignments or training.
Privacy Notice
Lumen is committed to protecting the privacy and security of personal information collected during the recruitment and hiring process. Our Applicant Privacy Notice explains how we collect, use, disclose, and protect applicant information, as well as how individuals may request access to or deletion of their personal data.
To review Lumen’s Global Employment Applicant and Talent Community Privacy Notice, please visit:
Disclaimer
The job responsibilities described above indicate the general nature and level of work performed by employees within this classification. It is not intended to include a comprehensive inventory of all duties and responsibilities for this job. Job duties and responsibilities are subject to change based on evolving business needs and conditions.
In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.
Please be advised that Lumen does not require any form of payment from job applicants during the recruitment process. All legitimate job openings will be posted on our official website or communicated through official company email addresses. If you encounter any job offers that request payment in exchange for employment at Lumen, they are not for employment with us, but may relate to another company with a similar name.
$158.5k - $172k
...velocity energy of a powerhouse startup. As a leading U.S. ordering and delivery marketplace,... .... About The Opportunity As a Senior Engineer on the Runtime Automation team, you will... ...high-impact position driving continuous reliability, deep system optimization, and...SuggestedFull timeWork at office3 days per week- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex...SuggestedFull timeLocal area
- Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions. As a Site Reliability Engineer II at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will use technology...SuggestedFull timeLocal area
$250k - $350k
...and India, where quantitative researchers, engineers, traders, and operational teams work together... ...that boost stability, throughput, and reliability Qualifications Minimum of 3 years’ experience in production support, site reliability, or infrastructure operations in...SuggestedFull time$164.6k - $288k
...Overview The SRE Community of Practice (CoP) Senior Implementation Lead is responsible for driving the adoption, standardization, and maturity of Site Reliability Engineering (SRE) practices across the organization. This role serves as a key enabler in scaling SRE principles...SuggestedVisa sponsorshipWork visa- ...building and running systems that must perform reliably under real-time market conditions. The culture is highly collaborative, engineering-driven, and focused on continuous... ...a related field 3+ years of experience in site reliability, systems engineering, or technical...
$190.8k - $267.1k
...unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability...Work experience placementHome officeFlexible hours$125.04k - $187.56k
...part of the U.S. family of brands, which also includes five leading omnichannel grocery brands – Food Lion, Giant Food, The... ...Digital and E-commerce, Technology and more. Overview The Site Reliability Engineer (SRE) III is responsible for ensuring the scalability, reliability...Full timeWork at officeRemote workFlexible hours- ...Senior Manager for Edge Operations/SRE in Chicago. This pivotal role involves leading edge infrastructure operations, collaborating across teams to ensure high availability and reliability of the platform. Candidates should have 10+ years in infrastructure roles, strong...
- ...commission eligible.CCC Intelligent Solutions Inc. (CCC) is a leading cloud platform for the multi-trillion-dollar insurance... ...more about CCC at **The Role**We are seeking a talented Sr. Site Reliability Engineering Developer to be part of the fast moving, innovative CCC...Night shift
$150k - $200k
...message the job poster from Selby Jennings Recruitment Consultant @ Selby Jennings | Financial Technology We are seeking a Site Reliability Engineer to join our team and assist with the design, development, and administration of our trading and research systems. This...Full timeWork at office$150k - $155k
...Site Reliability Engineer Hybrid (3 days onsite, 2 days remote) full‑time. No visa sponsorship. Base pay: $150,000 – $155,000 per year, subject to skills and experience. A prestigious company seeks a Site Reliability Engineer focused on observation, logging, and capacity...Full timeWork experience placementRemote workVisa sponsorship$130k - $170k
...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end software... ...in suitability and risk management with industry‑leading education and the latest technology, Supernova enables advisors...Full timeFlexible hoursShift work- ...A leading quantitative trading firm is seeking a Head of Site Reliability Engineering to help scale one of its most critical infrastructure organisations. You will work closely with senior engineering leadership before taking responsibility for a team of SREs. This is...Immediate start
- ...Okta is seeking a Staff Site Reliability Engineer to join the TCore team in Chicago. You will design scalable network solutions, maintain a highly available cloud edge for the identity platform, and automate infrastructure with Terraform and Chef. You will analyze data...
- ...on this job and more exclusive features. Direct message the job poster from Algo Capital Group Senior Site Reliability Engineer - Observability and Automation A leading high-frequency trading firm is seeking a mid to senior-level Site Reliability Engineer with deep...Full timeWork at officeFlexible hours
- ...Overview: Senior Site Reliability Engineer (SRE) Location: Chicago, IL (Onsite) Type: Contract Role Overview: We are seeking a Senior Site Reliability Engineer (SRE) with strong expertise in AWS infrastructure, automation, observability, and production...Contract work
$86k - $105k
...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best... ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$114k - $155k
...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and...Local area- ...Qualifications: 8+ years of software engineering experience, or equivalent demonstrated through... ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform... ...vendor resources. Willingness to work on-site at stated location in the job opening....For contractorsWork experience placement
$112.5k - $187.5k
...TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering... ...Operating with full autonomy, you will drive reliability strategy, lead high-risk technical initiatives, and set the engineering...Full timeTemporary workWork experience placementWork at officeFlexible hours2 days per week$128.5k - $214.1k
...We're looking for a Staff Site Reliability Engineer to join our team, focusing on the core systems that power global financial markets. This isn... ...deploy and maintain critical applications.* **Innovate** **and lead efforts** to prevent incidents, enhance operational...Work at officeWorldwide2 days per week$132.23k - $176.31k
...the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This...Full timeTemporary workRemote work$91.2k - $136.8k
Reliability Engineer - IE08GE We’re determined to make a difference and are proud to be an insurance... ...position will play a crucial role to lead infrastructure resilience in ensuring the... ...in Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps...Full timeTemporary workWork at office3 days per week$17 - $27.75 per hour
...deliver an exceptional customer experience * Serves as a Brand Ambassador embodying of Coach values and increasing brand awareness * Leads implementation of Company initiatives and support full operation of the business * Maintain a growth mindset for business and...Minimum wageShift work$22 - $25 per hour
...sourcing, sustainable practices, and philanthropic initiatives that reflect our values and long-term vision. The Role: The Floor Lead plays a pivotal role on our store leadership team, driving the success of the store by upholding exceptional customer service standards...Hourly payPart timeWorldwideShift workWeekend workAfternoon shift$77.6k - $106.7k
JOB DESCRIPTION Are You Ready to Make It Happen at Mondelēz International? Join our Mission to Lead the Future of Snacking. Make It With Pride. In this leadership role, you are an internal technical master at the line level who ensures the line team is performing...Full timeRelocation packageShift workWeekend work$19 per hour
We are looking for a reliable and experienced Lead to ensure all facility operations follow policies and procedures. They coordinate daily operations... ...the world’s largest providers of integrated facility, engineering, and infrastructure solutions. Every day, our over 100,00...Hourly payFull timeLocal areaShift work$24.95 - $26.75 per hour
Phlebotomist III Site Lead - Oak Lawn, IL, Monday to Friday 8:00 AM to 6:30PM, with rotational weekends 7:00 AM to 2:00 PM Pay range: $24.95 - $26.75 / hour Salary offers are based on a wide range of factors including relevant skills, training, experience, education...Full timePart timeWork at officeMonday to FridayFlexible hoursWeekend work$125k - $205k
...awareness, engagement, and revenue. Trusted by leading enterprise brands including Nike,... ...later.com. About the role: The Senior Engineer is responsible for driving large-scale projects... ...debugging, troubleshooting, and system reliability. Team / Collaboration Grow as a...Permanent employmentFull timeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Chicago, IL
- lead operating engineer Chicago, IL
- site reliability engineer Chicago, IL
- site reliability engineer sre Chicago, IL
- construction site safety Chicago, IL
- site recruiter Chicago, IL
- on site coordinator Chicago, IL
- website content developer Chicago, IL
- website coordinator Chicago, IL
- on-site clinical research associate (traveling/remote) Chicago, IL


