Lead Site Reliability Engineer
$105.79k - $141.05kLumen
Lumen is the trusted network for the AI‑powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high‑performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities.
At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where performance, security, and resilience matter.
This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you’re ready to take ownership, deliver meaningful impact, and help shape the future of AI‑ready connectivity, join us today.
The Role
We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is critical to ensuring the reliability, scalability, and efficiency of our systems, with a strong emphasis on AWS infrastructure, observability, automation, and AI-assisted engineering practices.
The Lead SRE requires an AI-native mindset, understands the software development lifecycle (from coding to support) and applies modern AI tools to enhance productivity, quality, and operational excellence. This role will shape how Lumen combines the latest technologies, including AI-driven automation, to modernize software delivery and application lifecycle management.
This role will collaborate with key stakeholders across the engineering organization — including product owners, developers, and testers — to design, optimize, and automate business and technical processes, while effectively navigating multiple teams within a large and complex organization.
Location
This role is designated as a fully remote position within the United States.
The Main Responsibilities
Production Support & Incident Management
- Implement AI systems and automations to assist during ongoing outages and triage potential ones. You will work with the development teams to ensure that they have all the data normally needed during an outage at their fingertips including preliminary analysis by AI.
- Provide Tier 3 support for issues across portal services by troubleshooting and resolving technical issues in test and production environments.
- Lead root cause analysis and post-mortem processes, incorporating AI-assisted analysis and pattern detection to ensure continuous improvement.
Performance Optimization
- Monitor system performance and proactively identify bottlenecks or degradation using AI-driven observability and anomaly detection tools.
- Implement tuning strategies across application layers, databases, and infrastructure.
- Drive initiatives to improve latency, throughput, and resource utilization.
Monitoring & Observability
- Deploy improved alerting for Lumen Connect in depth, focusing on outside in but including early indicators for fulfilment and other areas. Combining traditional monitoring with AI-based anomaly detection and noise reduction.
- Proactively monitor the errors and performance on Lumen Connect. Implement rules to detect deviations, implement improvements together with the teams.
- Design and maintain dashboards, alerts, and metrics using tools like Datadog, AppInsights, CloudWatch, or similar.
Automation & Infrastructure as Code
- Develop and maintain automation scripts and tools for deployment, scaling, and recovery.
- Develop and maintain automation scripts and tools for deployment, scaling, and recovery, leveraging AI-assisted code generation and validation tools
- Use Terraform, or similar IaC tools to manage AWS resources.
Reliability Engineering
- Perform an in-depth analysis of the overall system and its dependencies, implementing techniques to increase the global availability, reduce the reliance on unstable dependencies and guide ecosystem improvements.
- Champion SRE principles such as SLIs, SLOs, and error budgets.
- Advocate for resilient architecture and fault-tolerant design patterns, incorporating AI-assisted design reviews and architecture evaluation.
Collaboration & Communication
- Work closely with software engineers, DevOps, and product teams to align reliability goals.
- Document processes, runbooks, and best practices for knowledge sharing.
- Provide mentorship and guidance on reliability and operational excellence.
What We Look For in a Candidate
Required Qualifications:
- 5 years overall professional experience in SRE, DevOps, or infrastructure engineering roles.
- Experience with Terraform, or similar IaC tools to manage Cloud resources.
- Proficiency in scripting languages (Python, Bash, etc.) and automation frameworks.
- Experience with CI/CD pipelines and tools like GitHub Actions, Jenkins or GitLab CI.
- Solid understanding of monitoring and logging tools (e.g., CloudWatch, ELK, Datadog).
- Familiarity with containerization and orchestration (Docker, Kubernetes).
- Excellent AI and problem-solving skills, and a proactive mindset.
Preferred Qualifications:
- Experience in AWS services (EC2, CloudFront, EKS, RDS, S3, etc.).
- Certifications in AWS or related technologies are a plus.
- Experience of application development using Java Microservices and Spring Boot framework
- Experience with Agile/SCRUM Methodologies and development practices
Compensation
This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.
Location Based Pay Ranges
$105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY $111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI $116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA
Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process.
Learn more about Lumen's:
LI-Remote
LI-VK1
Requisition #: 342698
Life at Lumen
Life at Lumen is human and connected, even in a fast moving, AI‑focused organization. We set clear expectations and trust people to meet them. With real support and shared accountability, teams collaborate better, move faster, and deliver meaningful outcomes.
Our Lumen 8 behaviors guide how we interact, make decisions, and work together, shaping a culture built to perform and win.
To learn more about Life at Lumen and how we live the Lumen 8, please visit:
Background Screening
If you are selected for a position, there will be a background screen, which may include checks for criminal records and/or motor vehicle reports and/or drug screening, depending on the position requirements. For more information on these checks, please refer to the Post Offer section of our FAQ page . Job-related concerns identified during the background screening may disqualify you from the new position or your current role. Background results will be evaluated on a case-by-case basis.
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
Equal Employment Opportunities
We are committed to providing equal employment opportunities to all persons regardless of race, color, ancestry, citizenship, national origin, religion, veteran status, disability, genetic characteristic or information, age, gender, sexual orientation, gender identity, gender expression, marital status, family status, pregnancy, or other legally protected status (collectively, “protected statuses”). We do not tolerate unlawful discrimination in any employment decisions, including recruiting, hiring, compensation, promotion, benefits, discipline, termination, job assignments or training.
Privacy Notice
Lumen is committed to protecting the privacy and security of personal information collected during the recruitment and hiring process. Our Applicant Privacy Notice explains how we collect, use, disclose, and protect applicant information, as well as how individuals may request access to or deletion of their personal data.
To review Lumen’s Global Employment Applicant and Talent Community Privacy Notice, please visit:
Disclaimer
The job responsibilities described above indicate the general nature and level of work performed by employees within this classification. It is not intended to include a comprehensive inventory of all duties and responsibilities for this job. Job duties and responsibilities are subject to change based on evolving business needs and conditions.
In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.
Please be advised that Lumen does not require any form of payment from job applicants during the recruitment process. All legitimate job openings will be posted on our official website or communicated through official company email addresses. If you encounter any job offers that request payment in exchange for employment at Lumen, they are not for employment with us, but may relate to another company with a similar name.
- ...Administrator / SRE in Dallas to own production Java environments, middleware, and cloud automation. You will optimize performance, drive reliability, and mentor teammates while aligning with enterprise security and AI-enabled integrations. You will work across Java apps, IBM...Suggested
- ...The Depository Trust & Clearing Corporation (DTCC) is seeking a Senior Application Support Engineer (SRE) to enhance reliability, scalability, and performance of mission-critical applications. You will apply SRE principles across engineering, infrastructure, and operations...Suggested
- ...Site Reliability Engineer- W2 Role* Technical proficiency: Strong Proficiency in Java, Strong understanding of Database concepts (Oracle, SQL, Dynamo DB etc.) Industry standard SRE Tools like Prometheus, Grafana, Data Dog Etc Good to have skills: Cloud Concepts / AWS,...Suggested
- ...Caterpillar is seeking a Digital Technical Support Analyst to ensure platform stability and cloud service reliability. You will own incidents end-to-end, coordinate across engineering and product teams, and drive improvements in runbooks, monitoring, and incident response. The...Suggested
- ...interview process. Lantern is seeking an experienced Senior Site Reliability Engineer to champion the reliability, availability, and performance... ..., alerting, tracing) using Datadog and Azure Monitor Lead incident management processes using Rootly, including on-call...SuggestedTemporary workFlexible hours
- ...Role: Site Reliability Engineer 6+ months Contract role Remote About the Role We are looking for a dynamic and accomplished Site Reliability Engineer (SRE) who excels at solving complex reliability challenges and thrives in high-impact environments....Contract workRemote work
- ...cloud-native platforms to advanced release engineering practices, our teams are redefining how... ...accelerate development and improve reliability. Your work will directly influence how... ...Scrum teams with demonstrated success leading improvements (getting better/faster/happier...Full timeH1bWork at officeRemote workVisa sponsorshipFlexible hours2 days per week3 days per week
$100k - $115k
...Internal Developer Platform Engineer Analytic Partners is a global... ...is powered by our industry-leading platform and team of experts,... ...customers and optimizing for reliability, usability, and delivery velocity... ...in Platform Engineering, Site Reliability Engineering, DevOps...Temporary work$132.23k - $176.31k
...the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This...Full timeTemporary workRemote work$215k - $355k
About this role Wells Fargo is seeking an AI Solutions Lead to own AI solution delivery within Wealth & Investment Management. This is... ...a single shape-and-build role — a hybrid of a Forward Deployed Engineer and a product leader — who both defines the AI solution and rapid...Full timeWork experience placementWork at office$16 - $17 per hour
Title: Custodial Lead Job Description: Job Overview The Custodial Lead will be responsible for the cleanliness and sanitation of the areas assigned and provides some work direction to custodial staff. Roles & Responsibilities To perform this job successfully and...Hourly payFull timeImmediate startMonday to FridayShift work- ...Pay Rate: $40/Hr. W2 Experience: 3-5 Years Overview We are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in a DevOps‑driven environment...Long term contractContract workInternshipRemote work
$86.6k - $144.4k
...Helix division has an opportunity for a Senior Associate Software Engineer. In this role, you will need to develop and maintain state of... ...in DFW area, the selected candidate may be expected to work on‑site at our Las Colinas office a minimum of two (2) days per week, with...Full timeTemporary workH1bWork at officeRemote workWork from home2 days per week- ...SRE/Devops Engineer Locations can be any of Tampa, Jersey City or Dallas. Below is Detailed... ...Capacity & Performance Optimization: Lead capacity planning and performance analysis to ensure Risk platforms scale reliably under high load. Metrics & Continuous...
- A leading consulting firm is looking for an experienced SAP SuccessFactors Employee Central Senior Consultant to provide high-level support on multiple projects. You will lead requirements definition, business process improvement, and ensure high-quality solutions meet...Remote job
$40 per hour
...A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants...Long term contractInternshipRemote work$231.44k - $416.59k
Anticipated End Date: 2026-08-05 Position Title: Distinguished AI Engineer Job Description: Distinguished AI Engineer Location: This... ...translate business opportunities into scalable AI solutions Lead buy, build, and integrate decisions across commercial and open-source...Full timeTemporary workWork experience placementWork at officeLocal areaDay shift2 days per week1 day per week$55k - $151.47k
...Specialism IFS - Information Technology (IT) Management Level Senior Associate Job Description & Summary The Opportunity As an AI Engineer, you will be at the forefront of transforming raw data into actionable insights, enabling informed decision-making and driving...Work experience placementH1b- ...ONSITE Competitive salary Opportunity for advancement AI Engineer - Agentic AI | LLM | LangGraph | LangChain Location: Dallas,... ...behavior evaluation. * Optimize AI systems for scalability, reliability, security, and cost efficiency. * Collaborate with engineering...Full time
$119k - $224k
About this role: Wells Fargo is seeking a Lead Database Engineer in the Database and Middleware Operations Team under Infrastructure Operations & Reliability (IOR) group. In this role, you will: Being a Lead Database Engineer with deep technical expertise in database...Full timeWork experience placementRelocation- Hours of Work : 7:30a-4p Days Of Week : Monday-Friday Work Shift : Job Description : Your Job: In this highly technical allied imaging professional position, you'll collaborate with a multidisciplinary team to provide the very best imaging services, which...Monday to FridayShift work
- ...Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through... ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform... ...vendor resources Willingness to work on-site at stated location in the job opening...Contract workFor contractorsWork experience placement
- ...Job Description Job Description Forhyre is looking for engineers who can bring unique perspectives and innovative ideas to all areas... ...evangelize cloud best practices while building a culture of reliability and observability Engage in and improve the end to end lifecycle...
$77.4k - $135.4k
...apply strong SQL and Python expertise to enhance performance, reliability, and data accessibility while supporting advanced analytics and... ...solutions. Partner with application developers, data engineers, BI teams, and analytics partners to deliver integrated solutions...Full time$112.71k - $183.14k
...Caterpillar Inc., leveraging the latest technologies to build industry leading digital solutions for our customers and dealers. With over 1.5... ...rebuilds, maintenance, and repairs. As a Senior Software Engineer on the CAT Digital - Global Services Application Team, you will...Full timePart timeCasual workWork at officeWorldwideRelocation packageFlexible hours- Express Virtual Assistant ( Work At Home ) About the job Express Virtual Assistant ( Work At Home ) The Global Advertising and Brand Management (GABM) organization has a mission to create marketplace demand and drive commerce for American Express through differentiated ...Remote jobWork from home
$141.9k - $284.9k
RSM US LLP in Dallas, Texas, is looking for talented individuals to join their team, offering professional services to the middle market globally. The company fosters a culture that empowers both employees and clients, ensuring personal and professional growth. With...Flexible hours- ...About this role: Wells Fargo is seeking a Lead Site Reliability Engineer within the Consumer Technology (CT) organization- providing technology solutions to the LOBs and manage the application portfolios through the enablers of skills, stability, security, scalability...Permanent employmentFull timeWork experience placementRelocation package
$25 - $50 per hour
...Role Overview TSA is accepting applications for Lead and Supervisory Transportation Security Officers at airports in Dallas. These roles are ideal for individuals looking to step into leadership positions within airport security operations. TSA provides training to...Shift workNight shiftWeekend work$128.47k - $192.71k
...Top Candidates Will Have: Typical candidates will have 5+ years of direct project management experience working with software engineering teams in an agile environment Experience managing product backlogs using Agile software tools such as Azure DevOps, Mingle, Team...Full timePart timeRelocationRelocation packageFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Dallas, TX
- lead operating engineer Dallas, TX
- site reliability engineer Dallas, TX
- site reliability engineer sre Dallas, TX
- construction site safety Dallas, TX
- site recruiter Dallas, TX
- on site coordinator Dallas, TX
- website content developer Dallas, TX
- website coordinator Dallas, TX
- on-site clinical research associate (traveling/remote) Dallas, TX




