Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$105.79k - $141.05k
Full-time

Lumen

Lumen is the trusted network for the AI‑powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high‑performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities.

At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where performance, security, and resilience matter.

This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you’re ready to take ownership, deliver meaningful impact, and help shape the future of AI‑ready connectivity, join us today.

The Role

We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is critical to ensuring the reliability, scalability, and efficiency of our systems, with a strong emphasis on AWS infrastructure, observability, automation, and AI-assisted engineering practices.

The Lead SRE requires an AI-native mindset, understands the software development lifecycle (from coding to support) and applies modern AI tools to enhance productivity, quality, and operational excellence. This role will shape how Lumen combines the latest technologies, including AI-driven automation, to modernize software delivery and application lifecycle management.

This role will collaborate with key stakeholders across the engineering organization — including product owners, developers, and testers — to design, optimize, and automate business and technical processes, while effectively navigating multiple teams within a large and complex organization.

Location

This role is designated as a fully remote position within the United States.

The Main Responsibilities

Production Support & Incident Management

  • Implement AI systems and automations to assist during ongoing outages and triage potential ones. You will work with the development teams to ensure that they have all the data normally needed during an outage at their fingertips including preliminary analysis by AI.
  • Provide Tier 3 support for issues across portal services by troubleshooting and resolving technical issues in test and production environments.
  • Lead root cause analysis and post-mortem processes, incorporating AI-assisted analysis and pattern detection to ensure continuous improvement.

Performance Optimization

  • Monitor system performance and proactively identify bottlenecks or degradation using AI-driven observability and anomaly detection tools.
  • Implement tuning strategies across application layers, databases, and infrastructure.
  • Drive initiatives to improve latency, throughput, and resource utilization.

Monitoring & Observability

  • Deploy improved alerting for Lumen Connect in depth, focusing on outside in but including early indicators for fulfilment and other areas. Combining traditional monitoring with AI-based anomaly detection and noise reduction.
  • Proactively monitor the errors and performance on Lumen Connect. Implement rules to detect deviations, implement improvements together with the teams.
  • Design and maintain dashboards, alerts, and metrics using tools like Datadog, AppInsights, CloudWatch, or similar.

Automation & Infrastructure as Code

  • Develop and maintain automation scripts and tools for deployment, scaling, and recovery.
  • Develop and maintain automation scripts and tools for deployment, scaling, and recovery, leveraging AI-assisted code generation and validation tools
  • Use Terraform, or similar IaC tools to manage AWS resources.

Reliability Engineering

  • Perform an in-depth analysis of the overall system and its dependencies, implementing techniques to increase the global availability, reduce the reliance on unstable dependencies and guide ecosystem improvements.
  • Champion SRE principles such as SLIs, SLOs, and error budgets.
  • Advocate for resilient architecture and fault-tolerant design patterns, incorporating AI-assisted design reviews and architecture evaluation.

Collaboration & Communication

  • Work closely with software engineers, DevOps, and product teams to align reliability goals.
  • Document processes, runbooks, and best practices for knowledge sharing.
  • Provide mentorship and guidance on reliability and operational excellence.

What We Look For in a Candidate

Required Qualifications:

  • 5 years overall professional experience in SRE, DevOps, or infrastructure engineering roles.
  • Experience with Terraform, or similar IaC tools to manage Cloud resources.
  • Proficiency in scripting languages (Python, Bash, etc.) and automation frameworks.
  • Experience with CI/CD pipelines and tools like GitHub Actions, Jenkins or GitLab CI.
  • Solid understanding of monitoring and logging tools (e.g., CloudWatch, ELK, Datadog).
  • Familiarity with containerization and orchestration (Docker, Kubernetes).
  • Excellent AI and problem-solving skills, and a proactive mindset.

Preferred Qualifications:

  • Experience in AWS services (EC2, CloudFront, EKS, RDS, S3, etc.).
  • Certifications in AWS or related technologies are a plus.
  • Experience of application development using Java Microservices and Spring Boot framework
  • Experience with Agile/SCRUM Methodologies and development practices

Compensation

This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.

Location Based Pay Ranges

$105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY $111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI $116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA

Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process.

Learn more about Lumen's:

LI-Remote

LI-VK1

Requisition #: 342698

Life at Lumen

Life at Lumen is human and connected, even in a fast moving, AI‑focused organization. We set clear expectations and trust people to meet them. With real support and shared accountability, teams collaborate better, move faster, and deliver meaningful outcomes.

Our Lumen 8 behaviors guide how we interact, make decisions, and work together, shaping a culture built to perform and win.

To learn more about Life at Lumen and how we live the Lumen 8, please visit:

Background Screening

If you are selected for a position, there will be a background screen, which may include checks for criminal records and/or motor vehicle reports and/or drug screening, depending on the position requirements. For more information on these checks, please refer to the Post Offer section of our FAQ page . Job-related concerns identified during the background screening may disqualify you from the new position or your current role. Background results will be evaluated on a case-by-case basis.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Equal Employment Opportunities

We are committed to providing equal employment opportunities to all persons regardless of race, color, ancestry, citizenship, national origin, religion, veteran status, disability, genetic characteristic or information, age, gender, sexual orientation, gender identity, gender expression, marital status, family status, pregnancy, or other legally protected status (collectively, “protected statuses”). We do not tolerate unlawful discrimination in any employment decisions, including recruiting, hiring, compensation, promotion, benefits, discipline, termination, job assignments or training.

Privacy Notice

Lumen is committed to protecting the privacy and security of personal information collected during the recruitment and hiring process. Our Applicant Privacy Notice explains how we collect, use, disclose, and protect applicant information, as well as how individuals may request access to or deletion of their personal data.

To review Lumen’s Global Employment Applicant and Talent Community Privacy Notice, please visit:

Disclaimer

The job responsibilities described above indicate the general nature and level of work performed by employees within this classification. It is not intended to include a comprehensive inventory of all duties and responsibilities for this job. Job duties and responsibilities are subject to change based on evolving business needs and conditions.

In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.

Please be advised that Lumen does not require any form of payment from job applicants during the recruitment process. All legitimate job openings will be posted on our official website or communicated through official company email addresses. If you encounter any job offers that request payment in exchange for employment at Lumen, they are not for employment with us, but may relate to another company with a similar name.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Annandale, VA vacancy
  • $80k - $133k

     ...degree, Four (4) years additional experience will be needed. * Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems. * One(1)+ years of experience... 
    Suggested
    Permanent employment
    Contract work
    Temporary work
    Flexible hours

    Guidehouse Careers

    McLean, VA
    3 days ago
  •  ...Capital One is hiring a Manager, SRE Risk Advisory and Oversight to lead technical oversight over software engineering and SRE practices. This role involves conducting risk analyses on cloud implementations and collaborating with leadership to develop strategic recommendations... 
    Suggested

    Information Technology Senior Management Forum

    McLean, VA
    1 day ago
  • $103.5k - $150k

     ...Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and...  ...Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure... 
    Suggested
    Temporary work
    Work experience placement
    Local area
    3 days per week

    Medallia

    McLean, VA
    8 hours ago
  • $125k - $135k

    Site Reliability Engineer Job number: 880 This is a remote position. Ad Hoc is a technology company that empowers organizations to deliver scalable, impactful digital services. Using modern, agile methods, our team creates products that meet people's needs and... 
    Suggested
    Remote work
    Flexible hours

    Ad Hoc LLC

    McLean, VA
    5 days ago
  •  ...Site Reliability Engineer II Join the leader in providing smarter solutions for a safer world. The property technology space is growing rapidly, and Kastle Systems is leading the way. Kastle Systems is the leader in managed security, with a track record of introducing... 
    Suggested
    Remote work

    Kastle Systems

    Falls Church, VA
    3 days ago
  •  ...No clearance needed / 100% remote within the US  Staff Site Reliability Engineer / Cloud SME Location: 100% remote in the continental US...  ...Responsibilities Architecture & Transformation Leadership Lead the technical rearchitecting efforts, transforming a large-... 
    Long term contract
    Remote work

    ASCENDING LLC

    Fairfax, VA
    5 days ago
  • $132.23k - $176.31k

     ...the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This... 
    Full time
    Temporary work
    Remote work

    Lumen

    Annandale, VA
    2 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA....  ...Remote unavailable. Modality: On‑Site only. Must live within...  ...this role, you will take the lead on designing, deploying, and...  ...scalability, performance, and reliability across environments. What You... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    10 hours ago
  •  ...Communication : Excellent communicator. Expected to actively lead and triage proactively identified issues/incidents where...  ...a new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle... 
    Full time
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    10 hours ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology...  ...address performance and availability issues proactively. Lead the strategic effort to eliminate toil, identifying and... 
    Remote work

    Noctua Technology

    Washington DC
    3 days ago
  • Arena Technical Resources, LLC (ATR) is looking for a ServiceNow Developer to design, develop, and customize ServiceNow solutions. This role is fully remote and involves collaborating with product owners to build scalable IT solutions while ensuring compliance with ITIL...
    Remote work

    Arena Technical Resources

    Falls Church, VA
    4 days ago
  • $17 - $27.75 per hour

     ...deliver an exceptional customer experience * Serves as a Brand Ambassador embodying of Coach values and increasing brand awareness * Leads implementation of Company initiatives and support full operation of the business * Maintain a growth mindset for business and... 
    Minimum wage
    Shift work

    Tapestry

    McLean, VA
    5 days ago
  •  ...A leading security infrastructure firm in Washington, D.C. is seeking a hands-on Site Reliability Engineer (SRE) with expertise in Kubernetes and cloud infrastructure. The role emphasizes total ownership of security infrastructure while defending against advanced threats... 

    Cyrad Solutions LLC

    Washington DC
    1 day ago
  • $81.1k - $187k

     ...Infrastructure Engineer Takes proactive steps to design and architect...  ...and service to ensure reliability and functionality. Forecasts...  ...impact and develops knowledge of site reliability trends. Key...  ...your potential at a company leading the way in AI and cloud solutions... 
    Temporary work
    Flexible hours

    Oracle

    Vienna, VA
    2 days ago
  •  ...Detail Description: The AWS Site Reliability Engineer (SRE) is responsible for the operational health, availability, and performance of the AWS and Databricks environments built by the Platform Engineering team. You prepare and take ownership of "day two" operations... 

    InstantServe LLC

    Vienna, VA
    3 days ago
  • $106.3k - $221.1k

     ...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key... 
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    2 days ago
  •  ...including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with Enterprise...  ...Cloud journey and as a member of our team help lead Software automation and reliability for our platform... 
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    2 days ago
  • $95k - $171k

     ...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:...  ...Akamai powers and protects life online. Leading companies worldwide choose Akamai to... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    5 days ago
  • $165k - $230k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    2 days ago
  •  ...hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical...  ...technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident response in a constrained environment... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Twenty Inc.

    Arlington, VA
    20 hours ago
  •  ...Courtyard Tysons Corner Fairfax is seeking a Housekeeping Supervisor to lead and train a team of room attendants, housepersons, and lobby attendants. You will inspect performance, ensure high standards, and drive guest satisfaction while maintaining safety and productivity... 

    Courtyard Tysons Corner Fairfax

    McLean, VA
    4 days ago
  • $121.4k - $218.6k

     ...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages...  ...Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    3 days ago
  • $153k - $185k

     ...Senior Site Reliability Engineer El Segundo, California, United States About Varda Low Earth orbit is open for business. Varda is accelerating the development of commercial space infrastructure, from in-orbit pharmaceutical processing to reliable and economical... 
    Permanent employment
    Full time
    Immediate start
    Relocation package
    Flexible hours
    Weekend work

    Varda Space Industries

    Washington DC
    3 days ago
  •  ...Join to apply for the Lead Site Reliability Engineer role at Bridge Defense About Bridge Defense Bridge Defense is redefining how modern defense technology is delivered. Based in Washington, D.C., we are built for the dynamic mission environment facing the Department... 
    Full time
    Contract work
    Remote work
    Relocation

    Bridge Defense

    Washington DC
    3 days ago
  • $51.9 per hour

     ...This job is responsible for the reliability, availability, and...  ...efficiency. This role blends software engineering, clinical engineering, and...  ...cross-functionally with AHN site leaders and teams to navigate...  ...drills and exercises, as needed. Leads or participates in post-... 
    For contractors
    Local area

    Highmark Health

    Washington DC
    1 day ago
  •  ...SRE/DevOps Engineer Location: McLean, VA (5 Days mandatory) - Only locals/nearby F2F interview mandatory Developing appropriate DevOps channels throughout the organization. Evaluating, implementing and streamlining DevOps practices. Establishing a continuous... 
    Local area

    E-Solutions

    McLean, VA
    2 days ago
  • $131k - $164k

     ...Staff Site Reliability Engineer New York, New York, United States Position Overview We are seeking a highly skilled Staff Site Reliability...  ...they need to drive greater impact and accountability – to lead with purpose. Our employees are passionate, smart, and... 
    Work at office
    Local area
    Flexible hours

    Diligent

    Washington DC
    3 days ago
  •  ...candidate to join our talented Team. Job Title: SRE / DevOps Engineer Job Location: Mclean, VA Duration: 3-month...  ...possibility of extension Job Description: We are seeking a Site Reliability Engineer (SRE) with strong expertise in the client ecosystem... 

    Ampcus

    McLean, VA
    5 days ago
  •  ...WORK This senior role fosters collaboration with other senior engineers for the development of advanced data analytics solutions and...  ...mission objectives. WHO WE ARE At Lockheed Martin, we're a leading aerospace and defense company that's shaping the future of... 

    Lockheed Martin

    McLean, VA
    2 days ago
  • $99k - $225k

    Agentic AI Engineer The Opportunity: As an experienced engineer,...  ...will collaborate with product leads, solution architects, and engineers...  ...balancing solution quality, reliability, and cost. You will be part...  ...Resource page on our Careers site and reviewing Our Employee... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work
    Shift work

    Booz Allen Hamilton

    McLean, VA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!