Site Reliability Engineer III
$118.4k - $186.85kJLABHCM20
At Jefferson Lab, you’ll champion cutting-edge science and operational excellence while shaping the future of discovery. Join us and make your mark – where excellence meets purpose, and great minds truly matter.
The good-faith pay range for this role is $118,400 - $186,850 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.
What your job will be like:
As Lead Site Reliability Engineer on the High Performance Data Facility (HPDF) team, you will play a critical role in establishing and running the reliability practice for the facility's first systems on its path to operations. This is a technical role with manager responsibilities: you will supervise and develop a small team of site reliability engineers, and you will also design and build systems yourself, hands on in the code, the monitoring stack, and the incident response. You will design how the facility stays available and recovers, define and report on the service level objectives that measure how well it serves its users, serve as incident commander for significant incidents, and work day to day with staff at both Jefferson Lab and Berkeley Lab. HPDF is still in design, so there is room in this role to grow into influencing the technology choices the facility is built on. The users you support are research physicists and computational scientists, and helping them succeed is a core measure of this role.
In this job you will:
- Lead the design, implementation, and operation of monitoring, logging, alerting, and diagnostic tooling for HPDF compute, storage, network, and facility systems, contributing directly to that work as well as directing it.
- Supervise, mentor, and develop a small team of site reliability engineers: assign and review work, set expectations, give regular feedback, support technical growth, and plan and estimate the multi-person efforts assigned to the team.
- Establish and maintain the facility's operational framework, including on-call and escalation structure, incident management, change management, and scheduled maintenance, and keep operational records, runbooks, and documentation current.
- Design the facility's resilience model, including failure domain isolation, redundancy, graceful degradation, and disaster recovery objectives for a geographically distributed facility, and validate that design through testing.
- Define, implement, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in collaboration with the architecture team and scientific stakeholders, and hold facility operations to them.
- Serve as incident commander for significant incidents, own the postmortem process, and drive root cause prevention back into the design and operation of the systems.
- Drive reliability improvement through automation, process optimization, and the elimination of manual operations, using Python, Go, or shell and standard software development practices.
- Partner with the architecture team on HPDF technology selection from a reliability standpoint, lead evaluations of vendor and open source technologies, and represent HPDF site reliability engineering in the Berkeley Lab partnership.
Additional Responsibilities
- Participate in an on-call rotation as the facility moves toward operations.
Lead - Supervisory - Management
- Supervises a team of site reliability engineers
- Assigns and reviews work, sets performance expectations, provides regular feedback, conducts performance discussions, and supports the technical development of the team.
- Participates in hiring for the group. Does not hold fiscal or budget authority.
Experience
- Required: 10 or more years experience in Site Reliability Engineering, DevOps, systems engineering, or operations engineering, including at least two years leading or supervising engineers. Technical leadership of engineering teams or projects qualifies.
- Preferred: Supporting scientific computing, HPC, or research environments.
- Preferred: Establishing operational practice in a new or greenfield facility.
- Preferred: High availability or around the clock operations.
- Preferred: Serving as the reliability or availability authority during the design phase of a large system or facility, before it entered operations.
- Preferred: Evaluating vendor compute, storage, and network solutions against reliability requirements, including acceptance criteria and benchmarking.
- Preferred: Experience with containers and Kubernetes
- Preferred: Experience with configuration management and infrastructure as code tools (for example Ansible, Terraform, Puppet).
- Preferred: Experience with storage systems, data movement, or large scale data infrastructure.
- Preferred: Experience with IT service management practice and tooling (for example ServiceNow, ITIL).
- Preferred: Experience with HPC infrastructure and environments.
- Preferred: Supporting formal project milestone or gate reviews, such as DOE critical decision reviews, and defining KPPs or acceptance criteria.
Education
- Required: Bachelor's Degree Computer Science or Related Field
- Preferred: Master's Degree Computer Science or Related Field
Experience and Education Exchange
Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.
Knowledge, Skills, and Abilities
- Deep Linux systems expertise, with the ability to troubleshoot across the application, operating system, storage, and network layers and to guide others in doing so.
- Expertise in designing and operating monitoring and observability stacks (for example Prometheus, Grafana, ELK, OpenTelemetry) and in defining and reporting on SLOs and SLIs.
- Strong scripting and automation skills (Python, Go, or shell) with standard software development practices, together with the judgment to decide what is worth automating.
- Demonstrated ability to lead and mentor technical staff: setting expectations, assigning work, giving constructive feedback, and addressing performance issues in a constructive manner.
- Ability to establish and run incident response and operational process in a production environment.
- Ability to design for resilience, including failure domain isolation, redundancy, graceful degradation, and recovery objectives, and to validate the design through testing and analysis.
- Clear written and verbal communication, including the ability to present options, argue persuasively for proposals, and work productively with a community of scientific users and with colleagues across both laboratories.
- Familiarity with public cloud environments (AWS, Azure, GCP).
- Networking at scale: IPv4/IPv6, DNS, firewalls and access control lists, high speed interconnects, and data transfer protocols.
- Ability to review system and vendor designs from a reliability standpoint and to argue a technical position persuasively with architects, vendors, and scientific stakeholders.
- Load testing, performance analysis, and capacity modeling to validate design assumptions and identify bottlenecks in the data path.
- Ability to estimate cost and effort for multi-person projects and to plan staffing accordingly.
- Practical experience developing and deploying AI assisted or autonomous automation for operational work, and routine use of AI tools in day to day engineering.
About Jefferson Lab
Join a community with a common purpose of solving the most challenging scientific and engineering problems of our time. The Jefferson Lab campus is located in southeastern Virginia amidst a vibrant and growing technology community.
A career at Jefferson Lab is more than a job. You will be part of “big science” and work alongside top scientists and engineers from around the world unlocking the secrets of our visible universe. Managed by SURATech, LLC, Thomas Jefferson National Accelerator Facility is entering an exciting period of mission growth and is seeking new team members ready to apply their skills and passion to have an impact. You could call it work, or you could call it a mission. We call it a challenge. We do things that will change the world.
Total Rewards at Jefferson Lab
At Jefferson Lab, we believe that a comprehensive employee benefits program is an important and meaningful part of the compensation employees receive. Our benefits program includes, but is not limited to:
• Medical, Dental, and Vision Care Plans • Flexible Spending Accounts
• Paid Time-off and Leave Programs (Paid Parental, vacation, holidays, and sick leave)
• 401(k) Plan – 9% Lab Contribution; 100% vested • Flexible Work Arrangements
(Remote & Alternate Work Schedules available)
• Tuition Assistance, Training and Professional Development Programs
• Live near the waterways of the Chesapeake Bay region with access to nearby beaches,
mountains, and all major metropolitan centers on the East Coast
SURATech, LLC manages and operates the Thomas Jefferson National Accelerator Facility (Jefferson Lab). SURATech is an Equal Opportunity Employer.
SURATech is committed to providing reasonable accommodation for people with disabilities (unless doing so will result in an undue hardship). If you need a reasonable accommodation for any part of the employment process, please send an e-mail to View email address on aiapply.co or contact Human Resources by calling View phone number on aiapply.co and selecting option 1 between 8 am – 5 pm EST to provide the nature of your request.
Employment with SURATech is conditional upon DOE approval if at any time during your employment you are participating in a Foreign Government Talent Recruitment Program or Affiliated activity. Generally, such programs/activities include any foreign-state-sponsored attempt to acquire U.S.-funded scientific research through programs run or funded by the government that target scientists, engineers, students, academics, researchers, and entrepreneurs of all nationalities working or educated in the United States. This includes positions or appointments, both domestic and foreign, titled academic, professional, or institutional appointments whether or not remuneration is received and whether full-time, part-time or voluntary.
- ...-Job TitleSoftware Developer III, TRADOC G2LocationFort Eustis... ...the purpose of improving the reliability and maintainability of software... ...software information and engineering requirements which is necessitated... ...may take place at contractor site, Government site, or a...SuggestedFor contractors
- ...honesty, and innovation. The Software Engineer QA/QC will be responsible for designing,... ..., and ensure the overall quality and reliability of the software throughout the development... ...job openings or apply for a job on this site as a result of your disability. You can...SuggestedFull time
$95k - $117k
Job Title: Senior Utilities Reliability EngineerLocation: Newport News, VirginiaSalary: $95K -117K plus bonusJob Summary of the Senior Utilities Reliability Engineer: As a Sr. Utilities Reliability Engineer, you will play a crucial role in optimizing and managing utility...Suggested$95k - $117k
Senior Utilities Reliability EngineerLocation: Newport News, VirginiaCompensation: $95,000 to... ...seeking a Senior Utilities Reliability Engineer to manage and optimize utility systems across... ...standards.Schedule: Full-time, on-site, with travel to other facility locations...SuggestedFull time- ...established and growing capabilities across Intelligence, Analytics, Engineering, Mission Support, and Communications disciplines. Founded in 20... ...develop relationships with users at customer site and surface use case gapsBuild initial dashboard products and assist...SuggestedContract workFor contractors
$95k - $119k
...professionals ranges from skilled trades to project managers, engineers and software developers to solution architects, technical subject... ...medical, prescription drug, dental and vision plan choices, on-site health centers, tele-medicine, wellness resources, employee...Local areaRemote workRelocationRelocation package- ...Inc., headquartered in Newport News, Virginia, seeks a Software Engineer to work at unanticipated location(s) in the U.S. Will develop... ...software training for Mühlbauer’s semiconductor systems at customer’s sites;Provide software support service for customers; andComplete...Remote workWorldwideRelocation
- ...develop models of possible future configurations Create daily test metrics and reporting Occasionally perform other IT systems engineering activities such as requirements, design, installation, operation, sustainment, and support Apply technical principles,...Full timeWork experience placementRemote work
- ...established and growing capabilities across Intelligence, Analytics, Engineering, Mission Support, and Communications disciplines. Founded in 20... ...Proactively develop relationships with users at customer site and surface use case gaps Build initial dashboard products and...Contract workFor contractors
- ResponsibilitiesMake an impactControls Engineer serves as the primary controls and automation... ...(Sys Ops) team at our customer site. The role provides in-depth technical expertise... ...material handling equipment to ensure safe, reliable, and efficient operation of the facility....Permanent employmentLocal areaWorldwideFlexible hours
- ...track financially and for an on-time deploymentCollaborate with engineering and sales teams to design and quote automation systems and... ...customers and junior members with this knowledgeLead multiple on-site project teams throughout the installation of a project to a successful...Remote workWorldwide
- ...We recruit nationally and provide financial relocation assistance. Problem-solving with a purpose. As a Technical Solutions Engineer at Epic, you’ll work on software that impacts 305 million patients around the world. Alongside customer counterparts, you’ll tackle...Work at officeRelocationVisa sponsorshipRelocation package
- ...A leading healthcare software company is seeking a Technical Solutions Engineer to tackle complex problems that impact millions of patients worldwide. The role involves diagnosing issues, identifying solutions, and managing implementations across various locations. Ideal...WorldwideRelocationRelocation package
- ...relocation to the area. We recruit nationally and provide financial relocation assistance. Responsibilities As a Technical Solutions Engineer at Epic, you’ll work on software that impacts 305 million patients around the world. Together with customer counterparts, you’ll...RelocationVisa sponsorshipRelocation package
$100k - $130k
Job Title: Senior Project EngineerLocation: Newport News, VASalary: $100-130K plus bonusJob Summary of the Senior Project Engineer: The Senior Project Engineer will lead a small team on medium to large-scale projects or take charge of a critical portion of a more complex...Contract workFor contractorsWork at office$111k - $190k
OverviewLMI is seeking a Staff Software Engineer to build mission applications for a U.S. Army enterprise data platform. This engineer... ...make sound technical decisions, and turn complex user needs into reliable software.LMI is a new breed of digital solutions provider...Contract work- DescriptionSAIC is seeking a Systems Engineer to join our team in support of the Air Operations Center Weapons System (AOC WS) Falconer Program... ...Directive 8570.01-M for Information Assurance Technician Level III (CompTIA Security+).Experience in Systems Engineering activities...For contractors
$180k - $210k
OverviewThe Subject Matter Expert (SME), Senior Systems Engineer serves as the Government’s primary technical POC, providing overall leadership, technical guidance, and quality assurance.ResponsibilitiesKey Responsibilities:Advisory Leadership: Act as the main technical...Temporary workImmediate startFlexible hoursShift work$121k - $194k
OverviewLMI is seeking a Staff Security & Compliance Engineer to build security and authorization into a U.S. Army enterprise data platform.This is an engineering role focused on making secure delivery repeatable. The successful candidate will lead the platform’s authorization...Contract work- ...Software Engineer The Swift Group is a privately held, mission-driven and employee-focused services and solutions company headquartered in Reston, VA. Our capabilities include Software Development, Engineering & IT, Data Science, Cyber Enablement, Logistics, and Training...
- ...Software Engineer Soliel is an innovative, premium engineering company specializing in providing Enterprise Architecture, Network Design, Engineering, and Operations Support, Software Design and Development, Data Center Design, Deployment and Migration, and Systems...
$69k - $86k
...professionals ranges from skilled trades to project managers, engineers and software developers to solution architects, technical subject... ...medical, prescription drug, dental and vision plan choices, on-site health centers, tele-medicine, wellness resources, employee...Full timeLocal areaRemote workRelocationRelocation packageShift work$101k - $174k
OverviewLMI is building a secure mission data platform for the U.S. Army. We are seeking a Staff Data Engineer to own the design and delivery of the data pipelines and data products that power it.This is a senior individual-contributor role for an engineer who can make...Contract work$130k - $163k
...professionals ranges from skilled trades to project managers, engineers and software developers to solution architects, technical subject... ...medical, prescription drug, dental and vision plan choices, on-site health centers, tele-medicine, wellness resources, employee...Local areaRemote workRelocationRelocation package$99k - $122k
...professionals ranges from skilled trades to project managers, engineers and software developers to solution architects, technical subject... ...medical, prescription drug, dental and vision plan choices, on-site health centers, tele-medicine, wellness resources, employee...Local areaRemote workRelocationRelocation package- ...Mission Systems Engineer – Senior Ft. Eustis - Newport News VA areaSenior multi-disciplinary engineering support for Mission Systems, Platforms, and Technical Integration. Capable of overseeing staff and serving as Key Personnel.Key DutiesGenerate reports, procedures,...
$120k - $160k
..., VA, US Date Posted 2026-08-19 Category Engineering and Sciences Subcategory Systems Engineer... ...TS/SCI Potential for Remote Work ORA_ON_SITE Description SAIC is seeking a Systems Engineer... ...Information Assurance Technician Level III (CompTIA Security+). Experience designing...Full timeFor contractorsRemote workShift work- ...Senior Mission Systems EngineerEssnova is seeking a Senior Mission Systems Engineer to support multidisciplinary engineering, modeling and simulation, systems integration, and testing for U.S. Army aviation research and development.Key Responsibilities:Develop technical...
- Responsibilities Tasks include, but not limited to: • Collaborate with application engineers, create software requirements, and resolve customer issues • Develop code to analyze and optimize composite and metal structures for strength, stability, and manufacturability...
$84k - $101k
...professionals ranges from skilled trades to project managers, engineers and software developers to solution architects, technical subject... ...medical, prescription drug, dental and vision plan choices, on-site health centers, tele-medicine, wellness resources, employee...Full timeLocal areaRemote workRelocationRelocation packageShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer III. Be the first to apply!
- on-site clinical research associate (traveling/remote) Newport News, VA
- site safety Newport News, VA
- junior website developer Newport News, VA
- construction site safety Newport News, VA
- IT site lead Newport News, VA
- historic site Newport News, VA
- site services specialist Newport News, VA
- official site Newport News, VA
- site leader Newport News, VA
- site reliability engineer



