Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) - II

F5 Networks

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation. Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.Role SummaryWe are seeking a proactive and detail-oriented Site Reliability Engineer II (SRE II) to join our 24/7 Operations team in a hybrid capacity. In this role, you will provide round-the-clock, eyes-on-glass monitoring, proactive incident response, operational maintenance, and continuous compliance support across our hybrid environment—spanning bare-metal on-premises Red Hat Enterprise Linux (RHEL) servers, AWS (Commercial and GovCloud), and Kubernetes (EKS) infrastructure operating under strict FedRAMP standards.As an SRE II, your primary responsibility is maintaining platform availability, hardware reliability, and security posture through real-time telemetry monitoring via Prometheus and Grafana, log troubleshooting and root cause analysis in Kibana and Elasticsearch, rapid incident triage in Slack, execution of automated deployments and GitOps continuous delivery via GitLab CI/CD, ArgoCD, and Argo Workflows, and consistent enforcement of FedRAMP security controls (NIST SP 800-53). You will operate in a structured shift model covering weekdays and weekends to ensure 24/7/365 hybrid platform uptime.24/7 Shift & On-Call ExpectationsThis position requires active participation in a 24/7/365 operational shift model:• Round-the-Clock Coverage: Active "eyes-on-glass" monitoring during scheduled shifts (day, evening, night, and weekend rotations).• Weekend & Holiday Rotation: Scheduled weekend shifts and holiday coverage to ensure continuous operational readiness across both on-premises data centers and cloud regions.• On-Call Escalations: Primary and secondary on-call responsibilities during and outside regular shift windows to meet stringent FedRAMP Incident Response (IR) SLAs.• Real-time Collaboration: Continuous presence in operational Slack channels and ChatOps bridge rooms for instant incident mobilization.Key Responsibilities1. 24/7 Eyes-on Monitoring, Log Troubleshooting & Incident Triage• Maintain continuous monitoring of production telemetry across bare-metal RHEL servers, EKS clusters, and AWS cloud environments.• Act as first responder to high-priority alerts generated by Prometheus and Alertmanager, performing immediate triage and root cause investigation.• Perform deep-dive log troubleshooting using Elasticsearch and Kibana (building queries, analyzing container/application stdout logs, systemd/journald logs, and ingress traffic logs) to isolate error patterns and service failures.• Coordinate incident resolution bridges in Slack, engaging secondary on-call engineers, security teams, or infrastructure engineers as required.• Maintain rigorous adherence to Incident Response SLAs and document timeline logs for post-incident reviews (RCAs).2. Bare-Metal On-Premises RHEL Server Support & Hardware Operations• Provide operational support for bare-metal on-premises Linux servers (RHEL), including OS configuration, system maintenance, and day-to-day lifecycle administration.• Diagnose physical network interface bonding, link aggregation, VLAN tagging, and local server connectivity issues on bare-metal systems.• Coordinate vendor hardware dispatch requests for failing server components and oversee physical parts replacements.3. FedRAMP Security & Vulnerability Remediation• Execute day-to-day security operational duties aligned with FedRAMP High/Moderate (NIST SP 800-53) standards.• Perform routine vulnerability patching (CVE remediation) across bare-metal RHEL servers, cloud AMIs, Kubernetes worker nodes, and container images within strict regulatory timelines.• Apply operational system hardening based on DISA STIG guidelines and maintain FIPS 140 compliance configurations across both cloud and on-premise operating systems.• Ensure strict Role-Based Access Control (RBAC), SSH key management, and security boundary enforcement across operational environments.4. Kubernetes (EKS), AWS Infrastructure & Networking Operations• Support day-to-day operations and node maintenance for Amazon EKS clusters across AWS Commercial and AWS GovCloud environments.• Perform Layer 4 (L4) and Layer 7 (L7) secure load balancing troubleshooting, including AWS Application Load Balancers (ALB), Network Load Balancers (NLB), ingress controllers, SSL/TLS certificate termination, and traffic routing issues.• Perform Linux administration tasks including kernel parameter tuning (sysctl), storage expansion, log rotation, and system troubleshooting.• Execute infrastructure changes and updates using Infrastructure as Code (Terraform) in alignment with change control procedures.• Monitor key AWS cloud infrastructure components including VPCs, Security Groups, EC2, IAM, S3, and KMS.5. CI/CD & GitOps Deployment Operations (GitLab, ArgoCD, Argo Workflows)• Execute, monitor, and troubleshoot automated deployment pipelines using GitLab CI/CD.• Manage application state, synchronization, and rollouts across Kubernetes clusters using ArgoCD (GitOps paradigm).• Monitor, execute, and troubleshoot operational batch processes, system maintenance tasks, and automated pipelines using Argo Workflows.• Facilitate application releases and configuration rollouts using Helm charts, Kustomize, and GitOps workflows.• Validate build pipeline compliance, container security scanning results, and image signature verifications prior to production deployment.6. Documentation & Operational Runbooks• Maintain accurate, step-by-step incident runbooks, standard operating procedures (SOPs), and triage workflows for both cloud and bare-metal environments.• Write detailed post-incident reports and post-mortems for production impact events.• Identify manual operational overhead (toil) and implement shell scripts (Bash/Python) to streamline routine monitoring, hardware checks, and maintenance tasks.Technical Skills & QualificationsRequired Qualifications• Citizenship & Regulatory Standard: US Citizenship required due to FedRAMP / AWS GovCloud security requirements.• Experience: 3–5 years of hands-on experience in SRE, DevOps, Systems Administration, or Hybrid Infrastructure Support roles.• 24/7 Shift Availability: Willingness and capability to work in a 24/7 rotational shift schedule (including nights, weekends, and holidays).• Bare-Metal & On-Premises Linux Administration: Strong hands-on experience administering bare-metal Linux servers (Red Hat Enterprise Linux / RHEL), including hardware diagnostics, out-of-band management (IPMI, iDRAC, iLO), RAID configuration, LVM, and physical NIC bonding.• Log Troubleshooting (Kibana/Elasticsearch): Direct experience querying and analyzing log streams in Kibana and Elasticsearch (KQL/Lucene queries, index management, log pattern matching for microservices and cluster components).• L4/L7 Secure Load Balancing: Practical experience troubleshooting Layer 4 (NLB) and Layer 7 (ALB / Ingress Controllers) secure load balancing, mTLS, TLS termination, health checks, and secure traffic routing.• AWS & GovCloud: Solid operational experience with AWS core services (EC2, VPC, IAM, S3, KMS) and familiarity with AWS GovCloud operating models.• Kubernetes (EKS): Hands-on operational experience with Amazon EKS / Kubernetes (kubectl, Helm, pod lifecycle management, node pool maintenance).• Observability Tools: Experience working with Prometheus, PromQL metrics queries, Alertmanager, and Grafana dashboard visualization.• CI/CD & Continuous Delivery: Hands-on operational experience executing and troubleshooting deployments with GitLab CI/CD, managing GitOps syncing with ArgoCD, and orchestrating Kubernetes workflows using Argo Workflows.• ChatOps & Communication: Experience utilizing Slack for operational messaging, ChatOps commands, and incident war rooms.• FedRAMP / Security Hardening: Understanding of FedRAMP/NIST SP 800-53 controls, vulnerability patching cycles, and DISA STIG hardening on both bare-metal and cloud Linux systems.Preferred Qualifications• Experience with automation scripting using Python or Bash.• Basic working knowledge of Terraform for infrastructure maintenance.• Red Hat Certified System Administrator (RHCSA) or Red Hat Certified Engineer (RHCE).• AWS Certified SysOps Administrator or Certified Kubernetes Administrator (CKA).Key Performance Indicators (KPIs)• Monitoring MTTA (Mean Time to Acknowledge): Rapid response times for eyes-on alert notifications.• MTTR (Mean Time to Resolve): Efficient triage, log analysis in Kibana, and mitigation of production service disruptions across cloud and bare-metal environments.• Hardware Availability & Uptime: Rapid isolation and remediation of physical server component failures and RAID disk faults.• Deployment & GitOps Stability: High release success rates and swift remediation of failed ArgoCD syncs or Argo Workflows runs.• Shift Readiness & Escalation Precision: Clear shift handovers and timely escalation management during 24/7 windows.• CVE Remediation Timeliness: Strict compliance with patching SLAs for bare-metal OS, container images, and cloud AMIs.• Runbook Currency: Continuous refinement of operational runbooks based on shift learnings.#LI-KA1The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.The annual base pay for this position is: $124,800.00 - $187,200.00F5 maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, geographic locations, and market conditions, as well as to reflect F5’s differing products, industries, and lines of business. The pay range referenced is as of the time of the job posting and is subject to change.You may also be offered incentive compensation, bonus, restricted stock units, and benefits. More details about F5’s benefits can be found at the following link: . F5 reserves the right to change or terminate any benefit plan without notice. Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com View email address on click.appcast.io).Equal Employment OpportunityIt is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination. F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting View email address on click.appcast.io: San Jose; RestonType: Full time

Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) - II in Reston, VA vacancy
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...looking for a Site Reliability Engineer II with the right mix of software...  ...Site Reliability Engineering (SRE) team provides leadership, direction... 
    Suggested
    Ongoing contract
    Work at office
    Local area
    Shift work
    3 days per week

    Microsoft

    Reston, VA
    3 days ago
  • $103.5k - $150k

     ...together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the...  ...that power a highly reliable global SaaS platform. As an SRE II, you will help operate and improve the reliability, scalability... 
    Suggested
    Temporary work
    Work experience placement
    Local area
    3 days per week

    Medallia

    McLean, VA
    4 days ago
  • $114.4k - $171.6k

     ...Overview: Join a growing team securing both leading-edge protection solutions and enterprise infrastructure. As a Site Reliability Engineer II, a part of the Operational Support Systems (OSS) organization under the Security & Distributed Cloud organization ,... 
    Suggested
    Local area

    F5

    Reston, VA
    1 day ago
  •  ...Site Reliability Engineer Location: Occasional onsite visits to Reston VA (Zip code 20190). Duration-1 year plus...  ...interview in Reston VA. Zip code: 20190 Strong AWS SRE/Platform Engineering with Python/Java, Terraform/Ansible (IaC... 
    Suggested
    Long term contract
    Temporary work
    H1b
    Immediate start
    Relocation

    3B Staffing LLC

    Reston, VA
    11 hours ago
  •  ...SummaryWe are seeking an experienced, security-focused Senior Site Reliability Engineer (Senior SRE) to drive the reliability, architectural design, and...  ...audits (Moderate or High), NIST SP 800-53, or SOC 2 Type II audits.• Architectural Design Skills: Strong capability in... 
    Suggested
    Full time
    Local area

    F5 Networks

    Reston, VA
    1 hour ago
  •  ...Period of performance: Up to 2 years in duration MUST HAVES: Minimum of 8 years of experience as a Site Reliability Engineer with a strong understanding of SRE principles for highly scalable and reliable systems Possess a bachelor's degree Experience working... 
    Local area
    Relocation package
    3 days per week

    Beyond SOF

    Vienna, VA
    4 days ago
  • $117.2k - $176.7k

     ...level of U.S. government background investigation and clearance required for this role.Overview of the Role:Join our Site Reliability Engineering (SRE) team, where you'll work alongside Infrastructure and Research & Development (R&D) partners to keep Salesforce cloud services... 
    Full time
    Work experience placement

    Salesforce

    Herndon, VA
    1 day ago
  •  ...CyLogic is seeking a highly experienced Senior SRE / DevOps Engineer to build, operate, and improve scalable, reliable, and observable platforms supporting mission-critical applications such as Omnissa Workspace ONE. This role combines deep expertise in automation, CI/... 
    Full time

    CyLogic

    Ashburn, VA
    2 days ago
  • $120k - $140k

     ...in 1997 and started by ensuring the reliability and performance of mission-critical databases...  ...is building a next-generation Site Reliability Engineering team, and we're looking for talented,...  ...problem-solving environments. As an SRE, you'll design, deploy, and operate large... 
    Work from home

    Pythian

    Reston, VA
    2 days ago
  • $91.4k - $187k

    Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible... 
    Temporary work
    Flexible hours
    Shift work
    Weekend work

    Oracle Corporation

    Reston, VA
    1 day ago
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...than the Microsoft Defender engineering team. You will be building and...  ...looking for Software Engineer II to join the team. You will...  ...services to be scalable and highly reliable.Help deliver and improve... 
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Reston, VA
    4 days ago
  • $112k - $179k

     ...US Responsibilities Peraton is seeking a Lead Site Reliability Engineer to join our team of qualified, diverse individuals. The ideal...  ...Qualifications Bachelor's degree and 8-10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering... 
    Contract work
    Shift work

    Peraton

    Herndon, VA
    1 day ago
  • $69.55k - $125.73k

     ...Description Leidos is seeking a Software Engineer II to support the U.S. Department of Justice (DOJ), Office of Justice Programs (OJP),...  ...OCIO) in improving, maintaining, testing, and supporting secure, reliable, accessible, and user-centered digital products and services.... 
    Work at office
    Local area
    Immediate start

    Leidos

    Reston, VA
    2 days ago
  •  ...SRE/DevOps Engineer Location: McLean, VA (5 Days mandatory) - Only locals/nearby F2F interview mandatory Developing appropriate DevOps channels throughout the organization. Evaluating, implementing and streamlining DevOps practices. Establishing a continuous... 
    Local area

    E-Solutions

    McLean, VA
    1 day ago
  •  ...motivated candidate to join our talented Team. Job Title: SRE / DevOps Engineer Job Location: Mclean, VA Duration: 3-month...  ...possibility of extension Job Description: We are seeking a Site Reliability Engineer (SRE) with strong expertise in the client... 

    Ampcus

    McLean, VA
    4 days ago
  • $109.18k - $163.77k

     ...SummaryFreeWheel is seeking a Data Domain SRE to join the FreeWheel Data...  ...responsible for ensuring the reliability, scalability, and performance...  .... Working closely with data engineers and other operation sub-teams...  ...summary on our careers site for more details.... 
    Full time

    Comcast

    Reston, VA
    11 hours ago
  •  ...and system access requirements.AttainX is seeking a DevSecOps Engineer II to integrate security practices throughout the DevOps lifecycle...  ...controlled release and rollback procedures to support secure, reliable delivery.Configure and maintain DocuSign accounts, groups,... 
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Remote work
    Monday to Friday

    Attainx

    Herndon, VA
    1 day ago
  •  ...Reston, Virginia Type: Contract Job #3428 DevSecOps Engineer II Reston, VA Active TS/SCI with Polygraph...  ...lifecycle, ensuring the delivery of high-quality, secure, and reliable software solutions. Collaborate with development, operations... 
    Contract work

    Cornerstone Defense

    Reston, VA
    11 hours ago
  • $86.4k - $176.2k

     ...Join us to drive positive, lasting change that moves missions and the government forward! Job Description : As a DevSecOps Engineer II, you will design, build, and maintain a stable and efficient infrastructure to optimize service delivery across production, QA, and... 
    Live in
    Work at office
    Local area

    Accenture

    Reston, VA
    3 days ago
  • $100k - $185k

     ...looking for a Software Expert in Chantilly, VA. Software Expert II provides independent technical advisory support, evaluating...  ...communication skillsBachelor’s Degree in Computer Science, Software Engineering, Systems Engineering, or related fieldPreferred Qualifications:... 
    Contract work
    For contractors
    Work experience placement

    VTG

    Chantilly, Loudoun County, VA
    2 days ago
  • $86.4k - $176.2k

     ...DevSecOps Engineer II Reston, VA At Accenture Federal Services, nothing matters more than helping the US federal government make the nation stronger and safer and life better for people. Our 13,000+ people are united in a shared purpose to pursue the limitless potential... 

    Accenture Federal Services

    Reston, VA
    4 days ago
  • $92.52k - $138.79k

     ...insights you need to drive results. FreeWheel’s platform makes TV and video advertising work.Job DescriptionWe're looking for a Site Reliability Engineer to own cloud infrastructure, system reliability, and observability for the Freewheel BuyerCloud and Revenue Science teams.... 
    Full time
    Worldwide

    Comcast

    Reston, VA
    1 hour ago
  • $114.4k - $125.4k

     ...Are you passionate about ensuring the reliability and performance of mission-critical cloud...  ...? Salesforce is seeking a talented Site Reliability Engineer to join our dynamic team, supporting our...  ..., including concepts such as Safety II — looking at how things go right instead... 
    Full time
    Local area
    Shift work
    Night shift

    Salesforce

    McLean, VA
    3 days ago
  • $81.1k - $187k

     .... You’ll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers...  ...posted.Career Level - IC3Escalation points for junior Site Reliability Engineers during complex or high-impact incidents.Manage and... 
    Temporary work
    Monday to Friday
    Flexible hours
    Shift work
    Night shift

    Oracle Corporation

    Reston, VA
    4 days ago
  •  ...We are looking for a talented Senior Site Reliability Engineer to join our team to deliver world class search technologies to mobile devices. You will be working with a smart team of Engineers to lead and drive the stability, reliability, and observability of all Seekr... 
    Permanent employment
    Flexible hours

    Seekr

    Reston, VA
    1 day ago
  •  ...Site Reliability Engineer needed for a full time opportunity with SOC's direct client based in Herndon, VA. **Due to federal requirements, candidates must hold and possess an Active DOW TS/SCI security clearance to be considered for this role.** SOC is seeking... 
    Full time

    SOC

    Herndon, VA
    1 day ago
  • $87.1k - $157.45k

     ...throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across...  ...to come in and help us build systems that stay reliable when things get complicated. We need a Site Reliability Engineer who has experience building, deploying... 
    Local area
    Immediate start
    Work from home
    Flexible hours

    Leidos

    Reston, VA
    1 day ago
  •  ...2-3 days in Reston, VA   As Sr. .Net Developer II, you will serve as a senior full-stack engineer and technical lead, building and modernizing enterprise...  ...with stakeholders to deliver secure, scalable, and reliable solutions.   Due to federal security clearance... 
    Hourly pay
    Permanent employment
    Full time
    Work at office
    Local area

    Eliassen Group

    Reston, VA
    a month ago
  • $100k - $160k

     ...bring to the team; and empower our employees to create innovative and trusted results. We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer (SRE) will manage, monitor, and... 
    Temporary work

    Cathexis

    McLean, VA
    1 day ago
  • $129.2k - $174.8k

    The Platform Engineering and Emerging Technologies (PEET) team is hiring for a System Development Engineer to support AWS Cryptography...  ...experience- 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience-... 
    Internship
    Flexible hours
    Night shift
    Weekend work

    Amazon

    Herndon, VA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - II. Be the first to apply!