Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Data Site Reliability Engineer (SRE)

General Dynamics Information Technology

Public Trust: BI Full 6C (T4)
Requisition Type: Regular
Your Impact

Own your opportunity to support the missions that matter. From working with technologies like AI, cyber and cloud to careers in intelligence and health, we offer endless opportunities to apply your expertise to create a safer, smarter world while building new skills to propel your career forward.

Job Description

Seize your opportunity to make a personal impact supporting the Case Management Modernization (CMM) Program. The CMM program is an initiative to support the Administrative Office of the US Courts (AO) in developing a modern cloud-based solution to support all 204+ federal courts across the United States.

GDIT is your place to make meaningful contributions to challenging projects and grow a rewarding career. The Data Site Reliability Engineer (SRE) will work as part of the CMM Data Modernization and Governance team responsible for delivering integrated data governance, engineering, data platform, reporting, analytics, and Artificial Intelligence (AI)/Machine Learning (ML) capabilities that support operational decision-making and fulfill the AO's data and analytics objectives in support of the CMM program.

The successful candidate will be responsible for providing technical leadership for the day-to-day operational support, reliability, performance, and continuous improvement of the CMM data platforms, pipelines, applications, and analytics services. This role ensures that data services remain secure, available, reliable, and aligned with established service levels, data governance standards, architecture principles, and operational procedures.

THE DATA SITE RELIABILITY ENGINEER (SRE) WILL EXECUTE THE FOLLOWING RESPONSIBILITIES

  • Provide comprehensive real-time monitoring, incident and event management, capacity planning, and operational reporting to support application deployments, maintain system health, predict demand, and align cloud operations with evolving business and security objectives.

  • Maintain and audit user roles and responsibilities in cloud environments.

  • Integrate Single Sign On (SSO), Multi-Factor Authentication (MFA) and group identity management managed through the Judiciary Enterprise Network Information Exchange (JENIE) for enforcing least privilege access.

  • Adhere to guidelines prescribed by the Government and continuously assess and improve credential management processes for all user credentials.

  • Provide Disaster Recovery (DR) and Continuity of Operations (COOP) options. This must include high-availability options, including fault-tolerant and automated failover designs.

  • Integrate DevSecOps tools and processes seamlessly with enterprise systems (Integrated Development Environments (IDEs), ticketing, monitoring, etc.) to avoid fragmentation and ensure unified security posture.

  • Provide and manage a centralized secrets management system with automated rotation, access logging, and policy enforcement to securely store, manage, and control access to sensitive information and to prevent unauthorized access and data breaches for any administrative user account.

  • Integrate security tools (example: SAST, DAST, SCA, CSPM) into pipelines for continuous assessment and remediation.

  • Implement unified, automated, continuous monitoring (24/7/365) systems and tools for security, performance, and compliance across all environments, leveraging dashboards and alerting for real-time visibility. Provide supplemental monitoring of event response activities beyond normal business hours (7a.m – 6p.m Eastern Time). Systems and tools shall capture data without including a required response to alerts.

  • Ensure automated generation and management of Software Bill of Materials (SBOM) for all deployed artifacts, supporting transparency and compliance.

  • Provide diagnostics, metrics’ gathering, and performance tuning services.

  • Provide canary release function for end-user testing to support beta testing.

  • Configure an alert mechanism so that the support teams can react in an instance of unusual behavior.

  • Implement and operate a comprehensive incident and event management process, including integration with enterprise SIEM solutions, automated alerting, escalation workflows, and root cause analysis for all critical incidents.

  • Provide engineering support to ensure prompt detection, logging, diagnosis, escalation, and resolution of incidents to restore normal service operations as quickly as possible and minimize impact.

  • Perform systems support in identifying, analyzing, and eliminating the root causes of recurring incidents to minimize continued adverse impacts and potential degradation of services.

  • Make recommendations for the improvement of Incident and Problem management consistent with industry’s best practices for the cloud.

  • Maintain knowledge base of known issues, resolutions, and best practices for operational continuity.

  • Perform automated health checks across the full stack (Operating System, Application, Database and PaaS services) at agreed levels on an agreed frequency.

  • Provide a monthly issues management report. The report shall include cloud-related incidents, any stability and performance issues, configurations issues, quantity of tickets received, and time duration to resolve tickets.

  • Develop and implement thresholds, rules, and response procedures based on product team’s recommendation.

  • Monitor resource utilization (e.g., CPU, Memory, Disk Space) for the cloud hosted Virtual Machines (VMs) and other cloud services.

  • Manage the resolution procedures for any threshold breaches for cloud resources.

  • Improves system reliability, observability, automation, scalability, and operational resilience through engineering practices.

  • Monitors, maintains, and optimizes cloud infrastructure, databases, and platform services for reliability and performance.

  • Act as FinOps Analyst and perform cost optimization.

QUALIFICATIONS

  • Education: Bachelor's degree in Computer Science, Software Engineering, or related field. (Or equivalent experience.)

  • Experience: 5+ years’ experience in IT systems engineering, systems development, systems coding, and programming.

  • Deep expertise with AWS services, including monitoring, logging, compute, storage, and networking.

  • Proficiency in Infrastructure as Code (IaC) tools like Terraform, AWS CloudFormation, or Azure Bicep.

  • Hands-on experience with monitoring and APM tools such as CloudWatch, Azure Monitor, Datadog, Prometheus, Grafana, New Relic, etc.

  • Solid understanding of incident response, change management, and ITIL-based operational support.

  • Familiarity with CI/CD toolchains and automation platforms (Jenkins, GitHub Actions, GitLab, ArgoCD).

  • Strong scripting skills (Python, PowerShell, Bash) for automation and orchestration.

  • Advanced experience in providing DevSecOps implementation using GitOps, or similar tools.

  • Experienced in developing, testing, and maintaining containerized applications.

  • Expert knowledge of source version control, build/release tools and methodologies, CI/CD pipelines and the Software Build process.

  • Experience in building and maintaining CI/CD pipelines for large enterprises that consist of a large number of complex applications.

  • Ability to be flexible and work on several different products while supporting multiple teams

  • Experience with FinOps practices, cost modeling, forecasting, and optimization tools within cloud platforms.

  • Understanding of federal compliance and security frameworks (e.g., FedRAMP, NIST, JISF Rev 5).

  • Ability to analyze logs and metrics and conduct performance tuning for cloud-based services and applications.

  • Experience working across multiple product teams to get a grasp of a product and/or programs overall state of health.

  • ITIL, AWS SysOps, or Google Professional Cloud DevOps Engineer certifications are a plus.

COMMUNICATION & ORGANIZATIONAL SKILLS

  • Excellent presentation and communication skills.

  • Consultant mindset with the ability to work with high level customer stakeholders and build excellent customer relationships.

  • Experience identifying and applying industry tools, solutions, methods best practices, and emerging technologies.

  • Strong analytical skills and problem-solving skills with the ability to formulate and communicate recommendations for improvement.

  • Experience with process design and documentation methodologies, and design and production of quality deliverables, process and use case modeling, business case development.

  • Demonstrated ability to work effectively, independently, and as part of a team.

Security Clearance Level: Must be able to pass a background check to obtain a position of Public Trust.

Must be a US Person (Green Card Holder, US Permanent Resident Alien, Refugee, Asylee, or US Citizen).

Location: Remote.

GDIT IS YOUR PLACE
At GDIT, the mission is our purpose, and our people are at the center of everything we do.

● Growth: AI-powered career tool that identifies career steps and learning opportunities
● Support: An internal mobility team focused on helping you achieve your career goals
● Rewards: Comprehensive benefits and wellness packages, 401K with company match, and competitive pay and paid time off
● Flexibility: Full-flex work week to own your priorities at work and at home
● Community: Award-winning culture of innovation and a military-friendly workplace

OWN YOUR OPPORTUNITY
Explore an enterprise IT career at GDIT and you’ll find endless opportunities to grow alongside colleagues who share your desire to drive operations forward.

#GDITLA

Work Requirements

Years of Experience

5 + years of related experience

* may vary based on technical training, certification(s), or degree

Certification

Travel Required

Less than 10%

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Data Site Reliability Engineer (SRE) in Remote vacancy
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that...  ..., ELK stack) and logging systems for real-time data monitoring.  Hands on experience with rack mount servers... 
    Data
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    19 hours ago
  •  ...Site Reliability Engineer (SRE) Immediate need for a talented Site Reliability Engineer (SRE). This is a 12+ months contract opportunity with long...  ...partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or... 
    Data
    Contract work
    Local area
    Immediate start

    Pyramid Corporation

    Chicago, IL
    3 days ago
  •  ...Site Reliability Engineer (SRE) Devexperts works with respected financial institutions, delivering products and tailor-made solutions for retail...  ...automation, complex software development projects, market data products, and IT consulting services. Job Description... 
    Data
    Local area
    Remote work
    Flexible hours

    DevExperts

    United States
    2 days ago
  •  ...SRE Key Responsibilities • Design and manage multi-account AWS infrastructure (VPC,...  ...Exporter, Prometheus Push Gateway) • Support data platforms (Kafka/Kafka UI, Minion, Airflow...  ...in Bash, Python, Go, C#/.NET (Unity Game Engine) • Maintain developer experience (... 
    Data
    Remote work

    ACI Infotech

    United States
    19 hours ago
  •  ...Senior Site Reliability Engineer (SRE) We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-...  ...Kubernetes and/or OpenStack ~ Experience with high-performance data center processing, networking, and storage ~ Exposure to... 
    Data
    Remote work

    Mirantis

    United States
    19 hours ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and... 
    Data
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Shrewsbury, MA
    2 days ago
  •  ...Technology group delivers secure, reliable technology solutions that...  ...needs and implementing data standards and governance....  ...Application Support Engineer, you will help power DTCC'...  ...and settlement.Leveraging Site Reliability Engineering (SRE) principles, you will support... 
    Data
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Boston, MA
    1 day ago
  • $165k - $225k

     ...model training, and demanding data processing workloads.We...  ...workloads with enterprise-grade reliability and compliance. Your Role:...  ...Working closely with our systems engineers, network engineers, and...  ...Requirements Experience: 5+ years in SRE, DevOps, or infrastructure... 
    Data
    Remote work
    Flexible hours

    Moonlite

    Chicago, IL
    2 days ago
  • $87.72k - $109.65k

     ...Req ID: 381297 NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us....  ...thinking organization, apply now. We are currently seeking a Site Reliability Engineering (SRE) - Maryland, US to join our team in Baltimore, Maryland (... 
    Data
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA Services

    Baltimore, MD
    6 days ago
  •  ...products and ensuring they’re reliable and highly available in cloud...  ...technical problems and resolving data/configuration issues within...  ...projects Collaboration with cloud engineers in understanding new cloud...  ...~2+ years of related SRE experience ~ Apply core software... 
    Data
    Full time
    Work experience placement
    Remote work

    Tenable

    Remote
    29 days ago
  •  ...trillion in crypto transactions. We are looking for a Head of Site Reliability Engineering (SRE) who will serve as the principal leader in developing and...  ...hybrid infrastructure across GCP, AWS, and on-premise data centers, deploying sensitive components (using Kubernetes... 
    Data
    Full time
    Apprenticeship
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Blockchain

    United Kingdom
    22 days ago
  • $175k - $250k

     ...developed by our expert team of lawyers, engineers and research scientists. We’ve found...  ...Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the...  ...planning, graceful rollouts, and safe data access to maintain high reliability and... 
    Data
    Full time
    Relocation package

    Harvey

    Remote
    19 hours ago
  • $120k - $130k

     ...Tittle : SRE DevOps Engineer Location: Across USA any Location The pay range for this role is $120k - $130k per annum including...  ...Google Cloud ( GCS, BigQuery ). # Experience using GCP Data Engineering stack ( Composer, Dataflow, Dataproc ). #... 
    Data
    Remote work

    Yochana

    United States
    19 hours ago
  •  ...staffing experts can help you find the best job for you. Role: SRE DevOps Engineer Location: Austin TX Duration: 6 months Required...  ...Administration - o Designing and configuring clusters, analyzing data access patterns, identifying query performance issues,... 
    Data
    Permanent employment
    Contract work
    Remote work

    Tekfortune Inc

    Austin, TX
    4 days ago
  • $105.1k - $164.13k

     ...NextGen transformation strategy.Lead the reliability, performance and operations work-...  ...Management (SWIM) Flight Data Publication Service (SFDPS) program:...  ...define RMA improvement plans, implement Site Reliability Engineering (SRE) practices (SLIs/SLOs/SLAs), lead incident... 
    Data
    Permanent employment
    Full time
    Contract work
    Part time
    Local area
    Remote work

    Noblis

    Atlantic City, NJ
    5 days ago
  •  ...Req ID: 387092 NTT DATA strives to hire exceptional, innovative and passionate individuals...  ...apply now. We are currently seeking a SRE Reliability Engineer to join our team in Bangalore, Karnātaka (IN-KA), India (IN). Site Reliability Engineer (SRE) – Kubernetes /... 
    Data
    Permanent employment
    Work at office
    Remote work
    Flexible hours

    NTT DATA, Inc.

    Indiana
    3 days ago
  •  ...technical leadership to a growing team focused on applying software engineering practices to operations at scale. Monitor and report on...  ..., Python, Go, Perl or Ruby . Experience with algorithms, data structures, complexity analysis and software design.... 
    Data
    Full time

    PointClickCare

    Remote
    2 days ago
  •  ...company's first combined turbojet-ramjet engine and is now being scaled through its...  ...Hermeus is seeking a Senior Software or Site Reliability Engineer to join the Information Team and...  ...hardware and software in-the-loop systems. Data generated from engineering test will be... 
    Data
    Full time
    Remote work

    Hermeus

    Atlanta, GA
    10 days ago
  •  ...Job Description Job Description SRE Support Engineer - Observability While this position is...  ...company delivering large-scale cloud, data, and engineering solutions across 130+...  ...Slack and tickets, improving monitoring reliability, and reducing incident impact through... 
    Data
    Remote work

    Virtasant

    Austin, TX
    2 days ago
  • $1,000 per month

     ...Responsibilities Own reliability, deployments, observability, and incident...  ...Requirements Demonstrated DevOps/SRE depth and a genuine backend software engineering background with shipped...  ...vendor lock-in or migrating between data stores at scale. Benefits ~... 
    Data
    Full time
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Credit Genie

    Philadelphia, PA
    6 days ago
  •  ...products and ensuring they’re reliable and highly available in cloud...  ...technical problems and resolving data/configuration issues within...  ...projects Collaboration with cloud engineers in understanding new cloud...  ...~2+ years of related SRE experience ~ Apply core software... 
    Data
    Full time
    Work experience placement
    Remote work

    Tenable

    Remote
    23 days ago
  •  ...support PayPay’s exponential growth. As an SRE at PayPay, we strive towards ensuring high...  ...so that our users can have flawless and reliable service exceeding expectations. Considering...  ...Java, Go, etc with strong fundamentals in data structures, algorithms, problem solving... 
    Data
    Remote job
    Full time
    Work at office
    Visa sponsorship
    Relocation package
    Flexible hours

    PayPay

    Remote
    more than 2 months ago
  • $160k - $185k

     ...fitness journey and revolutionized the industry along the way. And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous improvement of reliability, availability, and performance across Planet... 
    Work at office
    Local area
    Remote work
    Work from home

    Planet Fitness

    Hampton, NH
    1 day ago
  •  ...Site Reliability Engineer (SRE) Location: Remote Shift Timings: 5:30 PM to 3:00 AM IST to ensure support for global operations. Job Description: We are seeking a skilled Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have... 
    Remote work
    Shift work

    InOrg Global

    United States
    1 day ago
  •  ...Site Reliability Engineer (SRE) Location: Remote (Secaucus, NJ) Duration: Contract Experience: 7+ Years Job Description 4+ years of experience with multiple APM tools and extensive experience with Dynatrace 4+ years of experience executing software load and performance... 
    Contract work
    Work experience placement
    Remote work

    Syntricate Technologies

    United States
    3 days ago
  • $175k - $185k

     ...Senior Site Reliability Engineer (SRE) Remote, US Branch is on a mission to empower workers with financial freedom. We do this by helping companies accelerate payments and providing working Americans with accessible, free financial services. We're committed to building... 
    Daily paid
    Remote work
    Home office
    Flexible hours

    Branch

    United States
    2 days ago
  •  ...Senior Site Reliability Engineer At Swile, we believe that good products can help reduce friction in daily professional life and boost employee...  ...Brazil. Your role as a Senior Site Reliability Engineer (SRE) centers around creatively solving problems, ensuring a balance... 
    Remote work

    Swile

    United States
    1 day ago
  •  ...in Computer Science, Information Technology, Engineering, or equivalent field ~3-5 years of experience in Site Reliability Engineering, Production Support, Platform Engineering...  ...application health ~ Understanding of SRE principles, including observability,... 
    Remote work

    Anveta

    United States
    3 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Remote work

    Noctua Technology

    Washington DC
    19 hours ago
  • $106.5k - $177.5k

     ...Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the... 
    Remote work

    Noctua Technology

    Washington DC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Data Site Reliability Engineer (SRE). Be the first to apply!