Data Site Reliability Engineer (SRE)
General Dynamics Information Technology
Public Trust: BI Full 6C (T4)
Requisition Type: Regular
Your Impact
Own your opportunity to support the missions that matter. From working with technologies like AI, cyber and cloud to careers in intelligence and health, we offer endless opportunities to apply your expertise to create a safer, smarter world while building new skills to propel your career forward.
Job Description
Seize your opportunity to make a personal impact supporting the Case Management Modernization (CMM) Program. The CMM program is an initiative to support the Administrative Office of the US Courts (AO) in developing a modern cloud-based solution to support all 204+ federal courts across the United States.
GDIT is your place to make meaningful contributions to challenging projects and grow a rewarding career. The Data Site Reliability Engineer (SRE) will work as part of the CMM Data Modernization and Governance team responsible for delivering integrated data governance, engineering, data platform, reporting, analytics, and Artificial Intelligence (AI)/Machine Learning (ML) capabilities that support operational decision-making and fulfill the AO's data and analytics objectives in support of the CMM program.
The successful candidate will be responsible for providing technical leadership for the day-to-day operational support, reliability, performance, and continuous improvement of the CMM data platforms, pipelines, applications, and analytics services. This role ensures that data services remain secure, available, reliable, and aligned with established service levels, data governance standards, architecture principles, and operational procedures.
THE DATA SITE RELIABILITY ENGINEER (SRE) WILL EXECUTE THE FOLLOWING RESPONSIBILITIES
Provide comprehensive real-time monitoring, incident and event management, capacity planning, and operational reporting to support application deployments, maintain system health, predict demand, and align cloud operations with evolving business and security objectives.
Maintain and audit user roles and responsibilities in cloud environments.
Integrate Single Sign On (SSO), Multi-Factor Authentication (MFA) and group identity management managed through the Judiciary Enterprise Network Information Exchange (JENIE) for enforcing least privilege access.
Adhere to guidelines prescribed by the Government and continuously assess and improve credential management processes for all user credentials.
Provide Disaster Recovery (DR) and Continuity of Operations (COOP) options. This must include high-availability options, including fault-tolerant and automated failover designs.
Integrate DevSecOps tools and processes seamlessly with enterprise systems (Integrated Development Environments (IDEs), ticketing, monitoring, etc.) to avoid fragmentation and ensure unified security posture.
Provide and manage a centralized secrets management system with automated rotation, access logging, and policy enforcement to securely store, manage, and control access to sensitive information and to prevent unauthorized access and data breaches for any administrative user account.
Integrate security tools (example: SAST, DAST, SCA, CSPM) into pipelines for continuous assessment and remediation.
Implement unified, automated, continuous monitoring (24/7/365) systems and tools for security, performance, and compliance across all environments, leveraging dashboards and alerting for real-time visibility. Provide supplemental monitoring of event response activities beyond normal business hours (7a.m – 6p.m Eastern Time). Systems and tools shall capture data without including a required response to alerts.
Ensure automated generation and management of Software Bill of Materials (SBOM) for all deployed artifacts, supporting transparency and compliance.
Provide diagnostics, metrics’ gathering, and performance tuning services.
Provide canary release function for end-user testing to support beta testing.
Configure an alert mechanism so that the support teams can react in an instance of unusual behavior.
Implement and operate a comprehensive incident and event management process, including integration with enterprise SIEM solutions, automated alerting, escalation workflows, and root cause analysis for all critical incidents.
Provide engineering support to ensure prompt detection, logging, diagnosis, escalation, and resolution of incidents to restore normal service operations as quickly as possible and minimize impact.
Perform systems support in identifying, analyzing, and eliminating the root causes of recurring incidents to minimize continued adverse impacts and potential degradation of services.
Make recommendations for the improvement of Incident and Problem management consistent with industry’s best practices for the cloud.
Maintain knowledge base of known issues, resolutions, and best practices for operational continuity.
Perform automated health checks across the full stack (Operating System, Application, Database and PaaS services) at agreed levels on an agreed frequency.
Provide a monthly issues management report. The report shall include cloud-related incidents, any stability and performance issues, configurations issues, quantity of tickets received, and time duration to resolve tickets.
Develop and implement thresholds, rules, and response procedures based on product team’s recommendation.
Monitor resource utilization (e.g., CPU, Memory, Disk Space) for the cloud hosted Virtual Machines (VMs) and other cloud services.
Manage the resolution procedures for any threshold breaches for cloud resources.
Improves system reliability, observability, automation, scalability, and operational resilience through engineering practices.
Monitors, maintains, and optimizes cloud infrastructure, databases, and platform services for reliability and performance.
Act as FinOps Analyst and perform cost optimization.
QUALIFICATIONS
Education: Bachelor's degree in Computer Science, Software Engineering, or related field. (Or equivalent experience.)
Experience: 5+ years’ experience in IT systems engineering, systems development, systems coding, and programming.
Deep expertise with AWS services, including monitoring, logging, compute, storage, and networking.
Proficiency in Infrastructure as Code (IaC) tools like Terraform, AWS CloudFormation, or Azure Bicep.
Hands-on experience with monitoring and APM tools such as CloudWatch, Azure Monitor, Datadog, Prometheus, Grafana, New Relic, etc.
Solid understanding of incident response, change management, and ITIL-based operational support.
Familiarity with CI/CD toolchains and automation platforms (Jenkins, GitHub Actions, GitLab, ArgoCD).
Strong scripting skills (Python, PowerShell, Bash) for automation and orchestration.
Advanced experience in providing DevSecOps implementation using GitOps, or similar tools.
Experienced in developing, testing, and maintaining containerized applications.
Expert knowledge of source version control, build/release tools and methodologies, CI/CD pipelines and the Software Build process.
Experience in building and maintaining CI/CD pipelines for large enterprises that consist of a large number of complex applications.
Ability to be flexible and work on several different products while supporting multiple teams
Experience with FinOps practices, cost modeling, forecasting, and optimization tools within cloud platforms.
Understanding of federal compliance and security frameworks (e.g., FedRAMP, NIST, JISF Rev 5).
Ability to analyze logs and metrics and conduct performance tuning for cloud-based services and applications.
Experience working across multiple product teams to get a grasp of a product and/or programs overall state of health.
ITIL, AWS SysOps, or Google Professional Cloud DevOps Engineer certifications are a plus.
COMMUNICATION & ORGANIZATIONAL SKILLS
Excellent presentation and communication skills.
Consultant mindset with the ability to work with high level customer stakeholders and build excellent customer relationships.
Experience identifying and applying industry tools, solutions, methods best practices, and emerging technologies.
Strong analytical skills and problem-solving skills with the ability to formulate and communicate recommendations for improvement.
Experience with process design and documentation methodologies, and design and production of quality deliverables, process and use case modeling, business case development.
Demonstrated ability to work effectively, independently, and as part of a team.
Security Clearance Level: Must be able to pass a background check to obtain a position of Public Trust.
Must be a US Person (Green Card Holder, US Permanent Resident Alien, Refugee, Asylee, or US Citizen).
Location: Remote.
GDIT IS YOUR PLACE
At GDIT, the mission is our purpose, and our people are at the center of everything we do.
● Growth: AI-powered career tool that identifies career steps and learning opportunities
● Support: An internal mobility team focused on helping you achieve your career goals
● Rewards: Comprehensive benefits and wellness packages, 401K with company match, and competitive pay and paid time off
● Flexibility: Full-flex work week to own your priorities at work and at home
● Community: Award-winning culture of innovation and a military-friendly workplace
Explore an enterprise IT career at GDIT and you’ll find endless opportunities to grow alongside colleagues who share your desire to drive operations forward.
#GDITLA
Work Requirements
Years of Experience
5 + years of related experience
* may vary based on technical training, certification(s), or degree
Certification
Travel Required
Less than 10%
$160k - $200k
Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that... ..., ELK stack) and logging systems for real-time data monitoring. Hands on experience with rack mount servers...DataLocal areaRemote work- ...Site Reliability Engineer (SRE) Immediate need for a talented Site Reliability Engineer (SRE). This is a 12+ months contract opportunity with long... ...partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or...DataContract workLocal areaImmediate start
- ...Site Reliability Engineer (SRE) Devexperts works with respected financial institutions, delivering products and tailor-made solutions for retail... ...automation, complex software development projects, market data products, and IT consulting services. Job Description...DataLocal areaRemote workFlexible hours
- ...SRE Key Responsibilities • Design and manage multi-account AWS infrastructure (VPC,... ...Exporter, Prometheus Push Gateway) • Support data platforms (Kafka/Kafka UI, Minion, Airflow... ...in Bash, Python, Go, C#/.NET (Unity Game Engine) • Maintain developer experience (...DataRemote work
- ...Senior Site Reliability Engineer (SRE) We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-... ...Kubernetes and/or OpenStack ~ Experience with high-performance data center processing, networking, and storage ~ Exposure to...DataRemote work
$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and...DataFull timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...Technology group delivers secure, reliable technology solutions that... ...needs and implementing data standards and governance.... ...Application Support Engineer, you will help power DTCC'... ...and settlement.Leveraging Site Reliability Engineering (SRE) principles, you will support...DataRemote workFlexible hours
$165k - $225k
...model training, and demanding data processing workloads.We... ...workloads with enterprise-grade reliability and compliance. Your Role:... ...Working closely with our systems engineers, network engineers, and... ...Requirements Experience: 5+ years in SRE, DevOps, or infrastructure...DataRemote workFlexible hours$87.72k - $109.65k
...Req ID: 381297 NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us.... ...thinking organization, apply now. We are currently seeking a Site Reliability Engineering (SRE) - Maryland, US to join our team in Baltimore, Maryland (...DataTemporary workWork at officeRemote workFlexible hours- ...products and ensuring they’re reliable and highly available in cloud... ...technical problems and resolving data/configuration issues within... ...projects Collaboration with cloud engineers in understanding new cloud... ...~2+ years of related SRE experience ~ Apply core software...DataFull timeWork experience placementRemote work
- ...trillion in crypto transactions. We are looking for a Head of Site Reliability Engineering (SRE) who will serve as the principal leader in developing and... ...hybrid infrastructure across GCP, AWS, and on-premise data centers, deploying sensitive components (using Kubernetes...DataFull timeApprenticeshipWork at officeRemote workWorldwideFlexible hours
$175k - $250k
...developed by our expert team of lawyers, engineers and research scientists. We’ve found... ...Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the... ...planning, graceful rollouts, and safe data access to maintain high reliability and...DataFull timeRelocation package$120k - $130k
...Tittle : SRE DevOps Engineer Location: Across USA any Location The pay range for this role is $120k - $130k per annum including... ...Google Cloud ( GCS, BigQuery ). # Experience using GCP Data Engineering stack ( Composer, Dataflow, Dataproc ). #...DataRemote work- ...staffing experts can help you find the best job for you. Role: SRE DevOps Engineer Location: Austin TX Duration: 6 months Required... ...Administration - o Designing and configuring clusters, analyzing data access patterns, identifying query performance issues,...DataPermanent employmentContract workRemote work
$105.1k - $164.13k
...NextGen transformation strategy.Lead the reliability, performance and operations work-... ...Management (SWIM) Flight Data Publication Service (SFDPS) program:... ...define RMA improvement plans, implement Site Reliability Engineering (SRE) practices (SLIs/SLOs/SLAs), lead incident...DataPermanent employmentFull timeContract workPart timeLocal areaRemote work- ...Req ID: 387092 NTT DATA strives to hire exceptional, innovative and passionate individuals... ...apply now. We are currently seeking a SRE Reliability Engineer to join our team in Bangalore, Karnātaka (IN-KA), India (IN). Site Reliability Engineer (SRE) – Kubernetes /...DataPermanent employmentWork at officeRemote workFlexible hours
- ...technical leadership to a growing team focused on applying software engineering practices to operations at scale. Monitor and report on... ..., Python, Go, Perl or Ruby . Experience with algorithms, data structures, complexity analysis and software design....DataFull time
- ...company's first combined turbojet-ramjet engine and is now being scaled through its... ...Hermeus is seeking a Senior Software or Site Reliability Engineer to join the Information Team and... ...hardware and software in-the-loop systems. Data generated from engineering test will be...DataFull timeRemote work
- ...Job Description Job Description SRE Support Engineer - Observability While this position is... ...company delivering large-scale cloud, data, and engineering solutions across 130+... ...Slack and tickets, improving monitoring reliability, and reducing incident impact through...DataRemote work
$1,000 per month
...Responsibilities Own reliability, deployments, observability, and incident... ...Requirements Demonstrated DevOps/SRE depth and a genuine backend software engineering background with shipped... ...vendor lock-in or migrating between data stores at scale. Benefits ~...DataFull timeTemporary workWork at officeImmediate startRemote workFlexible hours- ...products and ensuring they’re reliable and highly available in cloud... ...technical problems and resolving data/configuration issues within... ...projects Collaboration with cloud engineers in understanding new cloud... ...~2+ years of related SRE experience ~ Apply core software...DataFull timeWork experience placementRemote work
- ...support PayPay’s exponential growth. As an SRE at PayPay, we strive towards ensuring high... ...so that our users can have flawless and reliable service exceeding expectations. Considering... ...Java, Go, etc with strong fundamentals in data structures, algorithms, problem solving...DataRemote jobFull timeWork at officeVisa sponsorshipRelocation packageFlexible hours
$160k - $185k
...fitness journey and revolutionized the industry along the way. And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous improvement of reliability, availability, and performance across Planet...Work at officeLocal areaRemote workWork from home- ...Site Reliability Engineer (SRE) Location: Remote Shift Timings: 5:30 PM to 3:00 AM IST to ensure support for global operations. Job Description: We are seeking a skilled Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have...Remote workShift work
- ...Site Reliability Engineer (SRE) Location: Remote (Secaucus, NJ) Duration: Contract Experience: 7+ Years Job Description 4+ years of experience with multiple APM tools and extensive experience with Dynatrace 4+ years of experience executing software load and performance...Contract workWork experience placementRemote work
$175k - $185k
...Senior Site Reliability Engineer (SRE) Remote, US Branch is on a mission to empower workers with financial freedom. We do this by helping companies accelerate payments and providing working Americans with accessible, free financial services. We're committed to building...Daily paidRemote workHome officeFlexible hours- ...Senior Site Reliability Engineer At Swile, we believe that good products can help reduce friction in daily professional life and boost employee... ...Brazil. Your role as a Senior Site Reliability Engineer (SRE) centers around creatively solving problems, ensuring a balance...Remote work
- ...in Computer Science, Information Technology, Engineering, or equivalent field ~3-5 years of experience in Site Reliability Engineering, Production Support, Platform Engineering... ...application health ~ Understanding of SRE principles, including observability,...Remote work
$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing...Remote work$106.5k - $177.5k
...Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Site Reliability Engineer (SRE). Be the first to apply!
- data center engineer Remote
- senior cloud data engineer Remote
- data infrastructure engineer Remote
- sr information security engineer Remote
- hadoop big data developer Remote
- data visualization developer Remote
- remote data engineer Remote
- etl data engineer Remote
- data engineer analytics Remote
- sr data engineer Remote




