Site Reliability Engineer
Diné Development
The Site Reliability Engineer (SRE) / Subject Matter Expert (SME) – Computer Systems Engineer/Architect will provide senior-level reach-back expertise to support the reliability, scalability, performance, and operational resilience of the GEOMAP platform in secure cloud environments. This role focuses on improving service availability, monitoring, incident response, automation, and production stability across cloud-hosted and containerized systems supporting mission-critical geospatial capabilities for the U.S. Air Force.The Site Reliability Engineer will collaborate across development, DevSecOps, cloud, database, testing, and support teams to identify systemic issues, reduce operational risk, and implement engineering solutions that improve long-term platform reliability. *This position is contingent upon contract award.* Responsibilities
- Provide senior-level engineering support to improve reliability, availability, performance, and maintainability of GEOMAP cloud-hosted systems and services.
- Analyze production issues, recurring incidents, and operational trends to identify root causes and recommend durable corrective actions.
- Support the design and implementation of monitoring, alerting, logging, and observability solutions across applications, infrastructure, and containerized services.
- Develop and recommend automation approaches that reduce manual effort, improve deployment consistency, and increase system resilience.
- Partner with software engineers, DevSecOps engineers, Kubernetes engineers, database engineers, and production support personnel to improve service health and release readiness.
- Support incident response, problem management, service restoration, and post-incident reviews for high-priority operational issues.
- Evaluate system performance, capacity, and scalability needs and provide recommendations for optimization and operational risk reduction.
- Assist in defining service reliability objectives, operational metrics, and support models for sustained mission operations.
- Contribute to infrastructure and platform engineering efforts involving cloud environments, CI/CD pipelines, container orchestration, and secure deployment patterns.
- Support architecture reviews, technical assessments, and engineering analyses related to reliability, recoverability, and production operations.
- Develop or refine runbooks, standard operating procedures, reliability engineering practices, and technical documentation.
- Provide reach-back support for surge requirements, complex production investigations, and priority modernization or stabilization efforts as directed.
- Performs other related duties as assigned.
- Active Secret clearance required.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field; Master’s degree preferred.
- Minimum of 8 years of experience supporting enterprise systems, cloud platforms, site reliability engineering, production engineering, systems engineering, or related technical roles.
- Experience supporting AWS environments, including monitoring, performance tuning, troubleshooting, incident response, and operational sustainment.
- Experience with Linux administration, scripting, and troubleshooting distributed applications in production environments.
- Experience with containerized systems and orchestration platforms such as Kubernetes.
- Experience supporting CI/CD pipelines, release automation, infrastructure-as-code, and operational reliability in Agile or DevSecOps environments.
- Experience with monitoring, logging, and alerting tools used to support enterprise application performance and infrastructure visibility.
- Strong analytical, troubleshooting, documentation, and communication skills, with the ability to translate operational issues into engineering improvements.
- Ability to work effectively across cross-functional teams in a mission-focused DoD environment.
- Experience supporting AWS Cloud One or other secure federal cloud environments.
- Experience supporting geospatial or Esri-based platforms, including ArcGIS Enterprise or related technologies.
- Familiarity with service reliability practices such as SLIs, SLOs, error budgets, incident postmortems, and capacity planning.
- Experience with Risk Management Framework (RMF), STIG compliance, vulnerability remediation, and secure system hardening practices.
- AWS, Kubernetes, or other relevant cloud or reliability engineering certifications.
- Experience supporting technical refresh, platform modernization, or high-availability design initiatives in enterprise environments.
Diné Development Corporation (DDC) is a Navajo Nation owned family of companies that provides government agencies and commercial organizations with high-quality IT, professional, environmental, and research and development services. DDC is dedicated to empowering the Navajo Nation and communities we serve. Benefits
Eligible full-time employees receive a comprehensive benefits package, including medical, dental, vision, life and disability coverage, retirement savings with company match, paid time off, voluntary supplemental benefits, and access to an employee assistance program. The package also includes educational assistance, with tuition reimbursement. EEO Statement
This contractor and subcontractor shall abide by the requirements of 41 CFR 60-1.4(a), 60-300.5(a), and 60-741.5(a). These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity, national origin, or for inquiring about, discussing, or disclosing information about compensation, or any other basis prohibited by law. We participate in E-Verify.
- ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview-The ProCOM team is looking for a Site Reliability Engineering (SRE) who can help us solve problems, build our...SuggestedFull timePart timeImmediate startWorldwideFlexible hours
$96k - $163k
...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that...SuggestedFull timePart timeWorldwideFlexible hours$76k - $127k
...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that...SuggestedFull timePart timeWorldwideFlexible hours$96k - $163k
...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview The BizOps team is looking for a Senior Site Reliability Engineer who can help us solve problems and...SuggestedFull timePart timeWorldwideFlexible hoursShift work- ...and performance our customers have come to expect, and help raise the reliability bar as we grow. What you would do: Design, build, and operate the shared platform foundations engineers ship on every day: GCP infrastructure, Kubernetes, networking, routing,...SuggestedRemote workWorldwideFlexible hours
$147k - $168k
...Inc. as one of the most innovative and fastest-growing technology companies in the country. Role Summary As a Site Reliability Engineer at Filevine, you will improve the reliability, scalability, and operational maturity of the Filevine platform. You’ll...Full timeTemporary workWork experience placementWork at officeRemote work2 days per week3 days per week- ...Site Reliability Engineer Company: Milestone Systems Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Golang, Python, Linux, Shell scripting, Kubernetes, Docker, Terraform, CI/CD, GitOps, ArgoCD, Spinnaker, Prometheus, Datadog,...Full timeRemote work
- ...in Cupertino, California, invites an experienced CDN Solutions Engineer to join the Content Delivery Network Solutions team. You will... ...and collaborate with engineering groups across Apple to ensure reliable delivery at scale. The ideal candidate has 4+ years in CDNs and...
- ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation...
- ...provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their multi-... ...role involves leading programs in partnership with senior engineering leaders, requiring over 10 years of experience in cloud infrastructure...
- ...Motor Company seeks a Director of Cloud SRE to lead a team of engineering leaders and engineers, federating core SRE principles across... ...You will guide cross-domain collaboration, drive AI-enabled reliability, and ensure CI/CD integration while expanding reliability into...Remote work
- ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available... ...and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and infrastructure...
- ...ServiceNow in Santa Clara, CA, seeks a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, resilience, and toil... ...for global engineering teams. Embedded within the Site Reliability & Database Engineering organization, you will architect SRE tooling...
- ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage...
- ...Jobtailor is seeking an experienced Platform Architect/Lead to shape major cloud platform decisions and drive reliability across enterprise-scale workloads. You will own design, implementation, and ongoing improvements for critical systems, with heavy emphasis on observability...
$276.1k - $311.4k
...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a...Permanent employmentFull timeWork at officeWork from home- ...automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and...Local area
- ...Job Description The Opportunity: Versant's Sports & Entertainment Digital Products division is seeking a Senior Site Reliability Engineer to help drive the reliability, scalability, and usability of internal developer platforms, tooling, and engineering workflows...Local areaRemote workWorldwide
- ...Sight Machine is seeking a senior Cloud Infrastructure IC to lead reliability, automation, and scale across our platform. You will drive IaC... .../CD, observability and operate agentic AI systems, mentoring engineers and guiding architectural decisions while staying hands-on...
- ...exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of... ...and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase...
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business...
$145k - $175k
...more out of AI with trusted data products, a powerful analytics engine, and AI agents. This helps teams reduce risk and operating... ...use, maintaining flexibility without lock-in. The Senior Site Reliability Engineer Role Join our dynamic team at Qlik as a Senior...Temporary workImmediate startRemote workFlexible hoursRotating shift$152k - $195k
...including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/CD systems...Remote work$100k - $180k
...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$115k - $160k
...with Barclays to connect them with exceptional professionals for this role. Embark on a transformative journey as a Senior Site Reliability Engineer - AVP - Credit Trade Floor. At Barclays, our vision is clear –to redefine the future of banking and help craft innovative...Hourly payWork at office$194k - $237k
...employer, at the date of hire. This position is ineligible for employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours- ...ByteDance’s Infrastructure Engineering team in Seattle designs, builds, and operates global infrastructure spanning public and private... ...storage. Join a fast-paced, collaborative team focused on reliability, scalability, and continuous optimization, driving improvements...
$125k - $250k
...we are reimagining how developers build reliable, scalable, event-driven applications without... ...possible Partner closely with engineering teams to improve system resiliency and scalability... ...For 5+ years of experience in Site Reliability Engineering, DevOps,...Full timeImmediate startRemote workFlexible hours- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering... ...You will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes...
- ...Eyes on glass. Hands on the pipeline. Real ownership from day one. This isn't a watch-and-wait monitoring seat. Our client needs engineers who can read a Kibana query at 3am, know the difference between a blip and a breach, and act on it, on a FedRAMP-authorised cloud...Hourly payFor contractorsRemote workShift workNight shiftWeekend work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre United States
- site reliability engineering manager United States
- site reliability engineer United States
- site reliability engineer remote United States
- site recruiter United States
- site services specialist United States
- junior website developer United States
- official site United States
- on site coordinator United States
- site leader United States

