Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Diné Development

The Site Reliability Engineer (SRE) / Subject Matter Expert (SME) – Computer Systems Engineer/Architect will provide senior-level reach-back expertise to support the reliability, scalability, performance, and operational resilience of the GEOMAP platform in secure cloud environments. This role focuses on improving service availability, monitoring, incident response, automation, and production stability across cloud-hosted and containerized systems supporting mission-critical geospatial capabilities for the U.S. Air Force.The Site Reliability Engineer will collaborate across development, DevSecOps, cloud, database, testing, and support teams to identify systemic issues, reduce operational risk, and implement engineering solutions that improve long-term platform reliability. *This position is contingent upon contract award.* Responsibilities

  • Provide senior-level engineering support to improve reliability, availability, performance, and maintainability of GEOMAP cloud-hosted systems and services.
  • Analyze production issues, recurring incidents, and operational trends to identify root causes and recommend durable corrective actions.
  • Support the design and implementation of monitoring, alerting, logging, and observability solutions across applications, infrastructure, and containerized services.
  • Develop and recommend automation approaches that reduce manual effort, improve deployment consistency, and increase system resilience.
  • Partner with software engineers, DevSecOps engineers, Kubernetes engineers, database engineers, and production support personnel to improve service health and release readiness.
  • Support incident response, problem management, service restoration, and post-incident reviews for high-priority operational issues.
  • Evaluate system performance, capacity, and scalability needs and provide recommendations for optimization and operational risk reduction.
  • Assist in defining service reliability objectives, operational metrics, and support models for sustained mission operations.
  • Contribute to infrastructure and platform engineering efforts involving cloud environments, CI/CD pipelines, container orchestration, and secure deployment patterns.
  • Support architecture reviews, technical assessments, and engineering analyses related to reliability, recoverability, and production operations.
  • Develop or refine runbooks, standard operating procedures, reliability engineering practices, and technical documentation.
  • Provide reach-back support for surge requirements, complex production investigations, and priority modernization or stabilization efforts as directed.
  • Performs other related duties as assigned.
Qualifications
  • Active Secret clearance required.
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field; Master’s degree preferred.
  • Minimum of 8 years of experience supporting enterprise systems, cloud platforms, site reliability engineering, production engineering, systems engineering, or related technical roles.
  • Experience supporting AWS environments, including monitoring, performance tuning, troubleshooting, incident response, and operational sustainment.
  • Experience with Linux administration, scripting, and troubleshooting distributed applications in production environments.
  • Experience with containerized systems and orchestration platforms such as Kubernetes.
  • Experience supporting CI/CD pipelines, release automation, infrastructure-as-code, and operational reliability in Agile or DevSecOps environments.
  • Experience with monitoring, logging, and alerting tools used to support enterprise application performance and infrastructure visibility.
  • Strong analytical, troubleshooting, documentation, and communication skills, with the ability to translate operational issues into engineering improvements.
  • Ability to work effectively across cross-functional teams in a mission-focused DoD environment.
Preferred
  • Experience supporting AWS Cloud One or other secure federal cloud environments.
  • Experience supporting geospatial or Esri-based platforms, including ArcGIS Enterprise or related technologies.
  • Familiarity with service reliability practices such as SLIs, SLOs, error budgets, incident postmortems, and capacity planning.
  • Experience with Risk Management Framework (RMF), STIG compliance, vulnerability remediation, and secure system hardening practices.
  • AWS, Kubernetes, or other relevant cloud or reliability engineering certifications.
  • Experience supporting technical refresh, platform modernization, or high-availability design initiatives in enterprise environments.
About Us
Diné Development Corporation (DDC) is a Navajo Nation owned family of companies that provides government agencies and commercial organizations with high-quality IT, professional, environmental, and research and development services. DDC is dedicated to empowering the Navajo Nation and communities we serve. Benefits
Eligible full-time employees receive a comprehensive benefits package, including medical, dental, vision, life and disability coverage, retirement savings with company match, paid time off, voluntary supplemental benefits, and access to an employee assistance program. The package also includes educational assistance, with tuition reimbursement. EEO Statement
This contractor and subcontractor shall abide by the requirements of 41 CFR 60-1.4(a), 60-300.5(a), and 60-741.5(a). These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity, national origin, or for inquiring about, discussing, or disclosing information about compensation, or any other basis prohibited by law. We participate in E-Verify.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in United States vacancy
  •  ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview-The ProCOM team is looking for a Site Reliability Engineering (SRE) who can help us solve problems, build our... 
    Suggested
    Full time
    Part time
    Immediate start
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    20 hours ago
  • $96k - $163k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    20 hours ago
  • $76k - $127k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    20 hours ago
  • $96k - $163k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview The BizOps team is looking for a Senior Site Reliability Engineer who can help us solve problems and... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours
    Shift work

    Mastercard

    O Fallon, MO
    20 hours ago
  •  ...and performance our customers have come to expect, and help raise the reliability bar as we grow. What you would do: Design, build, and operate the shared platform foundations engineers ship on every day: GCP infrastructure, Kubernetes, networking, routing,... 
    Suggested
    Remote work
    Worldwide
    Flexible hours

    Sanity

    United States
    2 days ago
  • $147k - $168k

     ...Inc. as one of the most innovative and fastest-growing technology companies in the country. Role Summary As a Site Reliability Engineer at Filevine, you will improve the reliability, scalability, and operational maturity of the Filevine platform. You’ll... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    2 days per week
    3 days per week

    Filevine

    United States
    2 days ago
  •  ...Site Reliability Engineer Company: Milestone Systems Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Golang, Python, Linux, Shell scripting, Kubernetes, Docker, Terraform, CI/CD, GitOps, ArgoCD, Spinnaker, Prometheus, Datadog,... 
    Full time
    Remote work

    Milestone Systems Inc

    United States
    2 days ago
  •  ...in Cupertino, California, invites an experienced CDN Solutions Engineer to join the Content Delivery Network Solutions team. You will...  ...and collaborate with engineering groups across Apple to ensure reliable delivery at scale. The ideal candidate has 4+ years in CDNs and... 

    Apple

    Cupertino, CA
    1 day ago
  •  ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation... 

    Oracle

    Santa Clara, CA
    4 days ago
  •  ...provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their multi-...  ...role involves leading programs in partnership with senior engineering leaders, requiring over 10 years of experience in cloud infrastructure... 

    Menlo Ventures

    Bellevue, WA
    20 hours ago
  •  ...Motor Company seeks a Director of Cloud SRE to lead a team of engineering leaders and engineers, federating core SRE principles across...  ...You will guide cross-domain collaboration, drive AI-enabled reliability, and ensure CI/CD integration while expanding reliability into... 
    Remote work

    Ford Motor Company

    United States
    1 day ago
  •  ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available...  ...and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and infrastructure... 

    Neshent Technologies

    Los Gatos, CA
    4 days ago
  •  ...ServiceNow in Santa Clara, CA, seeks a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, resilience, and toil...  ...for global engineering teams. Embedded within the Site Reliability & Database Engineering organization, you will architect SRE tooling... 

    ServiceNow

    Santa Clara, CA
    4 days ago
  •  ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage... 

    Lambda

    San Francisco, CA
    1 day ago
  •  ...Jobtailor is seeking an experienced Platform Architect/Lead to shape major cloud platform decisions and drive reliability across enterprise-scale workloads. You will own design, implementation, and ongoing improvements for critical systems, with heavy emphasis on observability... 

    Jobtailor

    Florida, NY
    4 days ago
  • $276.1k - $311.4k

     ...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a... 
    Permanent employment
    Full time
    Work at office
    Work from home

    Lindus Health

    Sunnyvale, CA
    3 days ago
  •  ...automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action. Your CRDs are the schema the platform's predictors and... 
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    4 days ago
  •  ...Job Description The Opportunity: Versant's Sports & Entertainment Digital Products division is seeking a Senior Site Reliability Engineer to help drive the reliability, scalability, and usability of internal developer platforms, tooling, and engineering workflows... 
    Local area
    Remote work
    Worldwide

    Versant

    United States
    2 days ago
  •  ...Sight Machine is seeking a senior Cloud Infrastructure IC to lead reliability, automation, and scale across our platform. You will drive IaC...  .../CD, observability and operate agentic AI systems, mentoring engineers and guiding architectural decisions while staying hands-on... 

    Jobless

    Ann Arbor, MI
    1 day ago
  •  ...exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of...  ...and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase... 

    J.P. Morgan

    Palo Alto, CA
    20 hours ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business... 

    J.P. Morgan

    San Francisco, CA
    4 days ago
  • $145k - $175k

     ...more out of AI with trusted data products, a powerful analytics engine, and AI agents. This helps teams reduce risk and operating...  ...use, maintaining flexibility without lock-in. The Senior Site Reliability Engineer Role Join our dynamic team at Qlik as a Senior... 
    Temporary work
    Immediate start
    Remote work
    Flexible hours
    Rotating shift

    Qlik

    United States
    1 day ago
  • $152k - $195k

     ...including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/CD systems... 
    Remote work

    SecurityScorecard

    United States
    3 days ago
  • $100k - $180k

     ...Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    United States
    2 days ago
  • $115k - $160k

     ...with Barclays to connect them with exceptional professionals for this role. Embark on a transformative journey as a Senior Site Reliability Engineer - AVP - Credit Trade Floor. At Barclays, our vision is clear –to redefine the future of banking and help craft innovative... 
    Hourly pay
    Work at office

    Barclays

    New York, NY
    2 days ago
  • $194k - $237k

     ...employer, at the date of hire. This position is ineligible for employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    San Francisco, CA
    1 day ago
  •  ...ByteDance’s Infrastructure Engineering team in Seattle designs, builds, and operates global infrastructure spanning public and private...  ...storage. Join a fast-paced, collaborative team focused on reliability, scalability, and continuous optimization, driving improvements... 

    ByteDance

    Seattle, WA
    20 hours ago
  • $125k - $250k

     ...we are reimagining how developers build reliable, scalable, event-driven applications without...  ...possible Partner closely with engineering teams to improve system resiliency and scalability...  ...For 5+ years of experience in Site Reliability Engineering, DevOps,... 
    Full time
    Immediate start
    Remote work
    Flexible hours

    Orkes

    United States
    2 days ago
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering...  ...You will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes... 

    Socket

    San Francisco, CA
    2 days ago
  •  ...Eyes on glass. Hands on the pipeline. Real ownership from day one. This isn't a watch-and-wait monitoring seat. Our client needs engineers who can read a Kibana query at 3am, know the difference between a blip and a breach, and act on it, on a FedRAMP-authorised cloud... 
    Hourly pay
    For contractors
    Remote work
    Shift work
    Night shift
    Weekend work

    C-Serv

    United States
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!