Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

VDart Inc

Job Title: Site Reliability Engineer

Location: Bellevue, WA & Frisco, TX

Duration: / Term: Contract

Experience Desired: 7+ Years

Job Description:

This role serves as a vendor-provided Site Reliability Engineer responsible for improving and protecting the reliability, scalability, and performance of a platform currently hosted on Microsoft Azure that will be migrated to Tencent Kubernetes Engine (TKE) in the future. It manages availability, latency, performance, security, and capacity while enabling efficient, automated software delivery across both the current Azure environment and the upcoming TKE platform. The role differentiates by combining deep Azure cloud infrastructure expertise with Kubernetes-based container orchestration, positioning the team for a smooth cloud-to-TKE migration. Success is measured by improved system uptime, faster incident resolution, and a reliable, well-supported migration path. The work directly impacts IT service quality, operational resilience, and customer experience through the transition and beyond.

Main Responsibilities

  • Monitor, troubleshoot, and resolve incidents affecting availability, latency, and performance of current Azure-hosted workloads
  • Provision, configure, and manage Azure infrastructure (VMs, networking, storage, IAM) to support production and non-production environments
  • Support planning and execution of the platform's migration from Azure to Tencent Kubernetes Engine (TKE), including workload containerization and cutover activities
  • Design, build, and maintain CI/CD pipelines that support automated deployment and testing today on Azure and going forward on TKE
  • Build and maintain observability tooling - dashboards, alerts, logging, and health checks - to proactively identify and address system risks across both environments
  • Drive automation and infrastructure-as-code practices to reduce manual toil and improve deployment consistency
  • Collaborate with internal engineering teams and stakeholders to support incident response, capacity planning, and migration readiness

Required Qualifications

Education: Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.

Experience: 4+ years in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, including hands-on production experience with Microsoft Azure (compute, networking, storage, IAM). Experience with Tencent Kubernetes Engine (TKE) or comparable Kubernetes platforms strongly preferred, as the platform will migrate to TKE.

Technical Skills: Proficiency with CI/CD tooling (e.g., Azure DevOps, Jenkins, GitLab CI), containerization and orchestration (Kubernetes, Docker, TKE), infrastructure-as-code (Terraform, ARM/Bicep), scripting/automation (Python, Bash, PowerShell), and monitoring/observability platforms (e.g., Grafana, Prometheus, Azure Monitor).

Other: Strong troubleshooting and incident-response skills; experience supporting cloud platform migrations a plus; ability to work effectively as an embedded vendor resource within a client engineering team; on-call availability as required.

Preferred Qualifications

Azure certifications (e.g., AZ-104, AZ-400, AZ-500) and/or Kubernetes certifications (CKA/CKAD). Prior experience migrating workloads from a cloud VM-based platform to Kubernetes/TKE, including containerization of legacy services.

Key Skills:

SRE, Azure, Kubernetes, Terraform, observability

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Washington DC vacancy
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment... 
    Suggested
    Work experience placement

    Samprasoft

    Washington DC
    5 days ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Suggested
    Work experience placement
    Work at office

    Akamai

    Washington DC
    3 days ago
  •  ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington D.C. area, focusing on optimizing the availability, performance, and scalability of critical production services. The ideal... 
    Suggested

    Software Technology Inc

    Washington DC
    4 days ago
  •  ...Site Reliability Engineer ValidaTek is building teams of Site Reliability Engineers (SRE's) to support internal and external engineering and operations of a large scale and world-wide Enterprise IT environment that covers application hosting and support, enterprise... 
    Suggested

    ClearanceJobs

    Washington DC
    4 days ago
  •  ...solutions using a tailored Agile methodology. We are seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position will be a remote role open to US citizens residing in the... 
    Suggested
    Remote work

    Elevate Government Solutions

    Washington DC
    1 day ago
  • $230k - $250k

     ...Arlington, Virginia Secret Hybrid schedule Information Technology Overview GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining... 
    Full time
    Remote work
    Flexible hours

    GovCIO

    Arlington, VA
    3 days ago
  •  ...Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient, scalable, and highly performant... 
    Local area

    Tiger Analytics

    Washington DC
    2 days ago
  •  ...Catalyst, and In-Q-Tel. Mission | On Site | Full Time | Active TS/SCI with Full...  ...government customer site, ensuring the reliability and performance of Twenty's mission-critical...  ...technical ownership and customer-facing engineering: you'll define how we measure... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Technologies

    Arlington, VA
    2 days ago
  • $107k - $220k

     ...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track Key Performance Indicators (KPIs) and Service Level Objectives (SLOs), identify and resolve performance bottlenecks, and perform... 
    Full time
    Contract work
    Temporary work
    Work at office
    Visa sponsorship
    Work visa

    Avalore, LLC

    Arlington, VA
    2 days ago
  • $106.3k - $221.1k

     ...Senior Site Reliability Engineer At Accenture Federal Services, nothing matters more than helping the US federal government make the nation stronger and safer and life better for people. Our 13,000+ people are united in a shared purpose to pursue the limitless potential... 

    Accenture Federal Services

    Arlington, VA
    4 days ago
  • $112k - $218.4k

     ...: Our team is looking for a Senior Active Directory Site Reliability Engineer. Our mission is to improve the availability, latency, performance and security of the Identity systems behind Microsoft's cloud. Like traditional operations, we keep important revenue-critical... 
    Full time
    Local area

    Microsoft

    Washington DC
    6 days ago
  • $135k - $155k

     ...Responsibilities Own assigned infrastructure and software reliability problem statements through completion with guidance from senior engineers. Write, improve, and document efficient code and existing systems. Develop, configure, and maintain AWS accounts,... 
    Full time

    Xometry

    Bethesda, MD
    11 days ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms,... 
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    18 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering , this role will... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    6 days ago
  •  ...safety and security. Make an impact by using your expertise to protect our country from threats. Job Description SITE RELIABILITY ENGINEER (SRE) YOUR IMPACT Own your opportunity to support national defense. Your work will help keep critical operations secure... 

    General Dynamics Information Technology

    Washington DC
    more than 2 months ago
  •  ...Job Description Job Description Site Reliability Engineer II Metro DC · Hybrid · 24/7 FedRAMP Operations · Rotational Shift · Initial Contract till March 27. KEY REQUIREMENT This role requires US citizenship and residence on US soil. It sits within a FedRAMP... 
    Hourly pay
    Contract work
    For contractors
    Shift work
    Night shift
    Weekend work

    C-Serv

    Arlington, VA
    10 days ago
  •  ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical leader responsible for driving the operational excellence, resilience, and security of our core systems for a key Randstad client in the Washington D.C. area. This... 

    Software Technology Inc

    Washington DC
    4 days ago
  •  ...TENEX Staff Site Reliability Engineer TENEX is an AI-native, automation-first, built-for-scale Managed Detection and Response (MDR) provider. We are a force multiplier for defenders, helping organizations enhance their cybersecurity posture through advanced threat detection... 
    Work from home

    TenEx

    Washington DC
    3 days ago
  • $160k - $210k

     ...change and achieving remarkable growth in a rapidly evolving industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale... 
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Washington DC
    10 days ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Staff Site Reliability Engineer Opportunity Okta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    6 days ago
  • $207k - $284.9k

     ...is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    5 days ago
  • $182k - $250.8k

     ...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    5 days ago
  • $174k - $238k

     ...We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable,... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    5 days ago
  • $204k - $306k

     ...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Manager, Site Reliability Engineering San Francisco, California Secure Every Identity, from AI to Human Identity is the key to unlocking the potential... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Washington DC
    5 days ago
  •  ...We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs. The InfraSec team collaborates... 
    Full time
    Remote work
    Worldwide

    Mongodb

    Washington DC
    3 days ago
  • $151k - $297k

     ...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB's cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Washington DC
    4 days ago
  • $147k - $202.4k

     ...-defining work. We're all in on this mission. If you are too, let's talk. Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta

    Washington DC
    5 days ago
  • $128.5k - $190k

     ...Senior Site Reliability Engineer Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees... 
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    1 day ago
  •  ...is a recognized, award-winning leader in supply chain AI and a FedRAMP authorized provider to the federal government. Site Reliability Engineer Location: U.S. (Hybrid) This role requires U.S. citizenship and eligibility for a U.S. security clearance.... 
    Work at office
    Work from home
    Flexible hours

    Exiger

    McLean, VA
    5 days ago
  •  ...SRE Engineer Location: Washington, DC (Onsite) Duration: 08-17-2026 - 07-30-2027 Key Responsibilities Observability & Monitoring...  ...(RCA), and author comprehensive knowledge base articles. Reliability Engineering: Champion SRE metrics including Service Level... 

    Georgia IT Inc

    Washington DC
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!