Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (Azure + Data)

VDart Inc

Job Title: Site Reliability Engineer (Azure + Data)

Work Mode: Bellevue, WA - Hybrid

Contract

Skills

Mandatory Skills : Azure Infra Services, Azure Monitor, Azure Storage
Good to Have Skills : Azure Data Factory

Job Description :

  • 6 years of experience in infrastructure support site reliability engineering cloud operations or platform engineering including strong hands-on ownership of production Azure environments
  • Demonstrated expertise in reliability performance capacity cost optimization and high severity incident leadership
  • Deep knowledge of incident problem change security privacy compliance audit and governance processes
  • Strong experience with Azure access control certificates secrets monitoring logging CICD Infrastructure as Code and scripting
  • Solid understanding of data platform architecture and dependencies across ADF ADLS Gen2 Synapse Cosmos DB Azure Data Explorer SQL Server and Microsoft Fabric
  • Strong production ownership customer facing communication and the ability to drive operational discipline across onsite and offshore teams

Key Responsibilities

  • Infrastructure reliability and optimization Own the availability performance capacity and cost of production Azure infrastructure supporting high volume data platforms target 9999 availability SLA adherence and QoS while proactively addressing bottlenecks saturation and scaling risks
  • Azure architecture security and governance Design and operate landing zones subscriptions resource groups VNets peering ExpressRoute Azure Firewall Bastion DDoS protection Azure Policy Entra ID RBAC Managed Identities PIM and Conditional Access
  • Incident problem and change management Lead triage mitigation stakeholder communication root cause analysis corrective actions risk assessment approvals validation and rollback planning improve MTTR and prevent repeat incidents
  • Compliance and operational readiness Maintain S360 security privacy audit and governance compliance sustain accurate runbooks SOPs CENs and operational playbooks
  • Certificates secrets and dependencies Manage the lifecycle of certificates keys secrets identities and service dependencies track expirations automate renewals and secure service to service communication
  • Data platform infrastructure Support and optimize Azure Data Factory ADLS Gen2 Synapse Analytics Cosmos DB Azure Data Explorer SQL Server and Microsoft Fabric troubleshoot throughput and dependency issues and guide platform modernization and Fabric migration
  • Monitoring and observability Use Azure Monitor Log Analytics Application Insights and KQL platform metrics ing and cost dashboards to detect risks analyze trends and drive evidence based decisions
  • Automation and engineering practices Standardize infrastructure and operational workflows through Azure DevOps GitHub Actions ARM Bicep Terraform YAML PowerShell Azure CLI Python and Power Automate apply SRE practices including error budgets automated recovery and self-healing
  • Cost and performance management Right size resources optimize compute storage networking SQL and Cosmos capacity and implement budgets and reservation planning without compromising reliability
  • Customer and team collaboration Serve as the infrastructure SRE contact for Redmond customers communicate risks and optimization opportunities clearly and coordinate consistent execution across onsite offshore infrastructure SRE and data engineering teams

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (Azure + Data) in Washington DC vacancy
  • $175k - $200k

     ...Title : Site Reliability Engineer Location : Arlington, VA Hybrid : 3 days on-site / 2 days remote Contract : 6 month to perm Clearance...  ...Skills and Experience: Kubernetes certifications AWS/Azure certifications DevOps certifications ITIL preferred... 
    Suggested
    Permanent employment
    Contract work
    Remote work

    Insight Global

    Arlington, VA
    13 hours ago
  •  ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry...  ...Experience with multi-cloud environments (AWS, GCP, Azure). Chaos engineering experience (Gremlin, Chaos Mesh, or... 
    Suggested
    Full time
    Work at office

    TalentDome Staffing

    Washington DC
    2 days ago
  •  ...Lead Site Reliability Engineer Defense Tech / National Security US Defense Tech Startup The Company Early-stage defense technology leader building modern software for air-gapped, high-side, and accredited environments. Backed by a $99M defense contract to... 
    Suggested
    Full time
    Contract work

    Attis

    Washington DC
    2 days ago
  •  ...day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance...  ...of hands-on Cloud Operations and Site Reliability Engineering, operating production...  ...~ A second cloud (Google Cloud or Azure) is a plus, not a substitute. We... 
    Suggested
    Full time

    MangoApps

    Washington DC
    2 days ago
  •  ...capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE...  ...of customers across AWS, GCP, and Azure. You’ll work on everything from multi-...  ..., maintain, and optimize multi-region data and caching layers — including PostgreSQL... 
    Suggested

    Kong

    Washington DC
    4 days ago
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient...  ...• Develop and maintain Terraform modules for AWS and Azure environments• Create and manage Ansible playbooks for configuration... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    13 hours ago
  • $166k - $220k

     ...system that turns thousands of data streams into a realtime, 3D...  ...expectations. Our systems integration engineers internalize the nuances of...  ...THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD,...  ...cloud deployments in AWS, Azure and on premise. This role emphasizes... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    2 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir...  ...’s leading software for data-driven decisions and...  ...more.The RoleWe’re looking for Site Reliability Engineers who can help us...  ...hosting platforms like AWS, Azure, or GCP and/or experience with... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    13 hours ago
  • $112k - $179k

     ...About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in...  ...DevOps, or operations experience.Cloud Certifications (AWS, Azure, or equivalent).CISSP or CASP+DoD 8140/8570 Compliant.Strong... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    3 days ago
  • $135k - $154k

     ...ImpactAs a contributor in the APX platform engineering organization on the CloudNet team, you...  ...about achieving the high quality and reliability our customers demand. You will work closely...  ...-as-code on public cloud (Axon uses Azure, but other experience is fine)Experience... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    13 hours ago
  • $125k - $185k

     ...Changing CompanyPalantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data...  ..., and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    3 days ago
  • $165k - $270k

     ...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology...  ...techniquesUnderstanding of distributed databases and data modelingExperience with automatically managing dozens, hundreds... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    13 hours ago
  • $174k - $238k

     ...The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products...  ...sharing.Influence architecture and operational decisions through data-driven recommendations and engineering expertise.Drive... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    1 day ago
  • $174k - $239k

     ...If you are too, let's talk.Okta’s Technology, Data & Intelligence (TDI) team delivers the...  ...we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    3 days ago
  • $182k - $250.8k

     ...backbone of our platform's reliability and operational...  ...forward-thinking group of engineers and leaders who believe...  ...of continuous learning, data-driven decision-making,...  ...worldwide. As a Manager, Site Reliability Engineer,...  ...cloud platforms (AWS, Azure) and infrastructure as... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    4 days ago
  • $207k - $284.9k

     ...you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from...  ...too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports...  .../or have access to protected federal data. As a condition of employment for this... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    13 hours ago
  • $130k - $180k

    Job DescriptionEverforth ECS is seeking a Cloud Site Reliability Engineer (SRE)to work in our Arlington, VA office/remotely. Our Philosophy We believe...  ...cause after Uptime Goals & Reliability Define reasonable, data-driven SLOs and error budgets for critical services... 
    For contractors
    Work at office
    Remote work
    Shift work

    ECS Federal

    Arlington, VA
    3 days ago
  • $141.8k - $195k

    B2B SAAS data observability software. Join the company that's building the telemetry...  ...RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock...  ...of cloud platforms (prefer AWS and Azure) and container + orchestration technologies... 
    Remote work

    Cribl

    Washington DC
    1 day ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Remote work

    Noctua Technology

    Washington DC
    4 days ago
  •  ..., District of Columbia, United States About the job Sr. Site Reliability Engineer Our Client is currently hiring a full-time Sr. Site Reliability...  ...CI/CD pipeline solutions such as Team Foundation Server/Azure DevOps, Bitbucket, and GitHub. Familiarity with... 
    Full time
    Currently hiring
    3 days per week

    CruitZi

    Washington DC
    3 days ago
  • $120k - $140k

     ...Pythian, a multinational company, was founded in 1997 and started by ensuring the reliability and performance of mission-critical databases. We quickly earned a reputation for solving tough data challenges. We were there when the industry moved from on-premises to cloud... 
    Work from home

    Pythian

    Washington DC
    2 days ago
  •  ...Site Reliability Engineer III (AI Platform) Location: Mount Laurel, NJ (Onsite) Duration: Contract Experience: 4+ years About the...  ...Technical Skills Cloud Platforms AWS (required) Azure or GCP Containers & Orchestration Kubernetes (... 
    Contract work

    GCS Recruitment

    Laurel, MD
    1 day ago
  •  ...issues in Containerization, Docker, Kubernetes, AWS, PCF, Azure Analysis of issues via APM, NMON , Wireshark usage and analysis...  ...job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago... 
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    2 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or...  ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  • $10k

     ...Site Reliability Engineer (SRE) Skill Level 3 Wyetech is seeking an experienced Site Reliability Engineer Level 3 (SRE3) to work within a cloud environment supporting data-intensive analytics on a managed infrastructure. The position will work with technologies including... 
    Hourly pay
    Full time
    Contract work
    Temporary work
    Work experience placement
    Summer work
    Immediate start

    Wyetech LLC

    Laurel, MD
    1 day ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    1 day ago
  •  ...Join the Site Reliability Engineering (SRE) team to support high-engagement multimodal applications. You will be responsible for managing infrastructure...  ...of security groups and application exposure management. Familiarity with GCP or Azure is a plus. #J-18808-Ljbffr

    NextGen | GTA: A Kelly Telecom Company

    Laurel, MD
    3 days ago
  •  ...and agents. Role Overview We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is...  ...of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-... 

    GMI Cloud

    Arlington, VA
    19 hours ago
  •  ...seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The...  ...infrastructure and application alerting and diagnostics Optimize data flows and storage integrations Collaborate with... 
    Remote work

    Elevate Government Solutions

    Washington DC
    4 days ago
  • $106.3k - $221.1k

     ...missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability...  ...of the following fields: Computer Science, Cybersecurity, Data Science, Information Systems, Information Technology, or Software... 
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (Azure + Data). Be the first to apply!