Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GMI Cloud

About GMI

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities

  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  • Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  • Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  • Regional/international travel to GMI data center locations.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  • Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  • Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  • Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Vacancy posted 10 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Washington DC vacancy
  • $175k - $200k

     ...Title : Site Reliability Engineer Location : Arlington, VA Hybrid : 3 days on-site / 2 days remote Contract : 6 month to perm Clearance : Secret Salary : $175K-$200K Our client is hiring a Site Reliability Engineer with an active Secret clearance... 
    Suggested
    Permanent employment
    Contract work
    Remote work

    Insight Global

    Arlington, VA
    3 hours ago
  •  ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry: High-Growth Technology / SaaS About the Role We are seeking a highly skilled Senior Site Reliability Engineer to... 
    Suggested
    Full time
    Work at office

    TalentDome Staffing

    Washington DC
    2 days ago
  •  ...Lead Site Reliability Engineer Defense Tech / National Security US Defense Tech Startup The Company Early-stage defense technology leader building modern software for air-gapped, high-side, and accredited environments. Backed by a $99M defense contract to... 
    Suggested
    Full time
    Contract work

    Attis

    Washington DC
    2 days ago
  •  ...customers depend on every day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance of that platform...  ...Who you are: ~5+ years of hands-on Cloud Operations and Site Reliability Engineering, operating production-scale SaaS (... 
    Suggested
    Full time

    MangoApps

    Washington DC
    2 days ago
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Suggested
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    5 days ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 

    Kong

    Washington DC
    4 days ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    2 days ago
  • $165k - $270k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    5 days ago
  • $112k - $179k

     ...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    3 days ago
  • $135k - $154k

     ...where you matter.Your ImpactAs a contributor in the APX platform engineering organization on the CloudNet team, you are passionate about...  .... You are also obsessed about achieving the high quality and reliability our customers demand. You will work closely with sovereign... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    5 days ago
  • $141.8k - $195k

     ...their best work, grow fast, and bring their full selves to the herd.Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl... 
    Remote work

    Cribl

    Washington DC
    1 day ago
  •  ...Site Reliability Engineer III (AI Platform) Location: Mount Laurel, NJ (Onsite) Duration: Contract Experience: 4+ years About the Role We are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible... 
    Contract work

    GCS Recruitment

    Laurel, MD
    1 day ago
  •  ...plan Paid maternity leave 401(k) Get notified when a new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle, WA $115,000.00-$175,000.00 5 months ago Senior ServiceNow... 
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    2 days ago
  •  ...Washington D.C., District of Columbia, United States About the job Sr. Site Reliability Engineer Our Client is currently hiring a full-time Sr. Site Reliability Engineer (SRE), who will play a vital role in continuously driving improvements in observability, performance... 
    Full time
    Currently hiring
    3 days per week

    CruitZi

    Washington DC
    3 days ago
  • $120k - $140k

     ...digital transformation and operational excellence. Pythian, a multinational company, was founded in 1997 and started by ensuring the reliability and performance of mission-critical databases. We quickly earned a reputation for solving tough data challenges. We were there... 
    Work from home

    Pythian

    Washington DC
    2 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or...  ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  •  ...on one unified cloud. One cloud for compute, inference, and agents. Role Overview We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and... 

    GMI Cloud

    Arlington, VA
    10 hours ago
  •  ...Join the Site Reliability Engineering (SRE) team to support high-engagement multimodal applications. You will be responsible for managing infrastructure, observability solutions, and platform automation while ensuring the reliability and security of hosted applications... 

    NextGen | GTA: A Kelly Telecom Company

    Laurel, MD
    3 days ago
  • $10k

     ...Site Reliability Engineer (SRE) Skill Level 3 Wyetech is seeking an experienced Site Reliability Engineer Level 3 (SRE3) to work within a cloud environment supporting data-intensive analytics on a managed infrastructure. The position will work with technologies including... 
    Hourly pay
    Full time
    Contract work
    Temporary work
    Work experience placement
    Summer work
    Immediate start

    Wyetech LLC

    Laurel, MD
    1 day ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    1 day ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Full time
    Flexible hours

    Proofpoint

    Alexandria, VA
    1 day ago
  • $121.4k - $218.6k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    1 day ago
  •  ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington D.C. area, focusing on optimizing the availability, performance, and scalability of critical production services. The ideal... 

    Software Technology Inc

    Washington DC
    2 days ago
  • $130k - $160k

     ...customers. Our customers and partners trust us to deliver reliable, first-to-market solutions and safeguard the data we receive...  ...pioneers of market-changing solutions. We are seeking a Site Reliability Engineer to design, build, and maintain highly available systems and... 
    Local area

    LE038 Second Sight Solutions, LLC

    Washington DC
    2 days ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    1 day ago
  • $160k - $180k

     ...Site Reliability Engineer Location: Hybrid – Washington DC/Virginia/Maryland metro with the ability to travel to Patuxent River, MD, as needed (up to 20% of the time). Compensation: $160,000 - 180,000 per year, depending on experience and qualifications. Employment... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    Fortress Information Security

    Washington DC
    2 days ago
  • $106.3k - $221.1k

     ...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key... 
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    3 days ago
  • $150k - $180k

     ...what’s possible in remote sensing, you belong here at Umbra. About the Job We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's systems.... 
    Permanent employment
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    Umbra

    Arlington, VA
    2 days ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment... 
    Work experience placement

    Samprasoft

    Washington DC
    3 days ago
  •  ...technology solutions using a tailored Agile methodology. We are seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position will be a remote role open to US citizens residing in the... 
    Remote work

    Elevate Government Solutions

    Washington DC
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!