Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- ...and national security missions. Why Join? Founding engineering role: build and lead the deployment and DevSecOps function from... .... Keywords for Search (SEO): Lead DevSecOps Engineer, Site Reliability Engineer, SRE, Kubernetes, Linux System Administrator, CI CD...SuggestedContract work
$129.5k - $220.5k
...solutions for life’s essentials. Job Requisition #:035650 Division Reliability Engineer (Open)Job Description:Division Reliability EngineerBusiness:... ..., solve recurring problems, and help manufacturing sites implement sustainable improvements.You’ll partner with Maintenance...SuggestedFull timeTemporary workLocal areaWork from home$103.9k - $176.8k
...the Advantage That Helps Our Customers Win.OUR PURPOSE:Creating packaging solutions for life’s essentials.Job Requisition #: Reliability Engineer (Open)Job Description:Reliability EngineerLocation: Tacoma, WashingtonRelocation Assistance: May be consideredBuild...SuggestedFull timeTemporary workFor contractorsWork at officeLocal area- ...established and growing capabilities across Intelligence, Analytics, Engineering, Mission Support, and Communications disciplines. Founded in 20... ...Proactively develop relationships with users at customer site and surface use case gaps Build initial dashboard products and...SuggestedContract workFor contractors
$135.4k - $201.85k
...Senior Software EngineerWe are looking for a Senior Software Engineer to join our Universal Asset Insights team, located in Tacoma, WA... ...product roadmap & also engineering goals of building scalable, reliable & maintainable software. You are the ideal candidate if you are...SuggestedTemporary workFlexible hoursNight shift- ...Join to apply for the Technical Solutions Engineer role at Epic Join to apply for the Technical Solutions Engineer role at Epic Please note that this position is based on our campus in Madison, WI, and requires relocation to the area. We recruit nationally...Full timeWork at officeRelocationVisa sponsorshipRelocation package
$120k - $140k
...Forward Deployment Engineer Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human... ...Ensure solutions meet enterprise requirements for security, reliability, maintainability, quality, and production readiness. Troubleshoot...Hourly payFull timeLocal areaRemote work- ...Are you a seasoned Software Engineer seeking your next opportunity to make a significant impact? Do you bring curiosity and drive, coupled... ...across cross-functional teams. ~ Have experience delivering reliable, high-performance REST APIs. ~ Have a strong interest in...
- ...pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and... ...concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P...For contractorsRemote workFlexible hours
- Software Developer Location US Remote Rate: DOE Requirements • Experience with vanilla JavaScript, JavaScript frameworks like React, and modern CSS frameworks. • Experience developing Microservices, REST APIs, and other database-driven web applications ...Remote work
$100k
...get a Job? | SynergisticIT Who Should Apply? We're looking for recent grads in Mathematics, Statistics , Computer Science or Engineering or candidates with gaps in their career or people wanting to switch careers into tech. SynergisticIT is committed to supporting...H1b- ...Position Title: Lead Engineer Location: Remote (EST) Duration: Long Term Contract (only W2) Team: The team owns the platform... ...and simulation capabilities that support platform validation, reliability testing, and AI model development. Must Have Skills: Node...Long term contractRemote work
- ...years of experience. Full Description We are looking for a motivated and enthusiastic Python Developer to join our remote engineering team at Hiredbuddy. This is an entry-level role perfect for recent graduates or individuals with up to 2 years of experience who...InternshipRemote workHome office
- ...to customers. Role Description This full-time remote DevOps Engineer role is responsible for building, maintaining, and improving... ...implement infrastructure as code, manage cloud resources, and ensure reliable, secure, and scalable environments. Day-to-day tasks include...Full timeRemote work
$180k - $230k
...millions of Americans. About the role As Lead AI/ML Software Engineer for [R]AIMS, you will serve as a senior technical leader... ...scales, reducing incidental complexity and improving operational reliability without sacrificing capability. Optimize platform...Remote work- ...that are vital to national security. We are seeking a DevOps Engineer to support application deployment, cloud platform operations,... ...teams to improve deployment processes and platform reliability. The position focuses on Kubernetes , Docker , CI/CD...
- ...Join to apply for the Software Engineer - Python - Ubuntu Pro client - graduate level role at Canonical . Canonical is a leading provider of open source software and operating systems to the global enterprise and technology markets. Our platform, Ubuntu, is widely...Remote work
- ...Our client, a large federal systems integrator supporting a major federal agency, is hiring a Principal Software Engineer to lead design and delivery on a modernization effort spanning Java microservices, cloud deployment, and frontend UI work. This is a hands on technical...
$102.85k - $139.47k
...Senior Systems Engineer Join TRA Medical Imaging and help power the technology behind exceptional... ...and clinical teams throughout our multi-site imaging network. This role is ideal for a... ..., vendors, and IT teammates to deliver reliable technology solutions. What You Bring...Monday to Friday$57 per hour
...system performance and regular updates and patchesManage local site network equipment, expand capacity to meet demandCoordinate vendor... ...in Information Technology and/or IT Architecture and Engineering as a Systems Engineer to include:Windows Server administration...Local area- ...resolve client needs. Take2 is hiring an AWS Lakehouse Data Engineer who is eligible to be sponsored for a Public Trust Clearance.... ..., and reusable transformation frameworks. Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling...Remote work
- ...employed team of over 800 qualified and registered technical engineers, technicians, and artisans, Bidvest Facilities Management supports... ...with senior engineers and cross-functional teams to ensure reliable, secure, and efficient operation of technology platforms that...Full timeRemote work
- CACI International Inc. seeks a Cyber Operations Integrator to provide expert integration of cyber capabilities for ARSOF SAMMC, conducting cyber intelligence analysis, target development, and vulnerability assessments to support OCO/DCO operations. The role requires advising...
- ...(GSA), via the Technology Transformation Services (TTS) within the Federal Acquisition Service (FAS), seeks Principal AI Software Engineers to shape secure, ethical AI adoption across federal agencies. Multiple vacancies exist with potential for additional openings as needed...
$18 - $25 per hour
...October 2026. If you’re interested in Solutions, Consulting and Engineering this is the place for you! Why Solutions Consulting &... ...skills Ability to collaborate well with others Must provide reliable transportation & housing Certain states and localities...Remote jobHourly paySummer workInternshipWork at officeImmediate startShift work- ...Our client is seeking a hands-on, high-impact Senior React Engineer with deep expertise in modern front-end development using TypeScript, React, and the TanStack ecosystem. This is a heavily coding-focused position where you’ll spend the majority of your time architecting...
- ...that values your well-being. Job Summary The AI Software Engineer develops, integrates, deploys, and supports AI-enabled... ...independently turn defined business needs and technical designs into reliable production solutions. Working within the enterprise AI architecture...Local areaImmediate startRemote work2 days per week
- ...Position Overview A large grocery retailer is seeking a Lead AI Engineer / Agentic Commerce Lead to drive the architecture, development, and delivery of customer-facing AI agent experiences. This is a highly technical leadership role designed for an experienced engineer...Local area
- ...Senior .NET Engineer Location: Remote Department: Engineering Employment Type: Full-Time Build the Future of Transportation... ...who enjoys solving difficult technical problems, building reliable systems, and taking ownership from initial design through production...Full timeImmediate startRemote work
$240k - $300k
...AI Engineer $240k–$300k base + equity Remote, United States High-growth applied AI company building production agents to automate... ...around complex, fragmented enterprise data. Improve agent reliability, performance, and failure handling. Build reusable agent...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- official site Tacoma, WA
- site services specialist Tacoma, WA
- construction site safety Tacoma, WA
- IT site lead Tacoma, WA
- site leader Tacoma, WA
- site safety Tacoma, WA
- historic site Tacoma, WA
- junior website developer Tacoma, WA
- website coordinator Tacoma, WA
- on-site clinical research associate (traveling/remote) Tacoma, WA



