Site Reliability Engineer
GMI Cloud
About GMI
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.
Responsibilities
- Design, implement and maintain scalable AI/ML infrastructure solutions.
- Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
- Automate deployment, configuration and management of infrastructure resources.
- Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
- Implement CI/CD pipelines for infrastructure deployment and orchestration.
- Ensure security, compliance and best practices across infrastructure.
- Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
- Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
- Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
- Regional/international travel to GMI data center locations.
Qualifications
- Bachelor’s degree in Computer Science or related field.
- Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
- Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
- Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
- Familiarity with Linux system administration and scripting (Python, Bash).
- Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
- Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
- Strong troubleshooting skills and ability to analyze system logs and performance metrics.
- Excellent communication and teamwork abilities.
Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
- ...in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.Site Reliability EngineerOnsite: Atlanta, GAJob SummaryAt NCR Voyix, we're looking for a Site Reliability Engineer II to help build, support, and scale the cloud platforms that power our...SuggestedFull timeWorldwideFlexible hours
- ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises...SuggestedPermanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
- Job PurposeAt Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect companies around... ..., results-oriented people to join our team.We are seeking a Site Reliability Engineer to bring 3+ years of hands-on experience to our SRE...Suggested
- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...SuggestedContract work
- ...solving and decision-making abilities and the highest degree of professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The infrastructure cloud team is responsible for internal services that provide...Suggested
$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...Full timeTemporary workWork experience placementFlexible hours- ...Join to apply for the Site Reliability Engineer role at Motion Recruitment Join to apply for the Site Reliability Engineer role at Motion Recruitment Get AI-powered advice on this job and more exclusive features. Every year, nearly 200 million travelers...Contract workWorldwide
$123.4k - $222.53k
...Responsibilities Enhance system reliability and resilience by identifying issues and implementing preventive measures to reduce downtime... ...) ~ Acceptable areas of study include Computer Science, Engineering or related field (Required) ~4-7 years Working in operations...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours- ...operational efficiency, accelerate time-to-value, and deliver better customer experiences.About The RoleWe're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'...Work at officeLocal areaRemote workWork from homeWorldwideHome officeFlexible hours
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours- ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from...Full timeLive inWork at office
- ...Job Title :- Site Reliability Engineer (SRE) Employment Type :- W2 Duration :- Long Term Visa Type :- All Visa applicable which are ready for W2 Location :- Atlanta, GA (Onsite) Job Description We are seeking a highly skilled Site Reliability Engineer (SRE...
- ...Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence...Worldwide
$130k - $150k
...recruiter to learn more. Base pay range $130,000.00/yr - $150,000.00/yr Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) to join our team and help build and maintain scalable, reliable, and efficient systems. The ideal candidate will...Full timeRemote work$178.13k - $205.4k
...Salary Range: $178,131 - $205,400 About You Basic Qualification ~Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5) years of progressive, post‑baccalaureate experience in job offered...Work at officeRemote workFlexible hours$141.8k - $195k
...their best work, grow fast, and bring their full selves to the herd.Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl...Remote work- ...data, and human expertise. We deliver faster, smarter, more reliable insights to insurance carriers and single-family rental... ...at scale, you’re in the right place. The Role As a Site Reliability Engineer, you'll be responsible for the availability, scalability,...Flexible hours
$120k - $175k
...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues...Full timeRemote workWork visaFlexible hours$130k - $160k
...requires a team that works together with trust and cares deeply about carrying out our mission. About the Role The Site Reliability Engineering team at Todyl exists to make our platform reliable, secure, and easy for engineering teams to ship to. We do that by...Full timeTemporary workWork at officeLocal areaFlexible hoursShift work3 days per week- ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and...
- ...Site Reliability Engineer At Acuity, you will join an Agile team focused on building and supporting advanced platforms and applications that drive our business forward. We are seeking a Site Reliability Engineer (SRE) to help define and raise the reliability bar for...
$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...- ...Technical Support Specialist In Site Reliability Engineering (SRE) Mandatory skills: Scripting and programming languages like Python, Java, Ruby. Cloud and infrastructure management – AWS, Google cloud and Azure is a plus- CI/CD Automation, Database Management....
- ...Site Reliability Engineer We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate, and...
$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office$61.09k - $104.36k
...Site Reliability Engineer Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like, where you’ll be supported and inspired by a collaborative community of colleagues around the world, and where you’ll be able...Permanent employmentFull timeContract workLocal area- ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE. Required Skillset • Manage and optimize data streaming and API components in OpenShift Onpremise and AWS. • Proactively...Work experience placement
- ...ideal time and number for communication, and the expected pay rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months Contract Visa : US Citizens/ Green Card Need Local to...Contract workLocal areaImmediate start
$60 - $68 per hour
...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential and is located in Atlanta, GA (Onsite). Please review the job description below and contact me ASAP if you are interested...Contract workLocal areaImmediate start- ...and we're looking for new team members who want to be a part of this journey! We're looking for a proactive, hands-on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in fast-moving startup environments. You're someone who enjoys owning...Work experience placementFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Atlanta, GA
- site reliability engineer Atlanta, GA
- site reliability engineer remote Atlanta, GA
- official site Atlanta, GA
- site services specialist Atlanta, GA
- construction site safety Atlanta, GA
- IT site lead Atlanta, GA
- site recruiter Atlanta, GA
- site leader Atlanta, GA
- site safety Atlanta, GA



