Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GMI Cloud

About GMI

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities

  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  • Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  • Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  • Regional/international travel to GMI data center locations.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  • Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  • Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  • Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Raleigh, NC vacancy
  • $104.9k - $174.7k

     ...automation, AI, and platform transformation to improve efficiency, accuracy, and cost effectiveness.About the RoleSenior Site Reliability Engineer IIWe're looking for a Senior Site Reliability Engineer II to serve as a go-to technical resource for our Life Sciences SRE... 
    Suggested
    Full time
    For contractors
    Work at office
    Local area
    Immediate start
    Flexible hours

    RELX Group

    Raleigh, NC
    23 hours ago
  • IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Suggested
    Work at office
    Immediate start

    IXL Learning

    Raleigh, NC
    7 hours ago
  •  ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises... 
    Suggested
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Raleigh, NC
    4 days ago
  • $139.3k - $203.6k

     ...is a hybrid role***Meet the TeamThe Customer Experience (CX) engineering team is a group of extraordinary technical guides whose main...  ...and employee satisfaction scores at Cisco.Your ImpactAs a Site Reliability Engineer (SRE) on this CX engineering team, you will play a... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    Cisco

    Raleigh, NC
    13 hours ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Suggested
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Raleigh, NC
    1 day ago
  •  ...Site Reliability Engineer Job Overview: Site Reliability Engineer role focuses on reliability, scalability, and performance of enterprise platforms across cloud and on-prem environments. Position requires hands-on engineering with automation, observability, and cross... 

    MDA Edge

    Raleigh, NC
    3 days ago
  • Role Profile:We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a Senior SRE, you will be a senior hands‑on technical person help... 
    Full time
    Shift work

    London Stock Exchange Group

    Raleigh, NC
    4 days ago
  • $55k - $151.47k

     ...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our... 
    Full time
    H1b

    PwC

    Raleigh, NC
    2 days ago
  • $84.9k - $209.5k

     ...Description Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity...  ...new tools and develops and maintains advanced knowledge of site reliability trends. Responsibilities Key Responsibilities... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Raleigh, NC
    1 day ago
  • $121.4k - $218.6k

     ...thrives in a dynamic environment? Join our highly skilled Site Reliability team! Our team designs, develops, and manages applications...  ...and performance tuning. As a Senior Lead Site Reliability Engineer, you will be responsible for: Defining requirements as part... 
    Work experience placement
    Work at office

    Jobleads-US

    Raleigh, NC
    4 days ago
  •  ...Site Reliability Engineer (SRE) - Security Infrastructure Position Summary We are seeking an SRE to support reliability, scalability, and operational excellence for a large-scale network security transformation initiative. This role will focus on monitoring... 

    IS3 Solutions

    Cary, NC
    5 days ago
  •  ...Site Reliability Engineer Number of Position: 2 Only Fulltime I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. *** In case, if you are not... 
    Full time
    Work visa

    Syntricate Technologies

    Cary, NC
    2 days ago
  • $125k - $185k

     ...Job Title: CaaS Private Site Reliability Engineer Corporate Title: Vice President Location: Cary, NC Who we are: In short – an essential part of Deutsche Bank’s technology solution, developing applications for key business areas. Our Technologists drive... 
    Full time
    Work at office
    Work from home
    Shift work

    Deutsche Bank

    Cary, NC
    4 days ago
  • $65 - $70 per hour

     ...MatchPoint Solutions is a fast-growing global IT and Engineering services firm delivering innovative technology solutions to leading...  ...solutions in a collaborative, high-growth environment. Site Reliability Engineer (SRE) – Security Infrastructure Location: Onsite... 
    Local area
    3 days per week

    MatchPoint

    Cary, NC
    2 days ago
  •  ...technologies to enable scalable, secure, and reliable business operations. Applies strong...  ...infrastructure.3. Manages infrastructure engineering projects and processes aligned with...  ...benefit plans, please visit our Benefits site. Depending on the position and division,... 
    Permanent employment
    Full time
    Part time
    Work experience placement
    H1b
    Remote work
    Work visa
    Shift work
    Weekend work
    Day shift

    Truist

    Raleigh, NC
    2 days ago
  • $250k - $300k

    We're looking for a hands-on engineering leader to build and own the Release Engineering, SecDevOps, and Site Reliability Engineering (SRE) functions for Infinia. This is a foundational role: you'll define how our software is built, secured, released, and kept running... 
    Remote work

    DataDirect Networks

    Raleigh, NC
    2 days ago
  •  ...Hogan Software Engineer | 100% Remote (EST Hours) | W2 Contract (12 Months) About the Opportunity Optomi, in partnership with a...  ...troubleshooting production issues, and ensuring application performance and reliability. The ideal candidate will have a strong background in... 
    Contract work
    Remote work

    Optomi

    Raleigh, NC
    38 minutes ago
  • $144k - $198k

     ...BaxterJoin our dynamic team as a Senior Principal Software Systems Engineer in the R&D/Software organization, where you will play a pivotal...  ..., please speak with your recruiter or visit our Benefits site: Benefits | BaxterEqual Employment OpportunityBaxter is an equal... 
    Full time
    Temporary work
    Local area
    Work visa
    Relocation package
    Flexible hours

    Baxter International

    Raleigh, NC
    23 hours ago
  • $130k - $180k

    Position OverviewPower your future with Qualus as a Lead Relay Settings Engineer. In this role you will perform Protective Relay Design & Coordination: Design, specify, calculate settings, and coordinate protective relays and relay control schemes. Do you have 7+ years... 
    Temporary work
    Flexible hours

    Qualus

    Raleigh, NC
    3 days ago
  • $136.09k - $168.11k

     ...looking to move fast and make a significant impact in an exciting space, you're in the right place!We are seeking a Lead Solution Engineer to join our North America GTM team, specifically supporting our Industry Verticals organization across SLED (State, Local & Education... 
    Local area
    Flexible hours

    Freshworks

    Raleigh, NC
    1 day ago
  • $184k - $287.5k

    NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on NVIDIA CPU platforms. The team turns modern toolchains, build and code-health practices, performance-analysis workflows, and optimization techniques into repeatable improvements... 
    Full time
    Remote work

    Nvidia

    Raleigh, NC
    3 days ago
  • $184k - $287.5k

     ...inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely positioned... 
    Full time
    Remote work

    Nvidia

    Raleigh, NC
    3 days ago
  •  ...Reliability Engineer - Hydraulics & Pneumatics The Enviva team is driven by our shared vision for a renewable energy future. We are a fast...  ...capital projects to include large greenfield and brownfield sites and small capital projects. Safeguard all work, including... 
    For contractors
    Local area

    Enviva

    Raleigh, NC
    3 days ago
  •  ...mobile, RF, networking, industrial, business equipment, and automotive. ~ Signal degradation over life Position Reliability Engineer Location Raleigh, NC Responsibilities Test, Validation & Qualification Develop and execute reliability and... 
    Flexible hours

    Amphenol Communications Solutions

    Raleigh, NC
    3 days ago
  • $100k - $153k

     ...Position Overview J ob Title CaaS Private Site Reliability Engineer Corporate Title Assistant Vice President Location Cary, NC Who we are: In short – an essential part of Deutsche Bank’s technology solution, developing applications for key business... 
    Full time
    Work at office
    Work from home
    Shift work
    Cary, NC
    more than 2 months ago
  •  ...AI infrastructure, working with server, cloud, and platform engineering teams.Operationalize machine learning workflows and support AI...  ...implement system enhancements to improve performance, scalability, reliability, and cost efficiency.Collaborate across divisions to support... 
    Full time
    Work at office
    Remote work

    Applied Research Associates

    Raleigh, NC
    4 days ago
  • $175k - $200k

     ...desire to positively impact the environment and lives of others in a refreshing, vibrant, and inclusive culture.As a Senior AI Systems Engineer on the Advanced Systems team, you will architect and guide AI-enabled engineering systems that accelerate advanced aircraft... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work

    BETA TECHNOLOGIES

    Raleigh, NC
    1 day ago
  • $272k - $431.25k

     ...world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems....  ...with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce... 
    Full time
    Remote work
    Shift work

    Nvidia

    Raleigh, NC
    3 days ago
  • $120k - $160k

     ...PLATFORM ENGINEER LOCATION: Raleigh, NC or Beaverton, OR SALARY: $120K- $160K Full - Time Employee Hands-on role focuses on designing, building, and implementing scalable, secure infrastructure across commercial and CMMC Level 2 environments. The ideal... 
    Full time

    Confidential

    Raleigh, NC
    2 days ago
  • $130k - $170k

    About Us:BW Design Group is a fully integrated architecture, engineering, construction, system integration, and consulting firm committed to helping our clients realize their most critical goals from Strategy to Commercialization. As the only firm born from a manufacturing... 
    Full time
    Flexible hours

    Barry-Wehmiller

    Raleigh, NC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!