Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer

Full-time

42dot

About The Team & Mission

42dot의 AI 인프라 엔지니어는 여러 데이터 센터에 걸쳐 있는 수천 개의 GPU를 관리하며, 이를 효율적으로 오케스트레이션하는 고성능 AI 인프라를 운영합니다. 세계 최고 수준의 컴퓨팅 환경을 유지하기 위해 확장성, 모니터링 및 운영 최적화 전반에 기여하게 됩니다.

At 42dot, our AI Infrastructure Engineer manages the high-performance AI infrastructure orchestrating thousands of GPUs across multiple data centers. You will contribute to the scaling, monitoring, and operational optimization required to maintain a robust and world-class computing environment.

Responsibilities

  • Kubernetes 및 Slurm을 활용하여 여러 데이터 센터에 분산된 수천 개 규모의 대규모 GPU 클러스터 운영 및 유지 보수
  • GPU 하드웨어 및 소프트웨어 스택 전반의 장애를 모니터링하고 진단하여 고가용성 유지 및 신속한 장애 복구 수행
  • Python 또는 Shell을 활용한 자동화 도구 및 스크립트를 개발하여 반복적인 인프라 관리 업무를 효율화
  • GPU 리소스 쿼터(Quota) 관리 및 ML 개발자를 위한 기술 지원을 통해 컴퓨팅 자원의 최적 활용 보장
  • 대규모 자율주행 모델 학습을 위한 분산 학습 환경의 아키텍처 설계 및 성능 튜닝 참여
  • Operate and maintain a large-scale GPU cluster consisting of thousands of GPUs across multiple data centers using Kubernetes and Slurm.
  • Monitor and diagnose failures across the GPU hardware and software stacks to ensure high availability and rapid recovery.
  • Develop automation tools and scripts using Python or Shell to streamline repetitive infrastructure management tasks and improve operational efficiency.
  • Manage GPU resource quotas and provide technical support to ML researchers to ensure optimal utilization of computing resources.
  • Participate in the architectural design and performance tuning of distributed training environments for large-scale autonomous driving models.

Qualifications

  • Linux 운영체제에 대한 깊은 이해 (커널 동작, 프로세스 관리, 시스템 보안 등)
  • Docker 및 Kubernetes 등 컨테이너 기반 기술 및 오케스트레이션 실무 경험
  • TCP/IP, 등 네트워크 기본 원리에 대한 이해 및 기초적인 네트워크 트러블슈팅 능력
  • Python 또는 Shell을 활용하여 유지보수가 용이한 자동화/시스템 관리 스크립트 작성 역량
  • 복잡하고 거대한 시스템에서 근본 원인을 찾아 해결하는 논리적인 문제 해결 능력
  • 다양한 유관 부서 및 파트너와 원활하게 소통할 수 있는 커뮤니케이션 역량
  • Strong proficiency in Linux operating systems, including a solid understanding of kernel operations, process management, and system security.
  • Practical experience with containerization technologies (Docker) and orchestration (Kubernetes), including building, managing, and troubleshooting containerized environments.
  • Solid understanding of networking fundamentals, including TCP/IP and with the ability to perform basic network troubleshooting.
  • Ability to write clean and maintainable scripts in Python or Shell for automation and system administration.
  • Logical approach to problem-solving with the persistence to identify and resolve root causes in complex, large-scale systems.
  • Strong communication skills to effectively collaborate with cross-functional teams and external partners.

Preferred Qualifications

  • Prometheus, Grafana, Datadog 등을 활용한 대규모 클러스터의 관측성(Observability) 스택 구축 경험
  • AWS, GCP 등 퍼블릭 클라우드 플랫폼 상의 인프라 구축 및 운영 경험
  • 드라이버, CUDA, NCCL 등을 포함한 NVIDIA 가속 컴퓨팅 스택에 대한 지식
  • ML 모델 학습 라이프사이클 및 PyTorch, TensorFlow 등 딥러닝 프레임워크에 대한 이해
  • Kubernetes 또는 Slurm과 같은 대규모 워크로드 매니저 및 리소스 스케줄링 도구 활용 경험
  • Terraform 등 Infrastructure as Code(IaC) 도구를 활용한 복잡한 인프라 관리 경험
  • Experience in building observability stacks with Prometheus, Grafana, and Datadog for large-scale clusters.
  • Experience in building or operating infrastructure on public cloud platforms such as AWS or GCP.
  • Knowledge of the NVIDIA accelerated computing stack, including drivers, CUDA, and NCCL.
  • Familiarity with the ML model training lifecycle and deep learning frameworks such as PyTorch or TensorFlow.
  • Experience with large-scale workload managers or resource scheduling tools such as Kubernetes or Slurm.
  • Familiarity with Infrastructure as Code (IaC) tools such as Terraform to manage complex infrastructure.

※ Please review the following information before applying.

How to work in 42dot, About 42dot Way →
Vacancy posted 8 days ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in South Korea vacancy
  •  ...About The Team & Mission Senior Physical AI Engineer는 차세대 자율주행을 위한 End-to-End(E2E) Planning Model과 모델의 학습·검증을 위한 Closed-loop Simulation 환경 및 파이프라인의 설계·개발을 주도합니다. Trajectory Generation과 Decision-making을 수행하는 E2E Planning Model부터 Closed-loop Simulation까지, 실제 차량의 주행... 
    Suggested
    Full time

    42dot

    South Korea
    19 days ago
  •  ...About The Team & Mission Physical AI Engineer는 차세대 자율주행을 위한 End-to-End(E2E) Planning Model을 설계·개발하고, 모델의 학습·검증을 위한 Closed-loop Simulation 환경과 파이프라인을 개발합니다. Trajectory Generation과 Decision-making을 수행하는 E2E Planning Model부터 Closed-loop Simulation까지, 주행 상황을 바탕으로 안전한... 
    Suggested
    Full time

    42dot

    South Korea
    19 days ago
  •  ...About The Team & Mission IVI OS Engineer (Network)는 Android 기반 디바이스가 차량 내 다른 시스템들과 Ethernet을 통해 안정적으로 통신하고 제어될 수 있도록 Ethernet/Wifi/Telephony 프레임워크를 커스터마이징 및 최적화하는 역할을 수행합니다. 이 역할은 네트워크 조작 기술, HAL 구현 경험, Android Automotive OS의 네트워크 구조에 대한 깊은 이해를 요구합니다. Responsibilities... 
    Suggested
    Full time

    42dot

    South Korea
    18 days ago
  •  ...About The Team & Mission 42dot Software Update Engineer(Software Update Application) 는 SDV 차량의 소프트웨어를 안전하고 효율적으로 배포하기 위해 Vehicle Software Update 시스템을 설계·구현·운영합니다. 업데이트 패키지의 전달, 검증, 설치, 상태 관리 등 차량 소프트웨어 업데이트 전반의 흐름을 고려하여 안정적인 배포 환경을 구축하고, 다양한 차량 환경에서도 일관된 업데이트 경험을 제공... 
    Suggested
    Full time

    42dot

    South Korea
    11 days ago
  •  ...About The Team & Mission AD Framework Software Engineer는 Autonomous Driving System을 위한 핵심 Middleware Framework를 설계하고 개발합니다. 실시간 통신, 실행 프레임워크, 데이터 관리 및 공통 라이브러리 등 자율주행 소프트웨어의 기반이 되는 핵심 시스템을 구축하며, 안전성과 신뢰성을 갖춘 Automotive Grade Software를 제공합니다. 또한 다양한 Application Team... 
    Suggested
    Full time

    42dot

    South Korea
    18 days ago
  •  ...About The Team & Mission 42dot의 Release Engineer는 Pleos의 End-to-End 소프트웨어 Release Lifecycle을 총괄 운영합니다. 여기에는 자율주행 구성 요소, 소프트웨어 정의 차량(SDV) 기능, 클라우드 연결 서비스, 그리고 배포된 차량이 항상 최신의 가장 안전한 코드를 실행할 수 있도록 하는 OTA(Over-the-Air) 업데이트 파이프라인까지 모든 것을 포함합니다. 단순한 릴리스 관리를 넘어, 자기주도적으... 
    Full time

    42dot

    South Korea
    19 days ago
  •  ...inspired is expected and making a meaningful impact is rewarded. Purpose of the role As a core member of the Technology Infrastructure Services (TIS) team in MUFG Seoul Branch, this role ensures stable and secure infrastructure operations while supporting the Head... 
    Full time
    Work at office
    Local area

    MUFG Bank, Ltd.

    South Korea
    10 days ago
  •  ...ABOUT YOU We are looking for a Senior Backend Engineer who is detail-oriented, proactive, and highly collaborative to join our Engineering...  ...that power high-impact products and love creating seamless infrastructure that enables innovative payment, gaming, or commerce... 
    Full time
    Local area

    Xsolla

    South Korea
    24 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!