HPC Engineer
Institute of Foundation Models
Institute For Foundation Models
This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.
Responsibilities
- Monitor health, performance, and availability of large-scale GPU clusters.
- Respond to incidents and perform first-level triage.
- Support researchers and troubleshoot job failures.
- Execute operational runbooks and recovery procedures.
- Validate cluster deployments, upgrades, and maintenance activities.
- Track infrastructure utilization and operational metrics.
- Develop automation and monitoring tools.
- Contribute to documentation and reporting.
Education
Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.
Experience
- 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
- Strong Linux troubleshooting skills.
- Experience with scripting using Python or Bash.
Preferred Qualifications
- Slurm.
- GPU infrastructure.
- AWS, Azure, or GCP.
- Grafana, Prometheus, Datadog, or similar tools.
- Containers and Kubernetes.
- AI/ML infrastructure exposure.
- Research computing environments.
Salary Range
$150,000 - $300,000 a year
Benefits Include
- Comprehensive medical, dental, and vision benefits
- Bonus
- 401K Plan
- Generous paid time off, sick leave and holidays
- Paid Parental Leave
- Employee Assistance Program
- Life insurance and disability
Vacancy posted more than 2 months ago
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to HPC Engineer. Be the first to apply!
