Get new jobs by email
$70 per hour
...Jenkins, or ArgoCD. Proven ability to troubleshoot complex distributed systems, largely self-directed. Preferred Qualifications GPU infrastructure and AI/ML workloads: Ray, Kubeflow, ML flow, or similar. NVIDIA GPU orchestration: A100/H100 configuration,...SuggestedHourly payRemote work$20 per hour
...Ansible, Salt, Argo CD, etc.). Strong understanding of Linux systems and container runtimes (e.g., containerd, Docker) Experience with GPU workloads or high-throughput computing. Hands-on experience operating and optimizing High-Performance Computing (HPC) environments,...SuggestedPermanent employmentFull timeImmediate startRelocation packageFlexible hoursWeekend work- ...Position : SRE Solutions Architect AI, HPC & GPU Location: Remote Preferred: Santa Clara, CA | Washington | Oregon | Texas Duration: 12+ Months Contract Travel: Occasional Job Description: Seeking an experienced SRE Solutions Architect with strong expertise...SuggestedContract workRemote work
- ...observability, automation, and system-level debugging . Candidates should have deep expertise in at least one of the following areas: GPU-based hardware, including management, troubleshooting, and large-scale operations Software-Defined Networking (SDN), including OVN...Suggested
- ...Duration: 1 month Location: St. Louis, MO Role Overview Highly skilled AI Infrastructure Engineer to design, build, and operate scalable GPU-enabled Kubernetes platforms for AI/ML workloads. Must have skills: Kubernetes, Linux, Terraform, GPUs and NVIDIA stack...Suggested
- ...in Python, Go, or similar; comfortable owning complex systems end to end. Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments (Slurm) is a strong plus. You can architect systems, write the code, review others' work, and...Suggested
- ...supporting AI and analytics workloads. Collaborate with data engineering teams to optimize storage, compute, and data pipelines. Enable GPU infrastructure and high-performance computing environments for AI workloads. Support data governance, backup, disaster recovery, and...Suggested
- ...LLMs, building high-performance inference platforms using technologies such as vLLM, SGLang, TGI, Triton, and Ray Serve, and driving GPU utilization, latency, throughput, and cost optimization. This is a highly technical role requiring active involvement in designing, building...SuggestedContract workLocal area
- ...profiles, and Active Directory integration, to maintain optimal system health. Analyze workstation performance metrics related to CPU, GPU, memory, and overall system resources, identifying and resolving bottlenecks. Support Dell hardware troubleshooting, including...SuggestedWork at office
- ...operational stability Anticipates the infrastructure needs of the data platform and agent runtime - compute, storage, networking, and GPU serving/training capacity - and translates them into a funded, prioritized roadmap Makes decisions that influence resourcing,...Suggested