Get new jobs by email
$80.3k - $133.3k
...retrosynthesis, and AI for science more broadly. Excellent office, programming, and communication skills are essential. SPARC offers GPU/HPC infrastructure, large biomedical data assets, deep clinical partnerships, and direct opportunities to translate methodological advances...SuggestedWork experience placementWork at officeShift workDay shift$74.1k - $120.7k
...drug discovery, retrosynthesis, and AI for science more broadly. SPARC offers a collaborative, computationally rich environment with GPU/HPC infrastructure, large biomedical data assets, and tight partnerships across the UAB School of Medicine. Excellent office, programming...SuggestedTraineeshipWork experience placementWork at officeShift workDay shift- ...ArgoCD. Proven ability to troubleshoot complex distributed systems, largely self-directed. Preferred Qualifications: GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar. NVIDIA GPU orchestration: A100/H100 configuration, driver...Suggested
$20 per hour
...Salt, Argo CD, etc.). Strong understanding of Linux systems and container runtimes (e.g., containerd, Docker) Experience with GPU workloads or high-throughput computing. Hands-on experience operating and optimizing High-Performance Computing (HPC)...SuggestedPermanent employmentFull timeImmediate startRelocation packageFlexible hoursWeekend work- ...Position : SRE Solutions Architect AI, HPC & GPU Location: Remote Preferred: Santa Clara, CA | Washington | Oregon | Texas Duration: 12+ Months Contract Travel: Occasional Job Description: Seeking an experienced SRE Solutions Architect with strong...SuggestedContract workRemote work
- ...observability, automation, and system-level debugging . Candidates should have deep expertise in at least one of the following areas: GPU-based hardware, including management, troubleshooting, and large-scale operations Software-Defined Networking (SDN), including OVN...Suggested
- ...Duration: 1 month Location: St. Louis, MO Role Overview Highly skilled AI Infrastructure Engineer to design, build, and operate scalable GPU-enabled Kubernetes platforms for AI/ML workloads. Must have skills: Kubernetes, Linux, Terraform, GPUs and NVIDIA stack...Suggested
- ...supporting AI and analytics workloads. Collaborate with data engineering teams to optimize storage, compute, and data pipelines. Enable GPU infrastructure and high-performance computing environments for AI workloads. Support data governance, backup, disaster recovery, and...Suggested
- ...in Python, Go, or similar; comfortable owning complex systems end to end. Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments (Slurm) is a strong plus. You can architect systems, write the code, review others' work, and...Suggested
- ...LLMs, building high-performance inference platforms using technologies such as vLLM, SGLang, TGI, Triton, and Ray Serve, and driving GPU utilization, latency, throughput, and cost optimization. This is a highly technical role requiring active involvement in designing, building...SuggestedContract workLocal area
$80k - $105k
...is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation. Founded by...SuggestedFull timeWork at officeImmediate startRemote workFlexible hours$80k - $100k
...is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation. Founded by...SuggestedFull timeWork at officeImmediate startRemote workFlexible hours$118.4k - $196.8k
...months ahead in AI trends and prepare the infrastructure accordingly. Cost management skills , specifically optimizing cloud spend for GPU/TPU resources, are crucial. Exceptional communication skills are also required, demonstrated through the ability to document...SuggestedFull timeFor contractorsLive inLocal area- ...profiles, and Active Directory integration, to maintain optimal system health. Analyze workstation performance metrics related to CPU, GPU, memory, and overall system resources, identifying and resolving bottlenecks. Support Dell hardware troubleshooting, including...SuggestedWork at office
$152k - $287.5k
...innovate, develop, and integrate innovative AI solutions into the design and automation infrastructure that powers our chips. Every CPU, GPU, and Tegra SoC NVIDIA has shipped in the past four years passed through our toolchain on its way to production — over 200 product...SuggestedFull time- ...operational stability Anticipates the infrastructure needs of the data platform and agent runtime - compute, storage, networking, and GPU serving/training capacity - and translates them into a funded, prioritized roadmap Makes decisions that influence resourcing,...
$175k - $230k
DescriptionKforce has a client that is seeking a Senior Technical Product Manager GPU Orchestration in West Palm Beach, FL.Key Tasks:* Define and execute the roadmap for managed Kubernetes, managed Slurm services, SUNK, and Run:ai integration* Own the end-to-end cluster...$91.42k - $152.38k
...responsible AI practices. Experience with AI governance frameworks, LLM evaluation methodologies, and AI safety tooling. Experience with GPU infrastructure optimization and scalable inference architectures. Familiarity with multi-agent AI systems and autonomous workflows...Full timeTemporary workRemote workShift work$113.1k - $232.3k
...they are AI-infused, their distinct production failure modes such as drift, train/serve skew, latency and output variance, and token/GPU cost anomalies. Translate reliability needs, reference architectures, and operational requirements into service-level objectives, runbooks...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...libraries, especially PyTorch, TensorRT, or TensorRT-LLM. Demonstrated interest and experience in LLM’s. Deep understanding of GPU architecture. BONUS POINTS: Proficiency in enhancing the performance of software systems, particularly in the context of large...Full time
- ...certification: CCIE (RS, SP, or DC track), JNCIE-SP, or JNCIE-DC.Experience with AI/HPC fabric design and lossless transport requirements for GPU clusters.Ideal Candidate ProfileThe right person for this role has hands-on experience designing and deploying production backbone or...Full timeVisa sponsorshipFlexible hours3 days per week
- ...and curriculum learning.Proven experience building simulation environments in NVIDIA Omniverse, Isaac Sim, or Isaac Lab, or comparable GPU-accelerated simulation platforms.Direct experience generating synthetic data for model training, including sensor simulation,...Full timeLocal areaFlexible hours
- ...responsible usage across teamsTrack record of conference talks, published papers, or significant open-source contributionsExperience with GPU-accelerated inference and model serving optimizationFamiliarity with workflow orchestration and streaming architectures for real-time...Full timeFixed term contractWork at officeRemote workFlexible hours
- ...Coordinate to access compute and storage resources required to support large-scale model training and analytics workloads, including GPU accelerated environments, high-performance object storage, and parallel data processing capabilities.Fund usage-based compute and storage...
- ...safety-critical software development / DO-178 B or above industry standardsWorking knowledge of Bash/Python, Intel x86/ARM/FPGA/SoC, GPP, GPU, Embedded System Design, Release Engineering, Understanding of Change and Configuration Management, Debug/Trace Analysis, Real-time...Immediate start
$171.6k - $338.3k
...deployment strategies. 10+ years of experience architecting, implementing, and optimizing enterprise-scale hybrid infrastructure solutions (GPU platforms, server hardware, AI transformation), including working fluency in how rack power density and liquid cooling requirements...Local area- ...architecture design. Training techniques (pre-training, fine-tuning, RLHF), and optimization methods. Understanding of low-level details like GPU memory management, precision types (float16, bfloat16), and parallelization techniques. Proficiency in advanced training techniques...
$133.9k - $154.5k
...relationshipsBuild and lead the AI Platform Operations teamDefine the architecture, standards, and governance for AI/ML infrastructure, including GPU cluster design, compute resource planning, security controls, and observability across the platformDrive the strategy, standards, and...Full timeContract workTemporary workFlexible hours$125k - $140k
...West Palm Beach, FL that is seeking a Data Center Infrastructure Architect.Responsibilities:Architect and Design Network, Server, and GPU Infrastructure:* Architect and design the physical infrastructure to support data center scale deployments of networks, servers, and...For contractors- Job DescriptionSome markets you join. This one you help build.AI compute is outgrowing the ground. As GPU and AI-class systems move into space and other demanding high-reliability settings, the power electronics that feed them have to survive radiation, wide temperature...
