Get new jobs by email
$160k - $200k
...modes * Support and optimize data-intensive and MLOps workloads using technologies such as Kafka, Flink, Pinot, KServe, Kubeflow, Ray, GPU-enabled nodes, and model-serving pipelines * Build and refine observability patterns using Prometheus, Grafana, Fluent Bit, Loki,...SuggestedFull timeLive inRelocation- ...in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, OpenKyber delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and...Suggested
- ...data handling, regulatory expectations, and third-party data use Secure AI infrastructure and supply chain: Harden AI platforms, GPU and container workloads, model registries, and artifact stores Assess risks in third-party models, libraries, embeddings, and...SuggestedFlexible hours
- ...end-to-end Observability Platform roadmap across telemetry ingestion, querying, visualization, alerting, and retention for large-scale GPU clusters and multi-tenant cloud environments Define Vultr's observability strategy across bare metal, VMs, Kubernetes, and...SuggestedHourly payContract work
- ...and Optimization: Hands-on experience with LLMs, generative AI (prompt engineering, Agent frameworks, fine-tuning, RAG, LLMOps), and GPU optimizations. Observability and Feedback: Familiar with observability, telemetry, and user feedback loops for product improvement...SuggestedContract work
- ...optimization of AI-ready cloud environments: Modern AI workloads require specialized architectures, including high performance compute, GPU clusters, container orchestration, and distributed storage. An experienced architect ensures these environments are properly designed...SuggestedContract workRemote work
$148k - $216k
...utilization, data drift, and concept drift. Infrastructure Management: Provision and optimize cloud-based ML infrastructure (including GPU/CPU computing clusters) utilizing Infrastructure as Code (IaC) paradigms. Cross-Functional Collaboration: Work intimately with...SuggestedRemote workFlexible hours$132k - $191k
...Design and support cloud architectures for AI/ML workloads, including model training, inference, and high-performance compute (e.g., GPU/EDA burst capacity). Enable secure data pipelines, scalable compute environments, and integration of AI services while ensuring...SuggestedPermanent employment- ...infra /Kubernetes (Remote) Are you passionate about building scalable AI infrastructure and helping customers succeed with cutting-edge GPU platforms? We're looking for a Solutions Architect to join our team and work with enterprise customers deploying and optimizing AI/ML...SuggestedRemote work
- ...Level Troubleshooting: Investigating and troubleshooting problems and hardware faults that our automation can't determine within our GPU platforms. This will involve taking data from system logs, kernel logs, BMC redfish APIs, and if the data is not there, working with...SuggestedLong term contractWork from home
- ...RESPONSIBILITIES: Own the discovery and definition of customer requirements for AI infrastructure use cases, including training, inference, GPU clusters, bare metal, managed orchestration, networking, and storage Work directly with strategic customers to understand their...SuggestedHourly payContract workLocal area
- ...evaluate and guide the following areas: Future AI rack density and power consumption trends Impacts of next-generation GPU and AI chip architectures Optical networking and switching implications on infrastructure design AI workload impacts on utility...SuggestedWork at officeLocal areaWork visa
$80.2k - $166.1k
...• Technical foundation that allows you to quickly absorb the complexities of cloud architecture, distributed systems, and emerging GPU/AI infrastructure. • Work across large organizations and bring people together around shared goals. • Embrace a growth mindset, learn...SuggestedTemporary workFlexible hours$126.1k - $261.9k
...global deployment. Partner with the best You'll focus on the SmartNIC/DPU software and hardware that underpins our compute and GPU servers. These DPUs offload CPU resources, improve isolation, and accelerate services across Akamai Cloud. You'll help define and integrate...SuggestedPermanent employmentWork experience placementWork at officeWork from homeWorldwideFlexible hours$160k - $200k
...Support and optimize data-intensive and MLOps workloads using technologies such as Kafka, Flink, Pinot, KServe, Kubeflow, Ray, GPU-enabled nodes, and model-serving pipelines Build and refine observability patterns using Prometheus, Grafana, Fluent Bit, Loki,...SuggestedLive inRelocation$112.8k - $264.1k
...support for new HW platforms. We support both Bare Metal and Virtual machine instances across a diverse fleet of HW, including clustered GPU platforms. In addition, our customers demand high availability and security from our Cloud. We need a leader who can thrive in...Temporary workFlexible hours$126.2k - $264.1k
...area or are able to relocate permanently to the area. **** Preferred Qualifications Direct commissioning experience supporting GPU/high-density or liquid-cooled data halls (CDUs, leak detection, controls integration). Experience commissioning hyperscale or...Temporary workFor contractorsLive inLocal areaRelocationRelocation packageFlexible hours$89.2k - $209.5k
...team is responsible for deliver trusted, fast health determinations and customer‑initiated diagnostics that reduce false positives for GPU clusters, prevent unnecessary node returns, increase capacity for customers, protect revenue, and improve uptime—by providing an OCI‑...Temporary workFlexible hours$74.1k - $148.3k
...schedules. • Ensure that all work complies with OCI specifications, manufacturer warranty standards, and regional regulations. GPU Liquid-Cooled Rack Megaprojects • Serve as the technical and delivery lead for GPU-intensive data hall builds, managing low-voltage...Temporary workLive inLocal areaWorldwideRelocationRelocation packageFlexible hours
