Get new jobs by email
$148k - $216k
...utilization, data drift, and concept drift. Infrastructure Management: Provision and optimize cloud-based ML infrastructure (including GPU/CPU computing clusters) utilizing Infrastructure as Code (IaC) paradigms. Cross-Functional Collaboration: Work intimately with...SuggestedRemote workFlexible hours$132k - $191k
...Design and support cloud architectures for AI/ML workloads, including model training, inference, and high-performance compute (e.g., GPU/EDA burst capacity). Enable secure data pipelines, scalable compute environments, and integration of AI services while ensuring...SuggestedPermanent employment- ...data handling, regulatory expectations, and third-party data use Secure AI infrastructure and supply chain: Harden AI platforms, GPU and container workloads, model registries, and artifact stores Assess risks in third-party models, libraries, embeddings, and...Suggested
- ...Level Troubleshooting: Investigating and troubleshooting problems and hardware faults that our automation can't determine within our GPU platforms. This will involve taking data from system logs, kernel logs, BMC redfish APIs, and if the data is not there, working with...SuggestedLong term contractWork from home
- ...Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms. Hardware and GPU troubleshooting experience (nice to have). Exposure to OVN/OVS-based networking stack (nice to have). Strong communication skills....SuggestedNight shiftDay shift
- ...infra /Kubernetes (Remote) Are you passionate about building scalable AI infrastructure and helping customers succeed with cutting-edge GPU platforms? We're looking for a Solutions Architect to join our team and work with enterprise customers deploying and optimizing AI/ML...SuggestedRemote work
- ...as we shape the future of AI and beyond. Together, we advance your career. THE ROLE: AMD is looking for a Enterprise AI/HPC GPU architect to join our Datacenter System Architecture and Engineering team to develop world-class products around Instinct GPUs. In this...Suggested
- ...RESPONSIBILITIES: Own the discovery and definition of customer requirements for AI infrastructure use cases, including training, inference, GPU clusters, bare metal, managed orchestration, networking, and storage Work directly with strategic customers to understand their...SuggestedHourly payContract workLocal area
$175k - $220k
...on multi-functional teams to provide ethernet network expertise to server infrastructure builds, accelerated computing workloads and GPU enabled AI applications. Implementing tasks related to network configuration and validation for data centers. Create methods...SuggestedWorldwide- ...evaluate and guide the following areas: Future AI rack density and power consumption trends Impacts of next-generation GPU and AI chip architectures Optical networking and switching implications on infrastructure design AI workload impacts on utility...SuggestedWork at officeLocal areaWork visa
- ...runbooks Enforce IAM least-privilege policies, secrets management, and FinOps cost controls Collaborate with AI/ML engineers on GPU workloads, model serving, and inference pipelines Own technical communication during clinet interaction, translating...Suggested
$165k - $242k
...Strong understanding of design for mass manufacturing and reliability Preferred: ~ Experience in the hyperscaler space, specifically GPU systems ~ Prior experience with ODM/JDM design model ~10+ years of hardware development experience with a focus on servers ~...SuggestedPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hours$188k - $275k
...delivery of the intelligence that drives innovation. What You’ll Do: As a Staff Software Engineer, you’ll shape the backbone of our GPU-driven data centers—powering some of the most advanced workloads in AI and large-scale computing. This isn’t just about keeping the lights...SuggestedPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hours$106.3k - $234.6k
...Expertise in one the following areas: ~ Root Of Trust (TCG SRTM, DRTM) ~ x86 (Intel, AMD), ARM server platform architecture, UEFI ~ GPU platforms, rackscale systems, clustering ~ Baseboard Management Controllers ~ SmartNICs (DPUs) ~ Storage devices ~ Security...SuggestedTemporary workFlexible hours$125k - $180k
...Clear written and verbal communication skills in cross‑functional environments It will be an added bonus if you have Experience in GPU‑dense, AI, or high‑performance computing environments Exposure to firmware lifecycle management and large‑scale rollout validation...SuggestedRemote workFlexible hours$193.6k - $414.4k
...planning and execution for the compute, networking, storage, data, inference/training at large scale, orchestration, observability, and GPU/HPC capacity required to support high-volume GenAI workloads. Drive disciplined capacity forecasting, utilization management, and...Temporary workFlexible hours$95k - $171k
...activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...including model training, inference, vector search, and agentic orchestration. Build usage‑aligned pricing models that scale with compute, GPU, and AI service consumption. Conduct competitive analysis across AI platforms. Develop packaging and SKU strategy for AI platform and...Permanent employmentFlexible hours
$89.2k - $209.5k
...team is responsible for deliver trusted, fast health determinations and customer‑initiated diagnostics that reduce false positives for GPU clusters, prevent unnecessary node returns, increase capacity for customers, protect revenue, and improve uptime—by providing an OCI‑...Temporary workFlexible hours$102.3k - $209.5k
...building and managing the integrated master schedule (IMS) for large-scale OCI data center programs, from project initiation through GPU/rack handover . This role drives schedule predictability across a matrixed organization by aligning internal teams and external...Temporary workFor contractorsLive inRelocationRelocation packageFlexible hoursShift work- ...You will strive to automate the delivery of existing and new Ubuntu image products applied to all modern workloads from web servers to GPU‑aided AI for servers, VM’s and containers. As an engineering manager at Canonical your primary responsibility is to the people you...Work at officeWork from home
$102.3k - $209.5k
...construction audit practices, and formal quality management processes. Experience with AI infrastructure, high-density data halls, GPU deployments, liquid-cooled environments, or large-scale cloud infrastructure projects. Professional certifications such as CQM,...Temporary workFor contractorsFor subcontractorLive inRelocationRelocation packageFlexible hours
