Get new jobs by email
$132k - $191k
...Design and support cloud architectures for AI/ML workloads, including model training, inference, and high-performance compute (e.g., GPU/EDA burst capacity). Enable secure data pipelines, scalable compute environments, and integration of AI services while ensuring...SuggestedPermanent employment$148k - $216k
...utilization, data drift, and concept drift. Infrastructure Management: Provision and optimize cloud-based ML infrastructure (including GPU/CPU computing clusters) utilizing Infrastructure as Code (IaC) paradigms. Cross-Functional Collaboration: Work intimately with...SuggestedRemote workFlexible hours- ...infra /Kubernetes (Remote) Are you passionate about building scalable AI infrastructure and helping customers succeed with cutting-edge GPU platforms? We're looking for a Solutions Architect to join our team and work with enterprise customers deploying and optimizing AI/ML...SuggestedRemote work
- ...data handling, regulatory expectations, and third-party data use Secure AI infrastructure and supply chain: Harden AI platforms, GPU and container workloads, model registries, and artifact stores Assess risks in third-party models, libraries, embeddings, and...Suggested
- ...Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms. Hardware and GPU troubleshooting experience (nice to have). Exposure to OVN/OVS-based networking stack (nice to have). Strong communication skills....SuggestedNight shiftDay shift
- ...Level Troubleshooting: Investigating and troubleshooting problems and hardware faults that our automation can't determine within our GPU platforms. This will involve taking data from system logs, kernel logs, BMC redfish APIs, and if the data is not there, working with...SuggestedLong term contractWork from home
$175k - $220k
...on multi-functional teams to provide ethernet network expertise to server infrastructure builds, accelerated computing workloads and GPU enabled AI applications. Implementing tasks related to network configuration and validation for data centers. Create methods...SuggestedWorldwide- ...RESPONSIBILITIES: Own the discovery and definition of customer requirements for AI infrastructure use cases, including training, inference, GPU clusters, bare metal, managed orchestration, networking, and storage Work directly with strategic customers to understand their...SuggestedHourly payContract workLocal area
- ...evaluate and guide the following areas: Future AI rack density and power consumption trends Impacts of next-generation GPU and AI chip architectures Optical networking and switching implications on infrastructure design AI workload impacts on utility...SuggestedWork at officeLocal areaWork visa
- ...as we shape the future of AI and beyond. Together, we advance your career. THE ROLE: AMD is looking for a Enterprise AI/HPC GPU architect to join our Datacenter System Architecture and Engineering team to develop world-class products around Instinct GPUs. In this...Suggested
- ...runbooks Enforce IAM least-privilege policies, secrets management, and FinOps cost controls Collaborate with AI/ML engineers on GPU workloads, model serving, and inference pipelines Own technical communication during clinet interaction, translating...Suggested
$80.2k - $166.1k
...• Technical foundation that allows you to quickly absorb the complexities of cloud architecture, distributed systems, and emerging GPU/AI infrastructure. • Work across large organizations and bring people together around shared goals. • Embrace a growth mindset, learn...SuggestedTemporary workFlexible hours- ...smooth operation of multi-user computer systems consisting of Linux based application and license servers, virtual machines, and GP/GPU cluster based systems. Coordination with network administrators is required. Additional duties: # Setting up administrator and service...Suggested
- ...smooth operation of multi-user computer systems consisting of Linux based application and license servers, virtual machines, and GP/GPU cluster based systems. Coordination with network administrators required. Additional duties are: Setting up administrator and service...SuggestedInterim role
$102.3k - $209.5k
...operate the RDMA/RoCE network fabrics for OCI's largest AI and HPC customers. These fabrics are the foundation underneath OCI's AI, GPU and HPC services, and support major tier-0 vendors in the generative AI industry. If you're running an AI workload at OCI, we're running...SuggestedTemporary workFlexible hours- ...skills. Preferred Qualifications B.S. in Engineering, Facilities Management, or related field; advanced degree a plus. Experience with GPU clusters or AI‑driven data center environments. Methodical troubleshooting and technical leadership chops. Familiarity with Southaven...Shift workWeekend workAfternoon shift
- ...including model training, inference, vector search, and agentic orchestration. Build usage‑aligned pricing models that scale with compute, GPU, and AI service consumption. Conduct competitive analysis across AI platforms. Develop packaging and SKU strategy for AI platform and...Permanent employmentFlexible hours
$146.4k - $263.6k
...other SREs, influence architecture decisions with product engineering teams, and shape SRE practices for AI inference workloads and GPU infrastructure at scale. As a Senior II Site Reliability Engineer, you will be responsible for: Taking ownership of observability...Work experience placementWork at office
