Lead AI Systems Engineer: GPU Clusters & AI Ops
Semiconductor Engineering
Semiconductor Engineering in San Jose, CA seeks a senior AI Systems Engineer to lead the lifecycle of its AI infrastructure. This hands-on individual contributor will architect, build and operate high-performance GPU clusters, and drive deployment and optimization of advanced AI models and agentic services. The ideal candidate will own the end-to-end stack, ensuring reliability, scalability and seamless integration across platforms to support cutting-edge AI initiatives. #J-18808-Ljbffr Semiconductor Engineering
- AMD is seeking an AI Systems Engineer to design, deploy, and manage HPC/AI infrastructure, GPU clusters, and AI workload schedulers. You will collaborate across teams to deliver scalable, high-performance AI services on AMD hardware, with a focus on end-to-end reliability...Suggested
- AMD in San Jose, CA is seeking an AI Systems Engineer to join our IT compute platforms team. You will design, deploy, and manage HPC infrastructure, GPU clusters, and AI workload schedulers to enable scalable AI services on AMD hardware. You should have passion for large...Suggested
- ...computing, cloud, and AI. Whether you’re designing... ...of large-scale AI/ML clustered infrastructure. You will... ...team of multi-disciplined engineers that operates across... ...fundamentals: Linux operating systems, networking,... ...Kubernetes for HPC/AI (GPU operators, device plugins...SuggestedFlexible hours
- Semiconductor Engineering seeks an AI Systems Engineer in San Jose to lead the design, deployment, and optimization of our AI infrastructure. This senior IC will... ...through operations and support, including GPU clusters and agentic services. You will build and optimize...Suggested
- Bitdeer Technologies Group is seeking an L1 NOC/US Data Center operator to support NeoCloud's GPU DCs during 8AM-8PM PST shifts. You will monitor GPU clusters, networks, and storage, respond to alerts, and execute runbooks for common incidents across shore-to-APAC handoffs...SuggestedShift workNight shift
- AMD is seeking a Staff Software Developer to advance AI software on GPUs, focusing on high-performance C++ and low-level... ...-software co-design. The role requires deep expertise in GPU programming, LLMs, and AI systems, plus capability to mentor others and drive technical...
$176k - $276k
...looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure... ...intelligence and GPU computing. Provide insights on at-scale system design and tuning mechanisms... ...workflows and develop new, leading differentiated solutions. You...Full time- ...high-performance computing, cloud, and AI. Whether you’re designing next-gen processors... ...forward.THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms... ...Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON...
$160k - $198k
...team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and manage... ...specialized AI-centric bare-metal and GPU clouds (Nebius AI Cloud).Cloud-Native Orchestration... ...), paired with cloud-agnostic cluster abstractors like SkyPilot to manage...Local area$190k - $237k
...team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage... ...specialized AI-centric bare-metal and GPU clouds (Nebius AI Cloud).Cloud-Native Orchestration... ...), paired with cloud-agnostic cluster abstractors like SkyPilot to manage...Local area- A leading AI technology firm in California is seeking an experienced Senior Software Engineer to develop and optimize AI infrastructure software using state-of-the-art GPU systems. Candidates should have a Bachelor's degree in a technical field and a minimum of 5 years...
- NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models...
- ...seeking a highly skilled and experienced AI Systems Engineer to join our team. This is a hands-on,... ...role that will be pivotal in leading the development, operations, and support... ...architecting and building high-performance GPU clusters to deploying and optimizing our most advanced...
$144k - $180k
...team members. What You’ll Do As a Senior AI Systems Engineer, you will architect, deploy, and manage... ...specialized AI-centric bare-metal and GPU clouds (Nebius AI Cloud). Cloud-Native... ...Kubernetes), paired with cloud-agnostic cluster abstractors like SkyPilot to manage...Local area- ...discovery to powering AI and the... ...breakthroughs, or bringing leading edge products to market... ...platform that lets engineers ask natural-language questions across GPU design knowledge,... ..., RAG, agentic systems, evaluation, and efficient... ...observability, and cluster orchestration....Worldwide
- ...Austin, TX is seeking a Technical Marketing Engineer (TME) within the Software Product Management organization for AMD’s Data Center GPU Business Unit. You will shape the customer... ...teams design and deploy AMD-powered GPU clusters and networks, while producing high-quality...
- ...generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...looking for an influential software engineer who is passionate about improving... ...performance from the lowest-level GPU kernels to large-scale...
$144k - $180k
Archer Aviation in San Jose seeks a Senior AI Systems Engineer to architect, deploy, and manage the critical infrastructure for large-scale AI model training and inference, ensuring robust, low-latency performance. You will work with researchers and software engineers to...- AMD in San Jose, CA is hiring a hands-on ML engineer to lead GoldenEye’s retrieval, ranking, and answer-quality architecture. You will scale a prototype into a production platform, setting technical direction and mentoring engineers across hardware, software, and security...
- ...AI SYSTEMS ENGINEERAt AMD, we believe technology can change lives for the... ...ROLEAMD is seeking an AI Systems Engineer to help develop and optimize machine... ..., and enabling industry-leading AI inference performance across AMD NPU and GPU platforms.You will collaborate closely...Worldwide
$131k - $175k
...awards, such as Best Engineering Team, Best Company... ...world’s largest AI and cloud deployments. You will lead both technical design... ...scale AI and cloud clusters.What You’ll Do Lead... ...into high-density GPU environments, ensuring... ...and cabling systems, including fiber/copper...Remote workFlexible hours- ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture... ...a Technical Marketing Engineer (TME) within the... ...for AMD’s Data Center GPU Business Unit, you will... ...operate AMD-powered GPU clusters & networks to unlock performance...
- AMD seeks an AI Systems Engineer to advance ML workloads on AMD AI accelerators, bridging hardware and software from kernel design to production inference across NPU and GPU platforms. You will collaborate with compiler, runtime, silicon, and architecture teams, delivering...
$184k - $287.5k
...into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers... ...looking for talented Software Engineers to work on developing and deploying... ...building AI-driven software systems, ideally applied to complex...- AMD is seeking an AI Systems Engineer in San Jose to develop and optimize ML workloads on next‑gen AMD accelerators. You will design high‑performance... ..., optimize dataflow, and enable AI inference across NPU and GPU platforms. You will work across compiler, runtime, silicon,...
- ...Job Title: Senior AI/ML Ops Engineer Location: Cupertino Please share profiles for the... ...to engineering teams and leadership. Lead the triage and resolution of highly complex... ...experience building evaluation systems for ML or LLM applications: test harnesses...
- ...Micro Devices in San Jose, CA, seeks a hands-on PMTS-level ML engineer to lead GoldenEye’s retrieval, ranking, and answer-quality architecture... ...frameworks, ensuring accurate, cited answers with scalable, GPU-accelerated workloads and cross-team #J-18808-Ljbffr Advanced...
$184k - $287.5k
NVIDIA in Santa Clara, CA is seeking a senior systems/C++/CUDA engineer to develop first-of-its-kind storage and IO acceleration libraries. You will... ...speed-of-light performance, collaboration, and cutting-edge GPU storage technologies. Applications accepted until August 24,...- The era of pervasive AI has arrived. In this era, organizations... ...talented and driven ML performance engineer to optimize and scale state-of... ...gap between deep learning and systems performance, collaborating... ..., vLLM, or TensorRT. Strong GPU programming skills (CUDA, Triton...Full timeTemporary workLocal areaFlexible hours
- ...experiences—from AI and data centers,... ...gaming and embedded systems. Grounded in a culture... ...in enhancing GPU kernels, deep learning... ...and advanced engineering principles to drive... ...Seeking an Industry Leading Expert C++... ...utilization across clusters. Compiler Optimization...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead AI Systems Engineer: GPU Clusters & AI Ops. Be the first to apply!
- lead app. developer San Jose, CA
- lead infrastructure engineer San Jose, CA
- lead network engineer San Jose, CA
- lead engineer San Jose, CA
- lead operating engineer San Jose, CA
- lead system engineer San Jose, CA
- machine learning ai engineer San Jose, CA
- ai developer San Jose, CA
- senior ai engineer San Jose, CA
- ai engineer San Jose, CA


