Software Development Engineer — GPU Fleet Management & AI Infrastructure
Advanced Micro Devices Inc
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE :AMD is looking for an experienced software engineer to help build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure.Fleet Manager provides scheduling, workload orchestration, hardware health management, interactive development environments, and model-serving capabilities across GPU clusters. You will design and build production systems spanning distributed control planes, Kubernetes, GPU scheduling, inference infrastructure, and developer-facing APIs and tools.You will join a team working at the intersection of systems software, cloud infrastructure, and accelerated computing. Your work will directly influence how engineers and customers train, serve, debug, and operate workloads on current and future AMD GPU platforms.THE PERSON :The ideal candidate is a hands-on systems engineer who enjoys solving complex infrastructure problems and turning them into reliable, easy-to-use products.You have strong technical judgment, can reason about distributed-system failure modes, and are comfortable working across service, cluster, and hardware boundaries. You communicate clearly, collaborate effectively across organizations, and can lead substantial projects from architecture through production deployment.You care deeply about correctness, security, operability, and the experience of both end users and platform operators.KEY RESPONSIBILITIES :Design and develop Fleet Manager’s distributed control-plane services, APIs, schedulers, inference gateway, and command-line tools.Build reliable orchestration for GPU training, inference, custom jobs, and interactive development workloads.Develop scalable scheduling and admission-control capabilities, including priority, fairness, topology-aware placement, quotas, backfilling, and multi-node workload coordination.Implement durable reconciliation, lifecycle management, retries, idempotency, and recovery across PostgreSQL and external execution systems.Integrate Fleet Manager with Kubernetes and technologies such as Kueue, JobSet, container runtimes, storage systems, and observability platforms.Help evolve Fleet Manager into a portable orchestration layer capable of supporting Kubernetes, Slurm, Spur, and future execution environments.Develop GPU health, diagnostics, quarantine, and controlled-remediation capabilities using ROCm and AMD hardware telemetry.Build secure multi-tenant infrastructure with strong authentication, authorization, workload isolation, auditing, rate limiting, and least-privilege defaults.Improve the reliability and performance of AI inference services, including routing, streaming, load shedding, health detection, and usage metering.Define and maintain stable APIs, data models, compatibility contracts, and operational procedures.Diagnose complex failures across distributed services, Kubernetes, networking, storage, GPU runtimes, drivers, and hardware.Develop automated unit, integration, failure-injection, and production-readiness tests.Work with AMD architecture, driver, platform, security, and machine-learning software teams to enable current and future GPU products.Participate in new GPU, system, cluster, and software-stack bring-up.Provide technical leadership through design reviews, code reviews, mentoring, and cross-functional problem solving.PREFERRED EXPERIENCE :Strong systems-software development experience in Rust, C++, Go, or a comparable language. Production Rust experience is highly desirable.Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure.Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.Experience building reliable services using REST, streaming, WebSocket, or gRPC APIs.Experience with Kubernetes internals, controllers, operators, scheduling, resource management, or custom resources.Familiarity with workload scheduling technologies such as Kueue, JobSet, Slurm, or other batch and cluster schedulers.Experience with PostgreSQL-backed services, schema evolution, transactions, leader election, and optimistic concurrency.Experience developing command-line tools and stable, user-focused APIs.Understanding of container security, multi-tenant isolation, authentication, authorization, and secrets management.Experience with production observability, including metrics, structured logging, tracing, alerting, and incident diagnosis.Ability to write high-quality, maintainable code with careful attention to correctness, testing, and operational behavior.Experience with source control, continuous integration, automated testing, profiling, and debugging tools.Demonstrated ability to lead technically challenging projects and collaborate across organizational boundaries.Effective written and verbal communication skills.Experience in one or more of the following areas would be beneficial but is not required:AMD GPU architecture, ROCm, HIP, amd-smi, RCCL, or GPU device pluginsDistributed AI training and multi-node collective communicationLarge-model inference using platforms such as vLLM, SGLang, PyTorch, or similar runtimesGPU topology, capacity management, performance analysis, and hardware diagnosticsKubernetes networking, storage, admission control, and workload isolationOpenAI-compatible inference APIs, request routing, streaming, and rate limitingHigh-performance shared storage and large model or checkpoint managementBare-metal, virtualized, and cloud GPU infrastructureProduction security and threat modeling for multi-tenant compute platformsPREFERRED ACADEMIC CREDENTIALS :Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.This role is not eligible for visa sponsorship.#LI-G11 #LI-HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$109k - $160k
...Essential Cloud for AI™. Built for... ...combines superior infrastructure performance with... ...About the role A Software Engineer contributes to... ...’ll partner with fleet, product, and hardware... ...to evolve our GPU performance... ...Python software development. Hands-on experience...FleetPermanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours- ...Cloud, is a leader in AI cloud infrastructure serving tens of... ...superintelligence. One person, one GPU.If you'd like to... ...Infrastructure Engineering organization forges the... ...seasoned Staff Storage Software Engineer with deep... ...observability, compute, and fleet engineering teams to...FleetWork at officeLocal areaWork from homeFlexible hours
$248k - $391k
...potential of AI to define the... ...in which our GPU acts as the brains... ...a Principal Software Engineer to join our Configuration Management team and define... ...of enterprise infrastructure automation, configuration... ...fleets across data centers... ...leading the development and enterprise...FleetFull time$184k - $287.5k
...networking, systems, and software to solve some of... ...System Software Engineer to join NVIDIA’s GPU Performance and Power Management Software team.... ...systems, and AI workloads. In... ..., pre-silicon development, validation, silicon... ...and automation infrastructure for low-level...SuggestedFull time$152k - $241.5k
...Isaac Applications Engineering team and help... ...platform for Physical AI robots —... ...all of it: is our software ready to be used... ...it, and build the infrastructure that keeps it true... ...hardware.Experience managing GPU-backed CI infrastructure... ...-hosted runner fleets.Your base salary...FleetFull timeLive inNight shift- ...Cloud, is a leader in AI cloud infrastructure serving tens of... ...superintelligence. One person, one GPU.If you'd like to... ...is currently Tuesday.Engineering at Lambda is... ...for system deployment, management and maintenance.What... ...incident response using fleet management toolsParticipate...FleetWork at officeLocal areaWork from homeFlexible hours
$272k - $431.25k
...unlimited potential of AI to define the... ...in which our GPU acts as the... ...Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems... ...dependable, manageable, and... ...safely at rack and fleet scale. Build open...FleetFull timeRemote workShift work$160k - $253k
...unlimited potential of AI to define the... ...in which our GPU acts as the... ...together facilities infrastructure, hardware, software, simulation, and... ...Technical Marketing Engineer to show and... ...remediation, capacity management, and security.... ..., and fleet health.Ability to...FleetFull time$200k - $400k
...Figure is an AI Robotics company developing a... ...senior-level backend engineer who has scaled high-throughput... ...around cloud infrastructure and real-time streaming... ...sensor data across robot fleets and user sessions.... ...processing, connection management, data transport, and low...FleetFull timeWork at office- ...Netflix Cloud is the infrastructure every Netflix service runs... ...it, the Infrastructure Management org builds the platforms engineers use to provision, configure... ...as the backbone for AI/agentic workloads. And we... ...environment as agentic development outpaces governance. We'...Hourly payFull timeImmediate startFlexible hours
$160.36k - $240.54k
...cutting-edge AI with automotive... ...and commercial fleets to personally... ...is seeking a Software Engineer with expertise... ...in large-scale infrastructure, workload orchestration... ...feature management to accelerate... ...Nuro Driver™ development lifecycle.... ...thousands of GPU/CPU nodes across...FleetFull time- ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed... ..., equipment management, and daily operations.... ...building an AI-operated GPU cloud — a global fleet of self-built and OEM-... ...As an entry-level Software Engineer on the SRE / Monitoring...FleetFull timeContract workInternshipLocal area
$152k - $228k
...cutting-edge AI with... ...commercial fleets to personally... ...aspects of the software and hardware... ...will own the infrastructure that makes this... ...much more. Engineers across the... ...contention on the GPU? How does... ...for development, integration... ...systemd units, or managed bare-metal...FleetFull timeTemporary work$132k - $198k
...combining cutting-edge AI with automotive-... ...and commercial fleets to personally owned... ...for a skilled engineer to help build and... ...and (OTA) update infrastructure. Our team, Fleet connectivity... ..., telemetry, and software updates, which is... ...OTA updates. Manage individual project...FleetFull time$160k - $240k
...cutting-edge AI with... ...commercial fleets to personally... ...self-motivated engineers to build the... ...generation onboard infrastructure for... ...execution and state management. We actively... ...AI-assisted development, leveraging... ...with other software teams to build... ...with GPU programming...FleetFull time- ...their needs, we are looking for an AI Devops Infrastructure Engineer/GPU Infrastructure Engineer. Job... ...operationalize infrastructure supporting AI development, experimentation, training,... ...training, and inference. Design and manage infrastructure across GPU and...Contract work
$175k - $287k
...largest privately managed compute infrastructures in the world... ...providers. As a Staff Software Engineer on the Compute... ...daily, and the AI/ML... ...and efficiency, GPU compute optimization... ...for AI/ML, and fleet health at scale... ...software design, development, and algorithm-related...FleetFull timeFor contractorsWork experience placementWork at officeFlexible hours$168k - $264.5k
...unlimited potential of AI to define the next... ...era in which our GPU acts as the brains... ...Join NVIDIA’s CAD Infrastructure team and be part... ...a next-generation software platform for... ...endeavor, and we need engineers who thrive in a hands... ...active spec development to identify integration...Full time- ...world's largest AI chip, 56 times larger... ...faster than GPU-based hyperscale... ...Wafer-Scale Engine.We are hiring a Software Engineer to productionize... ...-scale AMD GPU infrastructure, to make this... ...new accelerator fleet, and drive... ...checking, capacity-management, and failure-recovery...Fleet
$127.1k - $185k
...talented early-career engineer to join our... ...2 distributed AI/ML systems. You'll work on software that enables... ...across massive GPU clusters, developing... ...learning infrastructure - building the... ...concepts, memory management) 2/ Parallel Computer... ...with Linux development environments...InternshipLocal areaFlexible hours$193.93k - $352.29k
...profound opportunity for AI to drive positive... ...and logistics fleets to personal... ...the Role Our software team is growing, and... ...looking for talented engineers to join us and be... ...tracing tools and infrastructure (perf, eBPF, Perfetto... ...(x86, ARM, GPU, FPGA, etc) ~ You...FleetImmediate startFlexible hours$258k - $387k
...opportunity for AI to drive positive... ...robotaxis and logistics fleets to personal... ...As a Principal Software Engineer, you will help define... ...Nuro's onboard infrastructure. We are looking for... ..., memory management, thread/process lifecycle... ...Experience with GPU programming,...FleetImmediate startFlexible hours$300 per month
...vertically integrated AI infrastructure company built... ...Infrastructure Engineering (DCIE) team is fundamental... ...for Crusoe’s fleet GPU’s and data center... ...and motivated Software Engineer to join... ...is focused on the development of software for the management of a fleet of GPU...FleetTemporary work- ...in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related... ...vulnerability-management experience; familiarity... ...leveraging AI-assisted development tools to improve software development, automation... ...Experience managing fleet wide software deployments...FleetFull timeWorldwide
- ...Role :- Site Reliability Engineer (SRE) Infrastructure & Agentic Automation... ...infrastructure management and cutting-edge agentic AI tooling, building robust... ...reliability across massive fleet environments, we want... ...Chef (or Cinc) cookbook development, serverless execution...Fleet
$165.6k - $296.4k
...than 25%Profession: Software EngineeringDiscipline... ...Intelligence (AI) Infrastructure Engineering Systems team at Microsoft... ...and service lifecycle management. Working with AI... ...consistent paths from development to production within... ...as model inference, GPU computing, model serving...Ongoing contractLocal area3 days per week$119.8k - $234.7k
...than 25%Profession: Software EngineeringDiscipline... ...Intelligence (AI) Infrastructure Engineering Systems team at Microsoft... ...and service lifecycle management. Working with AI... ...consistent paths from development to production within... ...as model inference, GPU computing, model serving...Ongoing contractLocal area3 days per week$193.93k - $352.29k
...opportunity for AI to drive... ...and logistics fleets to personal vehicles... ...Role Our software team is growing... ...looking for talented engineers to join us and... ...and Technical Infrastructure. Data... ...comprehensive management system for Nuro... ...modalities (CPU, GPU, FPGA) etc....FleetImmediate startFlexible hours$182k - $242k
...The Essential Cloud for AI™. Built for pioneers... ...CoreWeave combines superior infrastructure performance with deep... ...looking for a Senior Engineer to be a driving force... ...real time, whether a GPU fleet, a fabric, or a... ...for engineers, product managers, and executives, but as...FleetPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$193.93k - $352.29k
...profound opportunity for AI to drive positive... ...and logistics fleets to personal... ...The Autonomy ML Infrastructure team is responsible... ...Work with autonomy engineers to optimize, validate... ...robust, high quality software to increase our... ...profiling, and optimizing GPU ML compilers &...FleetWork experience placementImmediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Development Engineer — GPU Fleet Management & AI Infrastructure. Be the first to apply!
- software engineer internship remote San Jose, CA
- senior software engineer ruby on rails San Jose, CA
- software developer positions San Jose, CA
- intermediate software engineer San Jose, CA
- agile software developer San Jose, CA
- software engineer intern San Jose, CA
- part time software developer San Jose, CA
- rust software engineer San Jose, CA
- software engineer internship San Jose, CA
- entry level software engineer remote San Jose, CA



