Software Development Engineer — GPU Fleet Management & AI Infrastructure
Advanced Micro Devices Inc
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE :AMD is looking for an experienced software engineer to help build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure.Fleet Manager provides scheduling, workload orchestration, hardware health management, interactive development environments, and model-serving capabilities across GPU clusters. You will design and build production systems spanning distributed control planes, Kubernetes, GPU scheduling, inference infrastructure, and developer-facing APIs and tools.You will join a team working at the intersection of systems software, cloud infrastructure, and accelerated computing. Your work will directly influence how engineers and customers train, serve, debug, and operate workloads on current and future AMD GPU platforms.THE PERSON :The ideal candidate is a hands-on systems engineer who enjoys solving complex infrastructure problems and turning them into reliable, easy-to-use products.You have strong technical judgment, can reason about distributed-system failure modes, and are comfortable working across service, cluster, and hardware boundaries. You communicate clearly, collaborate effectively across organizations, and can lead substantial projects from architecture through production deployment.You care deeply about correctness, security, operability, and the experience of both end users and platform operators.KEY RESPONSIBILITIES :Design and develop Fleet Manager’s distributed control-plane services, APIs, schedulers, inference gateway, and command-line tools.Build reliable orchestration for GPU training, inference, custom jobs, and interactive development workloads.Develop scalable scheduling and admission-control capabilities, including priority, fairness, topology-aware placement, quotas, backfilling, and multi-node workload coordination.Implement durable reconciliation, lifecycle management, retries, idempotency, and recovery across PostgreSQL and external execution systems.Integrate Fleet Manager with Kubernetes and technologies such as Kueue, JobSet, container runtimes, storage systems, and observability platforms.Help evolve Fleet Manager into a portable orchestration layer capable of supporting Kubernetes, Slurm, Spur, and future execution environments.Develop GPU health, diagnostics, quarantine, and controlled-remediation capabilities using ROCm and AMD hardware telemetry.Build secure multi-tenant infrastructure with strong authentication, authorization, workload isolation, auditing, rate limiting, and least-privilege defaults.Improve the reliability and performance of AI inference services, including routing, streaming, load shedding, health detection, and usage metering.Define and maintain stable APIs, data models, compatibility contracts, and operational procedures.Diagnose complex failures across distributed services, Kubernetes, networking, storage, GPU runtimes, drivers, and hardware.Develop automated unit, integration, failure-injection, and production-readiness tests.Work with AMD architecture, driver, platform, security, and machine-learning software teams to enable current and future GPU products.Participate in new GPU, system, cluster, and software-stack bring-up.Provide technical leadership through design reviews, code reviews, mentoring, and cross-functional problem solving.PREFERRED EXPERIENCE :Strong systems-software development experience in Rust, C++, Go, or a comparable language. Production Rust experience is highly desirable.Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure.Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.Experience building reliable services using REST, streaming, WebSocket, or gRPC APIs.Experience with Kubernetes internals, controllers, operators, scheduling, resource management, or custom resources.Familiarity with workload scheduling technologies such as Kueue, JobSet, Slurm, or other batch and cluster schedulers.Experience with PostgreSQL-backed services, schema evolution, transactions, leader election, and optimistic concurrency.Experience developing command-line tools and stable, user-focused APIs.Understanding of container security, multi-tenant isolation, authentication, authorization, and secrets management.Experience with production observability, including metrics, structured logging, tracing, alerting, and incident diagnosis.Ability to write high-quality, maintainable code with careful attention to correctness, testing, and operational behavior.Experience with source control, continuous integration, automated testing, profiling, and debugging tools.Demonstrated ability to lead technically challenging projects and collaborate across organizational boundaries.Effective written and verbal communication skills.Experience in one or more of the following areas would be beneficial but is not required:AMD GPU architecture, ROCm, HIP, amd-smi, RCCL, or GPU device pluginsDistributed AI training and multi-node collective communicationLarge-model inference using platforms such as vLLM, SGLang, PyTorch, or similar runtimesGPU topology, capacity management, performance analysis, and hardware diagnosticsKubernetes networking, storage, admission control, and workload isolationOpenAI-compatible inference APIs, request routing, streaming, and rate limitingHigh-performance shared storage and large model or checkpoint managementBare-metal, virtualized, and cloud GPU infrastructureProduction security and threat modeling for multi-tenant compute platformsPREFERRED ACADEMIC CREDENTIALS :Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.This role is not eligible for visa sponsorship.#LI-G11 #LI-HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$108k - $162k
...highly skilled Sr. Systems & Infrastructure Engineer to join a dynamic, security-... ...cloud, and modern cloud-managed environments. This role spans... ...Microsoft 365 administration, AI-augmented tooling, and endpoint... ...across server and endpoint fleets.Contribute to identity...FleetPermanent employmentFull time$146.3k - $306.4k
...scale systems software, firmware... ...integration, and fleet automation... .... Champions engineering excellence:... ..., resource management, concurrency... ...together the data, infrastructure,... ...care. And with AI embedded across... ...Software Development:Contribute to... ...networking, and GPU subsystems,...FleetTemporary workFlexible hoursShift work- ...generation computing experiences—from AI and data centers, to PCs, gaming and embedded... ...are hiring a AI Research Scientist - Infrastructure Engineer, Reinforcement Learning, to own... ...and researcher‑facing APIs across large GPU fleets. You make RL scientists productive by...Fleet
- ...Cloud, is a leader in AI cloud infrastructure serving tens of... ...superintelligence. One person, one GPU.If you'd like to... ...Infrastructure Engineering organization forges the... ...seasoned Staff Storage Software Engineer with deep... ...observability, compute, and fleet engineering teams to...FleetWork at officeLocal areaWork from homeFlexible hours
$248k - $391k
...potential of AI to define the... ...in which our GPU acts as the brains... ...a Principal Software Engineer to join our Configuration Management team and define... ...of enterprise infrastructure automation, configuration... ...fleets across data centers... ...leading the development and enterprise...FleetFull time$160k - $253k
...unlimited potential of AI to define the... ...in which our GPU acts as the... ...together facilities infrastructure, hardware, software, simulation, and... ...Technical Marketing Engineer to show and... ...remediation, capacity management, and security.... ..., and fleet health.Ability to...FleetFull time$152k - $241.5k
...Isaac Applications Engineering team and help... ...platform for Physical AI robots —... ...all of it: is our software ready to be used... ...it, and build the infrastructure that keeps it true... ...hardware.Experience managing GPU-backed CI infrastructure... ...-hosted runner fleets.Your base salary...FleetFull timeLive inNight shift- ...Cloud, is a leader in AI cloud infrastructure serving tens of... ...superintelligence. One person, one GPU.If you'd like to... ...is currently Tuesday.Engineering at Lambda is... ...for system deployment, management and maintenance.What... ...incident response using fleet management toolsParticipate...FleetWork at officeLocal areaWork from homeFlexible hours
$272k - $431.25k
...as a Principal Software Engineer for DGX Cloud.... ...and AI and want to help... ...platform for ML/AI infrastructure? Do you thrive... ...DSX Kubernetes Fleet team within NVIDIA... ..., lifecycle management, and... ...to our massive GPU accelerated container... ...simplify the development, deployment, and...FleetFull timeWork experience placement$272k - $431.25k
...unlimited potential of AI to define the... ...in which our GPU acts as the... ...Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems... ...dependable, manageable, and... ...safely at rack and fleet scale. Build open...FleetFull timeRemote workShift work$176k - $276k
...’s Global Network Infrastructure (GNI) organization... ...enablement. We build software and automation to... ..., scaled, and managed across environments... ...a hands-on senior engineer to own the lifecycle... ...region Kubernetes fleets, including fleet-... ...vacancy. NVIDIA uses AI tools in its...FleetFull timeRemote workWeekend work$208k - $327.75k
...Senior Product Manager to architect... ...Enterprise AI. While the... ...role, own the software-defined... ...SuperPODs. Lead the development of products... ...like GPU partitioning... ...self-healing infrastructure. Thoughtfully... ...that keep the fleet at peak... ...of multiple engineering fields. As you...FleetFull timeNight shift$169.78k - $351k
...into a shared-mobility fleet will generate immense... ...Senior/Sr. Staff AI Infrastructure Engineer , Inference & Optimization... ...across target GPU architectures.... ...in Computer Science, Software Engineering, Systems... ...and memory bandwidth management. Demonstrated ability...FleetFull time- ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is... ...construction, equipment management, and daily operations... ...Scheduling & Orchestration Engineer to lead the workload... ...for eliminating "GPU stranding" and... ...our expensive compute fleets. This role is pivotal...FleetFull timeLocal area
- ...computing, cloud, and AI. Whether you’re... ...systems software for AMD GPUs with... ...runtimes, low-level GPU software, firmware... ...system interfaces, engineering practices, and validation... ...compiler infrastructure, GPU architecture... ...maintainability, or development velocity of...Local area
$193.3k - $261.5k
...in the world, and MONA (Management, Out-of-Band Networks,... ...data center powering the AI revolution.We build the in-house software that solves our hardest... ...for a Senior Software Development Engineer to be a technical... ...build these systems at fleet scale. If you want the...FleetInternshipLocal areaFlexible hoursDay shift$168k - $264.5k
...unlimited potential of AI to define the next... ...era in which our GPU acts as the brains... ...Join NVIDIA’s CAD Infrastructure team and be part... ...a next-generation software platform for... ...endeavor, and we need engineers who thrive in a hands... ...active spec development to identify integration...Full time- ...TeamNetflix Cloud is the infrastructure every Netflix service... ...it, the Infrastructure Management org builds the platforms engineers use to provision, configure... ...serving as the backbone for AI/agentic workloads. And... ...environment as agentic development outpaces governance. We'...Hourly payFull timeImmediate startFlexible hours
$160k - $275k
...world’s best AI models run as... ...seeking System Software Engineer to join our team... ..., memory management. High-speed connectivity... ...level: fleet-wide health aggregation... ...Linux systems development experience,... ...hardware or infrastructure... ...experience for GPU or accelerator...FleetDaily paidFull timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours$272k - $431.25k
...looking for a Principal Software Engineer to join our CSP... ...technical focal point for GPU firmware and GPU system... ...ensure they can reliably manage, update, and operate NVIDIA GPU firmware at fleet scale. You will drive work... ...vacancy. NVIDIA uses AI tools in its recruiting...FleetFull timeRemote work$160.36k - $240.54k
...profound opportunity for AI to drive positive... ...and logistics fleets to personal... ...RoleThe Autonomy ML Infrastructure team is responsible... ...Work with autonomy engineers to optimize, validate... ..., high-quality software to increase our confidence... ..., and optimizing GPU ML compilers &...FleetWork experience placementImmediate startFlexible hours- ...world's largest AI chip, 56 times larger... ...faster than GPU-based hyperscale... ...Wafer-Scale Engine.We are hiring a Software Engineer to productionize... ...-scale AMD GPU infrastructure, to make this... ...new accelerator fleet, and drive... ...checking, capacity-management, and failure-recovery...Fleet
$160.36k - $240.54k
...opportunity for AI to drive... ...and logistics fleets to personal vehicles... ...is seeking a Software Engineer with expertise... ...large-scale infrastructure, workload orchestration... ...feature management to accelerate... ...Nuro Driver development lifecycle.About... ...thousands of GPU/CPU nodes across...FleetImmediate startFlexible hours$160.36k - $240.54k
...opportunity for AI to drive... ...and logistics fleets to personal vehicles... ...About the RoleOur software team is growing... ...for talented engineers to join us and... ...and Technical Infrastructure.Data Platform:... ...comprehensive management system for Nuro... ...modalities (CPU, GPU, FPGA) etc.At Nuro...FleetImmediate startFlexible hours$152k - $228k
...opportunity for AI to drive... ...and logistics fleets to personal... ...of the software and hardware... ...will own the infrastructure that makes this... ...much much more.Engineers across the... ...on the GPU?How does a new... ...responsible for development, integration... ...units, or managed bare-metal infrastructure...FleetTemporary workImmediate startFlexible hours$127.1k - $185k
...talented early-career engineer to join our... ...2 distributed AI/ML systems. You'll work on software that enables... ...across massive GPU clusters, developing... ...learning infrastructure - building the... ...concepts, memory management) 2/ Parallel Computer... ...with Linux development environments...InternshipLocal areaFlexible hours$193.93k - $352.29k
...profound opportunity for AI to drive positive... ...and logistics fleets to personal... ...About the RoleOur software team is growing, and... ...looking for talented engineers to join us and be... ...tracing tools and infrastructure (perf, eBPF, Perfetto... ...(x86, ARM, GPU, FPGA, etc)You have...FleetImmediate startFlexible hours$258k - $387k
...opportunity for AI to drive positive... ...robotaxis and logistics fleets to personal... ...RoleAs a Principal Software Engineer, you will help... ...of Nuro’s onboard infrastructure. We are looking for... ..., memory management, thread/process lifecycle... ....Experience with GPU programming,...FleetImmediate startFlexible hours$175k - $287k
...largest privately managed compute infrastructures in the world... ...providers. As a Staff Software Engineer on the Compute... ...daily, and the AI/ML... ...and efficiency, GPU compute optimization... ...for AI/ML, and fleet health at scale... ...software design, development, and algorithm-related...FleetFor contractorsWork experience placementWork at officeFlexible hours$198k - $326k
...scaling LinkedIn's AI model training, feature engineering and serving with... ...data infra, compute software, and hardware to... ...the power of our GPU fleet with thousands of... ...queries.Model Training Infrastructure: As an engineer on... ...in and guide the development of containerized...FleetFor contractorsWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Development Engineer — GPU Fleet Management & AI Infrastructure. Be the first to apply!
- software engineer part time San Jose, CA
- software developer intern San Jose, CA
- software developer fintech San Jose, CA
- ngo software engineer San Jose, CA
- intel software engineer San Jose, CA
- machine learning software engineer San Jose, CA
- senior software engineer remote San Jose, CA
- junior software developer remote San Jose, CA
- software engineer full time San Jose, CA
- software engineer staff San Jose, CA

