Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Development Engineer — GPU Fleet Management & AI Infrastructure

Advanced Micro Devices Inc

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE :AMD is looking for an experienced software engineer to help build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure.Fleet Manager provides scheduling, workload orchestration, hardware health management, interactive development environments, and model-serving capabilities across GPU clusters. You will design and build production systems spanning distributed control planes, Kubernetes, GPU scheduling, inference infrastructure, and developer-facing APIs and tools.You will join a team working at the intersection of systems software, cloud infrastructure, and accelerated computing. Your work will directly influence how engineers and customers train, serve, debug, and operate workloads on current and future AMD GPU platforms.THE PERSON :The ideal candidate is a hands-on systems engineer who enjoys solving complex infrastructure problems and turning them into reliable, easy-to-use products.You have strong technical judgment, can reason about distributed-system failure modes, and are comfortable working across service, cluster, and hardware boundaries. You communicate clearly, collaborate effectively across organizations, and can lead substantial projects from architecture through production deployment.You care deeply about correctness, security, operability, and the experience of both end users and platform operators.KEY RESPONSIBILITIES :Design and develop Fleet Manager’s distributed control-plane services, APIs, schedulers, inference gateway, and command-line tools.Build reliable orchestration for GPU training, inference, custom jobs, and interactive development workloads.Develop scalable scheduling and admission-control capabilities, including priority, fairness, topology-aware placement, quotas, backfilling, and multi-node workload coordination.Implement durable reconciliation, lifecycle management, retries, idempotency, and recovery across PostgreSQL and external execution systems.Integrate Fleet Manager with Kubernetes and technologies such as Kueue, JobSet, container runtimes, storage systems, and observability platforms.Help evolve Fleet Manager into a portable orchestration layer capable of supporting Kubernetes, Slurm, Spur, and future execution environments.Develop GPU health, diagnostics, quarantine, and controlled-remediation capabilities using ROCm and AMD hardware telemetry.Build secure multi-tenant infrastructure with strong authentication, authorization, workload isolation, auditing, rate limiting, and least-privilege defaults.Improve the reliability and performance of AI inference services, including routing, streaming, load shedding, health detection, and usage metering.Define and maintain stable APIs, data models, compatibility contracts, and operational procedures.Diagnose complex failures across distributed services, Kubernetes, networking, storage, GPU runtimes, drivers, and hardware.Develop automated unit, integration, failure-injection, and production-readiness tests.Work with AMD architecture, driver, platform, security, and machine-learning software teams to enable current and future GPU products.Participate in new GPU, system, cluster, and software-stack bring-up.Provide technical leadership through design reviews, code reviews, mentoring, and cross-functional problem solving.PREFERRED EXPERIENCE :Strong systems-software development experience in Rust, C++, Go, or a comparable language. Production Rust experience is highly desirable.Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure.Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.Experience building reliable services using REST, streaming, WebSocket, or gRPC APIs.Experience with Kubernetes internals, controllers, operators, scheduling, resource management, or custom resources.Familiarity with workload scheduling technologies such as Kueue, JobSet, Slurm, or other batch and cluster schedulers.Experience with PostgreSQL-backed services, schema evolution, transactions, leader election, and optimistic concurrency.Experience developing command-line tools and stable, user-focused APIs.Understanding of container security, multi-tenant isolation, authentication, authorization, and secrets management.Experience with production observability, including metrics, structured logging, tracing, alerting, and incident diagnosis.Ability to write high-quality, maintainable code with careful attention to correctness, testing, and operational behavior.Experience with source control, continuous integration, automated testing, profiling, and debugging tools.Demonstrated ability to lead technically challenging projects and collaborate across organizational boundaries.Effective written and verbal communication skills.Experience in one or more of the following areas would be beneficial but is not required:AMD GPU architecture, ROCm, HIP, amd-smi, RCCL, or GPU device pluginsDistributed AI training and multi-node collective communicationLarge-model inference using platforms such as vLLM, SGLang, PyTorch, or similar runtimesGPU topology, capacity management, performance analysis, and hardware diagnosticsKubernetes networking, storage, admission control, and workload isolationOpenAI-compatible inference APIs, request routing, streaming, and rate limitingHigh-performance shared storage and large model or checkpoint managementBare-metal, virtualized, and cloud GPU infrastructureProduction security and threat modeling for multi-tenant compute platformsPREFERRED ACADEMIC CREDENTIALS :Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.This role is not eligible for visa sponsorship.#LI-G11 #LI-HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Software Development Engineer — GPU Fleet Management & AI Infrastructure in San Jose, CA vacancy
  • $108k - $162k

     ...highly skilled Sr. Systems & Infrastructure Engineer to join a dynamic, security-...  ...cloud, and modern cloud-managed environments. This role spans...  ...Microsoft 365 administration, AI-augmented tooling, and endpoint...  ...across server and endpoint fleets.Contribute to identity... 
    Fleet
    Permanent employment
    Full time

    Onto Innovation

    Milpitas, CA
    3 days ago
  • $146.3k - $306.4k

     ...scale systems software, firmware...  ...integration, and fleet automation...  .... Champions engineering excellence:...  ..., resource management, concurrency...  ...together the data, infrastructure,...  ...care. And with AI embedded across...  ...Software Development:Contribute to...  ...networking, and GPU subsystems,... 
    Fleet
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    3 days ago
  •  ...generation computing experiences—from AI and data centers, to PCs, gaming and embedded...  ...are hiring a AI Research Scientist - Infrastructure Engineer, Reinforcement Learning, to own...  ...and researcher‑facing APIs across large GPU fleets. You make RL scientists productive by... 
    Fleet

    AMD

    Santa Clara, CA
    16 hours ago
  •  ...Cloud, is a leader in AI cloud infrastructure serving tens of...  ...superintelligence. One person, one GPU.If you'd like to...  ...Infrastructure Engineering organization forges the...  ...seasoned Staff Storage Software Engineer with deep...  ...observability, compute, and fleet engineering teams to... 
    Fleet
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    16 hours ago
  • $248k - $391k

     ...potential of AI to define the...  ...in which our GPU acts as the brains...  ...a Principal Software Engineer to join our Configuration Management team and define...  ...of enterprise infrastructure automation, configuration...  ...fleets across data centers...  ...leading the development and enterprise... 
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $160k - $253k

     ...unlimited potential of AI to define the...  ...in which our GPU acts as the...  ...together facilities infrastructure, hardware, software, simulation, and...  ...Technical Marketing Engineer to show and...  ...remediation, capacity management, and security....  ..., and fleet health.Ability to... 
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...Isaac Applications Engineering team and help...  ...platform for Physical AI robots —...  ...all of it: is our software ready to be used...  ...it, and build the infrastructure that keeps it true...  ...hardware.Experience managing GPU-backed CI infrastructure...  ...-hosted runner fleets.Your base salary... 
    Fleet
    Full time
    Live in
    Night shift

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Cloud, is a leader in AI cloud infrastructure serving tens of...  ...superintelligence. One person, one GPU.If you'd like to...  ...is currently Tuesday.Engineering at Lambda is...  ...for system deployment, management and maintenance.What...  ...incident response using fleet management toolsParticipate... 
    Fleet
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $272k - $431.25k

     ...as a Principal Software Engineer for DGX Cloud....  ...and AI and want to help...  ...platform for ML/AI infrastructure? Do you thrive...  ...DSX Kubernetes Fleet team within NVIDIA...  ..., lifecycle management, and...  ...to our massive GPU accelerated container...  ...simplify the development, deployment, and... 
    Fleet
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

     ...unlimited potential of AI to define the...  ...in which our GPU acts as the...  ...Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems...  ...dependable, manageable, and...  ...safely at rack and fleet scale. Build open... 
    Fleet
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $176k - $276k

     ...’s Global Network Infrastructure (GNI) organization...  ...enablement. We build software and automation to...  ..., scaled, and managed across environments...  ...a hands-on senior engineer to own the lifecycle...  ...region Kubernetes fleets, including fleet-...  ...vacancy. NVIDIA uses AI tools in its... 
    Fleet
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    16 hours ago
  • $208k - $327.75k

     ...Senior Product Manager to architect...  ...Enterprise AI. While the...  ...role, own the software-defined...  ...SuperPODs. Lead the development of products...  ...like GPU partitioning...  ...self-healing infrastructure. Thoughtfully...  ...that keep the fleet at peak...  ...of multiple engineering fields. As you... 
    Fleet
    Full time
    Night shift

    Nvidia

    Santa Clara, CA
    16 hours ago
  • $169.78k - $351k

     ...into a shared-mobility fleet will generate immense...  ...Senior/Sr. Staff AI Infrastructure Engineer , Inference & Optimization...  ...across target GPU architectures....  ...in Computer Science, Software Engineering, Systems...  ...and memory bandwidth management. Demonstrated ability... 
    Fleet
    Full time

    DiDi Labs

    San Jose, CA
    3 days ago
  •  ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is...  ...construction, equipment management, and daily operations...  ...Scheduling & Orchestration Engineer to lead the workload...  ...for eliminating "GPU stranding" and...  ...our expensive compute fleets. This role is pivotal... 
    Fleet
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    20 days ago
  •  ...computing, cloud, and AI. Whether you’re...  ...systems software for AMD GPUs with...  ...runtimes, low-level GPU software, firmware...  ...system interfaces, engineering practices, and validation...  ...compiler infrastructure, GPU architecture...  ...maintainability, or development velocity of... 
    Local area

    AMD

    San Jose, CA
    3 days ago
  • $193.3k - $261.5k

     ...in the world, and MONA (Management, Out-of-Band Networks,...  ...data center powering the AI revolution.We build the in-house software that solves our hardest...  ...for a Senior Software Development Engineer to be a technical...  ...build these systems at fleet scale. If you want the... 
    Fleet
    Internship
    Local area
    Flexible hours
    Day shift

    Amazon

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

     ...unlimited potential of AI to define the next...  ...era in which our GPU acts as the brains...  ...Join NVIDIA’s CAD Infrastructure team and be part...  ...a next-generation software platform for...  ...endeavor, and we need engineers who thrive in a hands...  ...active spec development to identify integration... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...TeamNetflix Cloud is the infrastructure every Netflix service...  ...it, the Infrastructure Management org builds the platforms engineers use to provision, configure...  ...serving as the backbone for AI/agentic workloads. And...  ...environment as agentic development outpaces governance. We'... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Los Gatos, CA
    3 days ago
  • $160k - $275k

     ...world’s best AI models run as...  ...seeking System Software Engineer to join our team...  ..., memory management. High-speed connectivity...  ...level: fleet-wide health aggregation...  ...Linux systems development experience,...  ...hardware or infrastructure...  ...experience for GPU or accelerator... 
    Fleet
    Daily paid
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    13 hours ago
  • $272k - $431.25k

     ...looking for a Principal Software Engineer to join our CSP...  ...technical focal point for GPU firmware and GPU system...  ...ensure they can reliably manage, update, and operate NVIDIA GPU firmware at fleet scale. You will drive work...  ...vacancy. NVIDIA uses AI tools in its recruiting... 
    Fleet
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $160.36k - $240.54k

     ...profound opportunity for AI to drive positive...  ...and logistics fleets to personal...  ...RoleThe Autonomy ML Infrastructure team is responsible...  ...Work with autonomy engineers to optimize, validate...  ..., high-quality software to increase our confidence...  ..., and optimizing GPU ML compilers &... 
    Fleet
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    16 hours ago
  •  ...world's largest AI chip, 56 times larger...  ...faster than GPU-based hyperscale...  ...Wafer-Scale Engine.We are hiring a Software Engineer to productionize...  ...-scale AMD GPU infrastructure, to make this...  ...new accelerator fleet, and drive...  ...checking, capacity-management, and failure-recovery... 
    Fleet

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $160.36k - $240.54k

     ...opportunity for AI to drive...  ...and logistics fleets to personal vehicles...  ...is seeking a Software Engineer with expertise...  ...large-scale infrastructure, workload orchestration...  ...feature management to accelerate...  ...Nuro Driver development lifecycle.About...  ...thousands of GPU/CPU nodes across... 
    Fleet
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    16 hours ago
  • $160.36k - $240.54k

     ...opportunity for AI to drive...  ...and logistics fleets to personal vehicles...  ...About the RoleOur software team is growing...  ...for talented engineers to join us and...  ...and Technical Infrastructure.Data Platform:...  ...comprehensive management system for Nuro...  ...modalities (CPU, GPU, FPGA) etc.At Nuro... 
    Fleet
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago
  • $152k - $228k

     ...opportunity for AI to drive...  ...and logistics fleets to personal...  ...of the software and hardware...  ...will own the infrastructure that makes this...  ...much much more.Engineers across the...  ...on the GPU?How does a new...  ...responsible for development, integration...  ...units, or managed bare-metal infrastructure... 
    Fleet
    Temporary work
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    16 hours ago
  • $127.1k - $185k

     ...talented early-career engineer to join our...  ...2 distributed AI/ML systems. You'll work on software that enables...  ...across massive GPU clusters, developing...  ...learning infrastructure - building the...  ...concepts, memory management) 2/ Parallel Computer...  ...with Linux development environments... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $193.93k - $352.29k

     ...profound opportunity for AI to drive positive...  ...and logistics fleets to personal...  ...About the RoleOur software team is growing, and...  ...looking for talented engineers to join us and be...  ...tracing tools and infrastructure (perf, eBPF, Perfetto...  ...(x86, ARM, GPU, FPGA, etc)You have... 
    Fleet
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    16 hours ago
  • $258k - $387k

     ...opportunity for AI to drive positive...  ...robotaxis and logistics fleets to personal...  ...RoleAs a Principal Software Engineer, you will help...  ...of Nuro’s onboard infrastructure. We are looking for...  ..., memory management, thread/process lifecycle...  ....Experience with GPU programming,... 
    Fleet
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    16 hours ago
  • $175k - $287k

     ...largest privately managed compute infrastructures in the world...  ...providers. As a Staff Software Engineer on the Compute...  ...daily, and the AI/ML...  ...and efficiency, GPU compute optimization...  ...for AI/ML, and fleet health at scale...  ...software design, development, and algorithm-related... 
    Fleet
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    4 days ago
  • $198k - $326k

     ...scaling LinkedIn's AI model training, feature engineering and serving with...  ...data infra, compute software, and hardware to...  ...the power of our GPU fleet with thousands of...  ...queries.Model Training Infrastructure: As an engineer on...  ...in and guide the development of containerized... 
    Fleet
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Development Engineer — GPU Fleet Management & AI Infrastructure. Be the first to apply!