Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Software Development Engineer  - SGLang and Inference Stack

Advanced Micro Devices Inc

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:    As a core member of the team, you will play a pivotal role in optimizing and developing deep learning frameworks for AMD GPUs. Your work will be instrumental in enhancing GPU kernel performance, accelerating deep learning models, and enabling RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node systems. You will collaborate across internal GPU software teams and engage with open-source communities to integrate and optimize cutting-edge compiler technologies and drive upstream contributions that benefit AMD’s AI software ecosystem. THE PERSON:   Skilled engineer with strong technical and analytical expertise in GPGPU C++, Triton, TileLang or DSL development within Linux environments. The ideal candidate will thrive in both collaborative team settings and independent work, with the ability to define goals, manage development efforts, and deliver high-quality solutions. Strong problem-solving skills, a proactive approach, and a keen understanding of software engineering best practices are essential. KEY RESPONSIBILITIES:  Optimize Deep Learning Frameworks: Enhance performance of frameworks like TensorFlow, PyTorch, and SGLang on AMD GPUs via upstream contributions in open-source repositories. Develop and Optimize Deep Learning Models: Profile, analyze, code change and tune large-scale training and inference models for optimal performance on AMD hardware. Day-0 supports to many SOTA models, DeepSeek 3.2, Kimi K2.5, etc.GPU Kernel Development: Design, implement, and optimize high-performance GPU kernels using HIP, Triton, TileLang or other DSLs for AI operator efficiency. Collaborate with GPU Library and Compiler Teams: Work closely with internal compiler and GPU math library teams to integrate, optimize and align kernel-level optimizations with full-stack performance goals. Initiate and help with different level codegen optimizations.Contribute to SGLang Development: Support optimization, feature development, and scaling of the SGLang framework across AMD GPU platforms for LLM, multimodal serving and RL-training. Distributed System Optimization: Tune and scale performance across both multi-GPU (scale-up) and multi-node (scale-out) environments, including inference parallelism, prefill-decode disaggregation, Wide-EP and collective communication strategies. Graph Compiler Integration: Integrate and optimize runtime execution through graph compilers such as XLA, TorchDynamo, or custom pipelines. Open-Source Collaboration: Partner with external maintainers to understand framework needs, propose optimizations, and upstream contributions effectively. Apply Engineering Best Practices: Leverage modern software engineering practices in debugging, profiling, test-driven development, and CI/CD integration. PREFERRED EXPERIENCE:  Strong Programming Skills: Proficient in C++ and/or Python (PyTorch, Triton, TileLang), with demonstrated ability to code, debug, profile, and optimize performance-critical code. SGLang and LLM Optimization: Hands-on experience with SGLang or similar LLM inference frameworks is highly preferred. Compiler and GPU Architecture Knowledge: Background in compiler design or familiarity with technologies like LLVM, MLIR, or ROCm is a plus. Heterogeneous System Workloads: Experience running and scaling workloads on large-scale, heterogeneous clusters (CPU + GPU) using distributed training or inference strategies. AI Framework Integration: Experience contributing to or integrating optimizations into deep learning frameworks such as PyTorch, SGLang, vLLM, Slime, VeRL GPGPU Computing: Working knowledge of HIP, CUDA, Triton, TileLang or other GPU programming models; experience with GCN/CDNA architecture preferred. ACADEMIC CREDENTIALS:  Bachelor’s and/or Master’s Degree in Computer Science, Computer Engineering, Electrical Engineering, Physics or a related field. #LI-JG1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Sr. Software Development Engineer  - SGLang and Inference Stack in Santa Clara, CA vacancy
  • $193.3k - $261.5k

     ...builds AWS Neuron, the software development kit used to...  ...enabling unparalleled ML inference and training performance...  .... Working across the stack from PyTorch till the...  ...software boundary, our engineers build systematic...  ...inference serving with vLLM, SGLang, TensorRT or similar... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $152k - $241.5k

     ...application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by...  ...inference engines like vLLM and SGLang-ensuring they run best‑in‑class on...  ...-and by improving the underlying stack that enables high‑throughput, low‑latency... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...skilled and motivated software engineers to join us and build AI inference systems that serve large...  ...-performance inference stacks, optimize GPU kernels and...  ...(e.g., vLLM and SGLang).Familiarity with GPU programming...  ...AI research and development to create groundbreaking... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...career. TTHE ROLE: Senior level engineer who will be responsible for...  ...-training and Distributed Inference Performance on AMD GPU. You...  ...attainment across the entire software stack focusing on getting the best...  ...as PyTorch, JAX, vLLM, and SGLang. Strong technical leadership... 
    Senior

    AMD

    San Jose, CA
    1 day ago
  •  ...at the forefront of software and hardware innovation...  ...System Software Engineer, AI Inference ExecutionWhat you will...  ...helps productize the SW stack for our AI compute engine...  ...responsible for the development, enhancement, and...  ...TensorRT-LLM, vLLM, SGLang, etc.)Experience with... 
    Suggested
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...for a Senior System Software Engineer to work on Dynamo. NVIDIA...  ...Generative AI inference platform to make design...  ...GPUs.Contribute to the development of disaggregated serving...  ...engines (vLLM, SGLang, TRT-LLM) and expand...  ...self-hosted LLM serving stacks with agent harnesses... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment...  ...implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe specifications into... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $140k - $210k

     ...global scale, come make a difference at Fiserv. Job Title Sr. Full Stack Software Engineer About your role: Clover is a global leader in...  ...every day. What you’ll do: Lead the full-stack design, development, and implementation of scalable software applications and... 
    Senior
    Full time
    Work at office
    Monday to Friday
    Flexible hours

    Fiserv

    Sunnyvale, CA
    5 hours ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators....  ...run really fast on the Trainium hardware.As a Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $152k - $241.5k

     ...s TensorRT team as a Senior Software Engineer, and be at the forefront of...  ...enabling high-performance AI inference solutions for automotive safety...  ...doing:Lead the design and development of high-performance deep...  ...across the hardware and software stack to understand and leverage... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $157.5k - $225k

     ...the future of cybersecurity.RoleWe are looking for a Sr. Staff Software Engineer (Full-Stack) to join our team. This is a Hybrid (San Jose, CA) role...  ...grade solutions, with a strong emphasis on spec-driven development workflows using AI coding assistants like Claude or... 
    Senior
    Full time
    Work at office
    Local area
    Flexible hours

    Zscaler

    San Jose, CA
    3 days ago
  •  ...seeking a Principal GenAI Inference Optimization Engineer to join our Models and...  ...models, working across the software-hardware stack.THE PERSONThe ideal...  ...frameworks (e.g., vLLM, SGLang, Triton, or similar systems...  ...decoding.- Support development and optimization of scalable... 

    AMD

    San Jose, CA
    1 day ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What...  ...NVIDIA's LLM inference stack. The first is GPU kernel microbenchmarking...  ...and C++ skills with proven software engineering fundamentals....  ...frameworks such as TRT-LLM, SGLang, or vLLM and clear... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance...  ...explain performance across the full stack, from GPU silicon through the software runtime, and drive competitive...  ...engines including vLLM, SGLang, and emerging serving runtimes.... 
    Senior

    AMD

    Santa Clara, CA
    2 days ago
  •  ...-leading training and inference speeds; over 10 times...  ...a Senior Performance Engineer to join our Product team...  ...expert on how Cerebras stacks up against alternative...  ...stacks (vLLM, SGLang, TensorRT-LLM), GPU kernel...  ...who are serious about software make their own hardware... 
    Senior
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $224k - $356.5k

     ...AI applications like NemoClaw, LLM inference via NIM, Hermes agents, and deep...  ...looking for a deeply technical systems software engineer who will own AI stack readiness on DGX Station. You will...  ...jobs, inference serving alongside development, and resource isolation via MIG or... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $182.5k - $260.5k

     ...platform, its Zero Trust Engine, and the powerful...  ...Scientist, you own the inference and optimization layer...  ...Cutting-edge, unusual stack. The hard, interesting...  ...hands-on in ML/AI (model development, fine-tuning, and...  ...inference runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime... 
    Senior

    Netskope

    Santa Clara, CA
    3 days ago
  •  ...organizations that keep the world running. Job Title: Full-Stack Engineer Location: Onsite, Sunnyvale, California (5 days a week in...  ...have a strong foundation in both front-end and back-end development, with a focus on UI using React, TypeScript, and other modern... 
    Senior
    Full time
    Work at office
    Immediate start

    Illumio

    Sunnyvale, CA
    2 days ago
  • $152k - $241.5k

     ...large language model inference? Join NVIDIA’s TensorRT...  .... We build the software stack that enables Large Language...  ...kernel and operator development for critical...  ...Electrical/Computer Engineering, or a closely related...  ...TensorRT-LLM, vLLM, SGLang, MLC-LLM, or FlashInfer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems...  ...with ML/DL systems development preferableStrong experience...  ...and runtimes such as vLLM, SGLang, and MLC.Strong Python and... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $182k - $226k

     ...fossil fuels. We are seeking a Senior Software Engineer to join Electric Hydrogen's Digital...  ...optimize performance, accelerate product development, and improve operational efficiency. This...  ...customer-facing and internal full-stack web applications using modern web technologies... 
    Senior
    Temporary work

    Electric Hydrogen

    San Jose, CA
    1 day ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...Function of Position:As a Sr Software Engineer Navigation on the...  ...and robot data to CV/ML inference. This role spans the entire development cycle—from prototyping approximate...  ...real-time navigation stack, ensuring accuracy and... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    4 days ago
  • $250k - $300k

     ...each layer of the stack — from electrons...  ...means owning the inference stack end to end:...  ...directly with customer engineering teams to tailor...  ...like vLLM and SGLang to the CUDA...  ...and support the software and product features...  ...~ Professional development & tuition reimbursement... 
    Senior
    Temporary work

    Crusoe

    Sunnyvale, CA
    16 days ago
  • $224k - $356.5k

     ...seeking a deeply technical software manager to lead production AI inference for NVIDIA Inference Microservices...  ...production-ready software stack, combining optimized inference engines, model profiles/recipes,...  ...source ecosystems (i.e vLLM, SGLang, TensorRTLLM, Dynamo, Triton... 

    Socket.dev

    Santa Clara, CA
    3 days ago
  •  ...staffing solutions. Be it core Java, full-stack Java, Web/UI designers, Big Data or...  ...Job Description • Hands-on Design development of end-to-end application features using...  ...and clear code.  • Supports and develops engineers by providing advice, coaching by reviewing... 
    Senior
    Full time

    Jobsbridge

    Santa Clara, CA
    4 hours ago
  • $250.8k - $286.2k

    Senior Lead Software Engineer, Full Stack Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems...  ...regularly worked. San Jose, CA: $250,800 - $286,200 for Sr. Mgr, Software Engineering Candidates hired to work... 
    Senior
    Full time
    Part time
    Internship
    Local area

    Capital One

    San Jose, CA
    5 hours ago
  •  ...at Fiserv.Job TitleSr. Full Stack DeveloperAbout your role:Fiserv...  ...daily. As a Senior Fullstack Engineer on the Loyalty Program team,...  ...will own your services from development through deployment, mentor junior...  ...effective use of AI-powered software engineering and productivity... 
    Senior
    Full time
    Work at office
    Worldwide
    Monday to Friday

    Fiserv

    Sunnyvale, CA
    3 days ago
  •  ...scale, come make a difference at Fiserv.Job TitleSr Full Stack EngineerAbout your role:To accelerate our leadership in the...  ...our digital omni-commerce capabilities. As a Full Stack Software Development Engineering - Advisor I, you will serve as a critical technical... 
    Senior
    Full time
    Work at office
    Local area
    Monday to Friday

    Fiserv

    Sunnyvale, CA
    3 days ago
  • $152k - $241.5k

     ...autonomous cars.We are now looking for a highly motivated Full-Stack Web Applications Engineers to join this dynamic and innovative Hardware...  ...building scalable web services for VLSI designs. This software development team is a multifaceted Agile software team with high production... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $184k - $287.5k

     ...challenges in the field. RL requires inference, rollout generation, and...  ...building an RL Frameworks engineering team to develop the open-...  ...on. The team spans the full software stack, from collaborating closely...  ...performance inference engines (vLLM, SGLang, TensorRT-LLM) into RL... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Software Development Engineer  - SGLang and Inference Stack. Be the first to apply!