Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Software Development Engineer  - SGLang and Inference Stack (Santa Clara)

Part-time

AMD

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:    As a core member of the team, you will play a pivotal role in optimizing and developing deep learning frameworks for AMD GPUs. Your work will be instrumental in enhancing GPU kernel performance, accelerating deep learning models, and enabling RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node systems. You will collaborate across internal GPU software teams and engage with open-source communities to integrate and optimize cutting-edge compiler technologies and drive upstream contributions that benefit AMD’s AI software ecosystem. THE PERSON:   Skilled engineer with strong technical and analytical expertise in GPGPU C++, Triton, TileLang or DSL development within Linux environments. The ideal candidate will thrive in both collaborative team settings and independent work, with the ability to define goals, manage development efforts, and deliver high-quality solutions. Strong problem-solving skills, a proactive approach, and a keen understanding of software engineering best practices are essential. KEY RESPONSIBILITIES:  Optimize Deep Learning Frameworks: Enhance performance of frameworks like TensorFlow, PyTorch, and SGLang on AMD GPUs via upstream contributions in open-source repositories. Develop and Optimize Deep Learning Models: Profile, analyze, code change and tune large-scale training and inference models for optimal performance on AMD hardware. Day-0 supports to many SOTA models, DeepSeek 3.2, Kimi K2.5, etc.GPU Kernel Development: Design, implement, and optimize high-performance GPU kernels using HIP, Triton, TileLang or other DSLs for AI operator efficiency. Collaborate with GPU Library and Compiler Teams: Work closely with internal compiler and GPU math library teams to integrate, optimize and align kernel-level optimizations with full-stack performance goals. Initiate and help with different level codegen optimizations.Contribute to SGLang Development: Support optimization, feature development, and scaling of the SGLang framework across AMD GPU platforms for LLM, multimodal serving and RL-training. Distributed System Optimization: Tune and scale performance across both multi-GPU (scale-up) and multi-node (scale-out) environments, including inference parallelism, prefill-decode disaggregation, Wide-EP and collective communication strategies. Graph Compiler Integration: Integrate and optimize runtime execution through graph compilers such as XLA, TorchDynamo, or custom pipelines. Open-Source Collaboration: Partner with external maintainers to understand framework needs, propose optimizations, and upstream contributions effectively. Apply Engineering Best Practices: Leverage modern software engineering practices in debugging, profiling, test-driven development, and CI/CD integration. PREFERRED EXPERIENCE:  Strong Programming Skills: Proficient in C++ and/or Python (PyTorch, Triton, TileLang), with demonstrated ability to code, debug, profile, and optimize performance-critical code. SGLang and LLM Optimization: Hands-on experience with SGLang or similar LLM inference frameworks is highly preferred. Compiler and GPU Architecture Knowledge: Background in compiler design or familiarity with technologies like LLVM, MLIR, or ROCm is a plus. Heterogeneous System Workloads: Experience running and scaling workloads on large-scale, heterogeneous clusters (CPU + GPU) using distributed training or inference strategies. AI Framework Integration: Experience contributing to or integrating optimizations into deep learning frameworks such as PyTorch, SGLang, vLLM, Slime, VeRL GPGPU Computing: Working knowledge of HIP, CUDA, Triton, TileLang or other GPU programming models; experience with GCN/CDNA architecture preferred. ACADEMIC CREDENTIALS:  Bachelor’s and/or Master’s Degree in Computer Science, Computer Engineering, Electrical Engineering, Physics or a related field. #LI-JG1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Sr. Software Development Engineer  - SGLang and Inference Stack (Santa Clara) in Santa Clara, CA vacancy
  • $184k - $287.5k

     ...skilled and motivated software engineers to join us and build AI inference systems that serve...  ...inference stacks, optimize GPU kernels...  ...engines (e.g., vLLM and SGLang).Familiarity with...  ...AI research and development to create...  ...SummaryLocation: US, CA, Santa ClaraType: Full time... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $182.5k - $260.5k

     ...across offices in Santa Clara, St. Louis,...  ...Scientist, you own the inference and optimization...  ...Cutting-edge, unusual stack. The hard,...  ...systems and backend engineers to ship...  ...in ML/AI (model development, fine-tuning, and...  ...inference runtimes (vLLM/SGLang, TensorRT-LLM, ONNX... 
    Senior
    Part time
    Work at office

    Netskope

    Santa Clara, CA
    5 hours ago
  •  ...forefront of software and hardware innovation...  ...with LLM inference on...  ...spans the full stack: from pathfinding...  ...research and engineering team that moves...  ...frameworks (vLLM, SGLang, TensorRT-LLM,...  ...and business development to translate POCs...  ...benefits in Santa Clara, CA.Equal Opportunity... 
    Suggested
    Part time

    d-Matrix

    Santa Clara, CA
    5 hours ago
  • $224k - $356.5k

     ...applications like NemoClaw, LLM inference via NIM, Hermes agents,...  ...technical systems software engineer who will own AI stack readiness on DGX Station...  ...serving alongside development, and resource isolation...  ...SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time... 
    Senior
    Full time
    Part time
    Local area

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...the field. RL requires inference, rollout generation,...  ...an RL Frameworks engineering team to develop the open...  ...team spans the full software stack, from collaborating closely...  ...engines (vLLM, SGLang, TensorRT-LLM) into RL...  ...SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType:... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...outstanding AI systems engineers to develop...  ...technologies in the inference systems software stack! We build innovative...  ...with ML/DL systems development preferableStrong experience...  ...such as vLLM, SGLang, and MLC.Strong Python...  ...: US, CA, Santa Clara; US, GA, Remote; US... 
    Senior
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...Deep Learning engineer to bring advanced...  ...into AI stacks, including PyTorch...  ...-LLM, vLLM, SGLang, JAX, etc. You...  ...100K GPUs to inference down at microsecond...  ...to the development of innovative...  ...(aka systems software fundamentals)Adaptability...  ...: US, CA, Santa Clara; US, TX,... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers...  ...all layers of the hardware/software stack from GPU architecture to...  ...and multimodal model inference as part of NVIDIA Inference...  ...law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full time... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ...We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance...  ...explain performance across the full stack, from GPU silicon through the software runtime, and drive competitive...  ...engines including vLLM, SGLang, and emerging serving runtimes.... 
    Senior
    Part time

    AMD

    Santa Clara, CA
    5 hours ago
  • $224k - $356.5k

     ...progress.We’re looking for a Senior Full-Stack Software Engineer to join the AI Hub team within the...  ....Experience leveraging AI-assisted development tools (e.g., Cursor).Ways to Stand Out...  ...protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType: Full time... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service...  ...focus on the cloud-native stack for datacenter products like...  ...years of professional software development experience in distributed...  ....SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, Remote... 
    Senior
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $195.2k - $361.2k

     ...actually own. You optimize inference engines (llama.cpp, vLLM) for constrained...  ...related STEM field8+ years software development backgroundStrong in C++ and...  ...Location: US, California, Santa ClaraAdditional Locations:...  ...: US, California, Santa Clara; US, Oregon, Hillsboro; US,... 
    Senior
    Full time
    Part time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    5 hours ago
  • $224k - $356.5k

     ...Networking Systems & Software Architecture...  ...into production AI stacks. The team's...  ...research and production engineering!What you will be...  ...such as vLLM, SGLang, and TensorRT-LLMContributing...  ...training and inference patterns.Ways to...  ...: US, CA, Santa Clara; US, TX, Austin;... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $272k - $431.25k

     ...looking for a Principal Engineer to join our CSP...  ...configuration, software, or workload differences...  ...the full software stack impacts...  ...environmentsUnderstanding of inference workload...  ...vLLM, TensorRT-LLM, SGLang, continuous batching...  ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin;... 
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    10 hours ago
  • $184k - $287.5k

     ...built. We are seeking a Senior Software Engineer focused on container and...  ...strategy for NVIDIA Inference Microservices (NIMs) and our...  ...inference backends (vLLM, SGLang, TRT-LLM)Background in benchmarking...  ....SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $190k - $300k

     ...are at the forefront of software and hardware...  ...possibilities of AI. Location:Santa Clara, CA, headquarters, but...  ...role: Principal Software Engineer, SDK & Lowering...  ...compiler teams.Create development tools to be used in kernel...  ...integral to the software stack.Take ownership of... 
    Part time
    Work experience placement

    d-Matrix

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on NVIDIA CPU platforms...  ...experience.Strong C and C++ development, debugging, and code-review skills...  ...law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX,... 
    Senior
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ...learning models, and training/inference performance across multi-GPU...  ...technologies and advanced engineering principles to drive continuous...  ...goals, scope, and own development efforts while collaborating...  ...integrating graph compilers. Software Engineering Best Practices:... 
    Senior
    Part time

    AMD

    Santa Clara, CA
    5 hours ago
  •  ...the forefront of software and hardware innovation...  ...of AI. Location:Santa Clara, CA headquarters,...  ...StaffSoftware Engineer, Developer and...  ...cutting edge AI inference accelerators. You...  ...for the design, development, enhancement, and...  ...hardware and software stack. Minimum... 
    Senior
    Part time

    d-Matrix

    Santa Clara, CA
    5 hours ago
  •  ...OPPORTUNITYWe're looking for a senior software engineer who combines deep systems...  ...distributed training and inference.You'll join a core team of...  ...efficient on GPUs—tuning stacks, kernels, and workflows in...  ...where it accelerates development and tuning of the ROCm ecosystem... 
    Senior
    Part time
    Shift work

    AMD

    Santa Clara, CA
    5 hours ago
  • $224k - $356.5k

     ...silicon to the full-stack AI systems that...  ...skilled Deep Learning Engineer to develop systems...  ...researchers, software engineers to bring...  ...transformer architectures, inference bottlenecks, and...  ...using vLLM, SGLang.Hands-on experience...  ...SummaryLocation: US, CA, Santa Clara; US, WA,... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $200k - $322k

     ...Technical Marketing Engineer focused on Enterprise AI Software, and accelerating...  ...AI software stack and the developers...  ...agentic AI blueprints, inference platforms, Kubernetes...  ...across model development, inference, RAG, agentic...  ...: US, CA, Santa Clara; US, RemoteType: Full... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...highly motivated software professional to work...  ..., custom kernel development, and cluster-...  ...both training and inference pipelines.Collaborate...  ..., Computer Engineering, Electrical Engineering...  ...TensorRT, vLLM, sgLang, Nemo, Megatron,...  ...: US, CA, Santa Clara; US, TX, Austin;... 
    Senior
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ...a Forward Deployed Research Engineer to build, evaluate, and deploy...  ...role. You'll combine deep software engineering, applied AI research...  ...EvaluationAI Training or Inference InfrastructureDeep...  ...systems experienceLOCATION:Santa Clara, CA#LI-BW1#LI-hybridBenefits... 
    Senior
    Part time

    AMD

    Santa Clara, CA
    5 hours ago
  • $224k - $356.5k

     ...performance across various inference frameworks....  ...Manager, you will lead the engineering team within NVIDIA’s Dynamo...  ...vLLM, TRT-LLM, and SGLang in partnership with...  ...experience.8+ overall years of software engineering experience...  ...: US, CA, Santa Clara; US, TX, Austin; US, AL... 
    Full time
    Part time
    Local area
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...edge hardware and software innovation to...  ...team of innovative engineers dedicated to solving...  ...across the stack - from GPU operator...  ...plugins to distributed inference serving and cloud...  ...shape design and development decisions.What we...  ...: US, CA, Santa Clara; US, WA, SeattleType... 
    Senior
    Full time
    Part time
    Worldwide

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...revolution, building the software and systems that power...  ...for a Senior Software Engineer to lead the bring-up,...  ...training and inference workloads across NVIDIA...  ...and inference/training stacks to ensure state-of-the...  ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR... 
    Senior
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ...the forefront of software and hardware...  ...working onsite at our Santa Clara, CA,...  ...PrincipalSystem Software Engineer - AI Inference ExecutionWhat you...  ...the SW stack for our AI compute...  ...responsible for the development, enhancement, and...  ...TensorRT-LLM, vLLM, SGLang, etc.)Experience... 
    Part time
    3 days per week

    d-Matrix

    Santa Clara, CA
    5 hours ago
  • $224k - $356.5k

     ...team is building the software stack that makes large...  ...leading open-source LLM inference frameworks —...  ...assessment, inference recipe development, performance...  ...Computer Science, Computer Engineering, Electrical...  ...SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US... 
    Senior
    Full time
    Part time
    Local area

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...enablement of Network Industry Software Vendors (ISVs). These ISVs are developing a network stack for distributed inference which will be used to...  ...Computer Science, Electrical Engineering, Software Engineer, or...  ...law.SummaryLocation: US, CA, Santa ClaraType: Full time... 
    Senior
    Full time
    Part time
    Work experience placement

    Nvidia

    Santa Clara, CA
    5 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Software Development Engineer  - SGLang and Inference Stack (Santa Clara). Be the first to apply!