Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Accelerator Software Principal Engineer - Runtime Library

$182k - $273k

Jobleads-US

AI Accelerator Software Principal Engineer - Runtime Library

Ampere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.

As a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.

Join us at Ampere and work alongside a passionate and growing team.

About the Role

As an AI Accelerator Principal Software Engineer - Runtime Library, you will lead the design, development, and optimization of AI runtime software that enables multiple state-of-the-art deep learning models to run efficiently on Ampere’s deep learning accelerators. You will work at the intersection of systems software, performance engineering, and AI enablement, helping deliver high-throughput, low-latency inference and a strong foundation for future model and framework support.

What You’ll Achieve

  • Build and evolve an AI Runtime Library for Ampere accelerators that supports execution, scheduling, and lifecycle management of deep learning workloads across multiple model types and popular frameworks.
  • Own end-to-end acceleration paths , going deep into the full SW/HW stack—including:
    • Inference serving and integration layers
    • Compiler/runtime interfaces and graph/IR execution flows
    • Runtime library architecture (APIs, memory management, operators, execution engines)
    • Communication mechanisms and device/host orchestration
  • Drive HW/SW co-design and optimization to improve:
    • Throughput (tokens/requests per second)
    • Latency (kernel execution and scheduling efficiency)
    • Memory efficiency (buffering, paging, reuse, caching)
    • Overall compute utilization and scaling behavior
  • Contribute to AI co-processor/accelerator software enablement , partnering closely with hardware and systems teams to ensure runtime and kernel strategies match accelerator capabilities and constraints.
  • Collaborate cross-functionally to integrate runtime components into Ampere platform stacks, ensuring robust deployment on target environments and consistent performance in production-like workloads.

About You

  • BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or related technical field & 8 years of related experience; or MS degree & 6 years; or PhD & 3 years
  • Proven experience developing user-mode drivers and/or runtime libraries for GPUs or deep learning accelerators in Linux or RTOS environments.
  • Strong expertise in C/C++ and systems-level programming (memory, threading, synchronization, performance profiling).
  • Demonstrated background in AI framework enablement, with hands‑on experience in one or more of:
    • PyTorch (operator/runtime integration, graph execution, correctness/performance work)
    • llama.cpp (inference/runtime execution patterns)
    • ONNX (graph handling, interoperability, execution engines)
  • Strong performance engineering skills, including profiling/diagnostics and optimization of execution pipelines, data movement, and compute kernels.
  • Ability to operate effectively in a collaborative environment—owning complex components while partnering with compilers, hardware, and platform teams.

What We’ll Offer

At Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long-term incentive, and comprehensive benefits. The full base pay range for this role is between $182,000 and $273,000, except in the San Francisco Bay Area where the range is between $195,000 and $292,000.

Our benefits include health, wellness, and financial programs that support employees through every stage of life.

Benefit highlights include:

  • Premium medical insurance, dental insurance, vision insurance, as well as income protection and a 401K retirement plan, so that you can feel secure in your health and financial future.
  • Unlimited Flextime and 10+ paid holidays so that you can embrace a healthy work-life balance.
  • A variety of healthy snacks, energizing espresso, and refreshing drinks to keep you fueled and focused throughout the day.

And there is much more than compensation and benefits. At Ampere, we foster an inclusive culture that empowers our employees to do more and grow more. We are excited to share more about our career opportunities with you through the interview process. Our benefits include health, wellness, and financial programs that support employees through every stage of life.

Ampere is an inclusive and equal opportunity employer and welcomes applicants from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, national origin, citizenship, religion, age, veteran and/or military status, sex, sexual orientation, gender, gender identity, gender expression, physical or mental disability, or any other basis protected by federal, state or local law.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Accelerator Software Principal Engineer - Runtime Library in Santa Clara, CA vacancy
  •  ...Ampere Computing LLC is seeking an AI Accelerator Principal Software Engineer - Runtime Library to lead design, development, and optimization of AI runtime software for Ampere accelerators, enabling efficient inference across models and frameworks. You will own end... 
    Suggested

    Jobleads-US

    Santa Clara, CA
    1 day ago
  • $165.2k - $223.6k

     ...Neuron is the complete software stack for the AWS...  ...scale machine learning accelerators and the servers that...  ...Development Engineer for the Neuron Runtime Team, you will be responsible...  ...performance runtime libraries and drivers for...  ...learning applications and AI accelerators. You... 
    Suggested
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $150k - $180k

    Santa Clara, CASoftware Engineering - Runtime Platform /Full-time /HybridPlusAI is a Physical AI company pioneering AI-based virtual driver software for factory-built autonomous trucks. Headquartered...  ...and DSV are working with Plus to accelerate the deployment of next-generation... 
    Suggested
    Full time

    Plus.ai

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA’s accelerated computing platform is foundational...  ...to modern HPC and AI. At the center of...  ...are CUDA Core Libraries that enable...  ...scalable GPU-accelerated software. We are hiring a Senior Software Engineer to develop the...  ...algorithms, and language/runtime infrastructure... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA’s accelerated computing platform is foundational...  ...to modern HPC and AI. At the center of...  ...are CUDA Core Libraries that enable...  ...scalable GPU-accelerated software.We are hiring a Senior Software Engineer to advance the...  ..., algorithms, and runtime infrastructure on... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $120.8k - $193.3k

    DescriptionJob Title: Principal Engineer, Apps EngineeringJob...  ...Job Description: SiMa.ai is seeking a...  ...of complex hardware/software integration issues and...  ...TensorFlow, ONNX, ONNX Runtime, and OpenCV.Experience...  ...platforms, SoCs, or AI accelerators.Strong understanding... 
    Full time
    Work at office

    SiMa Technologies

    San Jose, CA
    3 days ago
  • $193.3k - $261.5k

     ...models of these custom-designed accelerator SoCs for use by AWS internal...  ...for a Senior SoC Modeling Engineer to join the team and deliver...  ...verification, emulation, and software teams to build, debug, and deploy...  ...growing suite of generative AI services and other cutting-... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $184k - $287.5k

     ...graphics, PC gaming, and accelerated computing for more...  ...unlimited potential of AI to define the next era...  ...cutting‑edge hardware and software innovation to deliver...  ...of forward‑thinking engineers tackling some of the...  ...Kubernetes‑based accelerated runtime stack (control and... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    13 hours ago
  • $176k - $298k

     ...innovating memory and storage solutions that accelerate the transformation of information into...  .... What's Encouraged Daily: AI-First Engineering: Deploy AI tools (LLMs, code/RTL...  ...offs—to be AI-consumable. Build prompt libraries, agents, and scripts that eliminate recurring... 
    Full time
    Local area
    Immediate start

    SwiftCruit

    San Jose, CA
    3 days ago
  • $130k - $220k

    Santa Clara, CASoftware Engineering - Motion Planning /Full...  ...is a Physical AI company pioneering AI-based virtual driver software for factory-built autonomous...  ...are working with Plus to accelerate the deployment of next-...  ...optimize the production runtime pipeline for ML-based autonomous... 
    Full time

    Plus.ai

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...are the GPU Communications Libraries and Networking team at NVIDIA...  ...HPC. We're seeking a Senior Software Architect to help co-design...  ...communication technologies to accelerate AI and HPC workloads.Explore...  ...at least one communication runtime (MPI, NCCL, NVSHMEM, OpenSHMEM... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...reliable, scalable experiences. We are looking for engineers who pair strong software fundamentals with practical AI-assisted development habits to move faster,...  ...account journeys Use AI-assisted workflows to accelerate implementation, code review, testing, debugging,... 
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $175.8k - $293k

     ...Information Job Name Principal Innovator - USA (B)...  ...Forbes Global 100 to accelerate business value,...  ...looking for a Principal AI Engineer to architect, build,...  ..., orchestration runtimes, and inference serving...  ...scalable, production-grade software, with significant... 

    Jobleads-US

    Santa Clara, CA
    13 hours ago
  • $224k - $356.5k

    We are looking for a Software Engineering Manager to lead a team responsible...  ...readiness, enabling CUDA Math Libraries on new and specialized...  ...organizations are revolutionizing AI, data analytics, and...  ...product engineering teams in accelerated computing domains. If this sounds... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering...  .... It combines software and systems...  ...production environments.As a Principal SRE, you will shape...  ...of NVIDIA’s AI Platform Runtime and lead reliability...  ...intelligent automation that accelerate platform operations,... 
    Full time

    Nvidia

    Santa Clara, CA
    13 hours ago
  • $184k - $287.5k

    NVIDIA’s accelerated computing platform is foundational...  ...to modern HPC and AI. At the center of...  ...are CUDA Core Libraries that provide the...  ...abstractions, and runtime capabilities needed...  ...GPU-accelerated software.We are hiring a Senior Software Engineer to advance the C++... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...NVIDIA IT’s Enterprise AI & Automation team...  ...efficiency and accelerate business results across engineering, IT, supply chain,...  ...and sales. We need a Principal or Distinguished Engineer...  ...memory, controlled runtime environments,...  ...maintainership, widely adopted libraries, or public... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...the world's largest AI chip, 56 times...  ...Cerebras Wafer-Scale Engine. Our work spans...  ...frameworks, compilers, runtimes, and low-level...  ...capabilities.Contribute to software architecture and...  ...-level assembly, accelerator programming, or a...  ..., or kernel libraries.Why Join CerebrasPeople... 

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $224k - $356.5k

     ...and help build and improve GPU and CPU accelerated software libraries. Projects like nvComp, NPP, nvJPEG,...  ...vision and growth. Starting from powering AI, data analytics, image processing,...  ...leadership and guidance to library engineers working with you,find opportunities to... 
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $140k - $215k

     ...the world’s most advanced AI-native platform. We work on...  ...and continuously accelerate execution, build expertise...  ...the Role:This is a systems software role on the Core Libraries team, which owns the shared...  ...sensor platform.As a Senior Engineer on this team, you will design... 
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    4 days ago
  • $200k

     ...Overview We are looking for a Principal AI SoC Runtime Software Architect to own the software...  ...across the SoC’s CPU cores, AI accelerator, vision and multimedia engines, and other embedded processors...  ...-facing APIs, user-space libraries, and low-level driver, firmware... 
    Full time
    Contract work
    Flexible hours

    Velaura AI

    Santa Clara, CA
    11 days ago
  • $183.6k - $297k

     ...Integrity, and Inclusion. We weave AI into the fabric of everything...  ...to alerts.Our users—software developers—need to troubleshoot...  .... We are seeking experienced engineers who can craft intuitive, fast...  ...paced environment that fosters accelerated learning, continuous growth,... 
    Full time
    Remote work
    Flexible hours

    Palo Alto Networks

    San Jose, CA
    1 day ago
  • $184k - $287.5k

     ...motivated Deep Learning engineer to bring advanced...  ...and Distributed Runtime technologies into AI stacks, including PyTorch...  ...in AI toolkits will accelerate enabling those for...  ...principles (aka systems software fundamentals)...  ...specific communication libraries (e.g., NCCL, MPI, UCX... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...graphics, PC gaming, and accelerated computing for more than...  ...unlimited potential of AI to define the next era of...  ...looking for an experienced Software Engineer to develop our core libraries for Agentic Applications...  ...level API calls through runtime internals, language... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...graphics, PC gaming, and accelerated computing for more...  ...potential of AI to define the next...  ...a hands-on senior engineer to expand Warp's...  ...integrate Warp into their software, resolve technical...  ...integrate CUDA-X libraries and other...  ...models, compilers, runtimes, and kernel libraries... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $165.2k - $223.6k

     ...builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI...  ...includes an ML compiler, runtime, collectives library, and application...  ...are looking for software engineers to help build and fine...  ...shape the direction of AI acceleration technology... 
    Local area
    Work from home
    Flexible hours
    Shift work

    Amazon

    Cupertino, CA
    1 day ago
  • $193.3k - $261.5k

    Annapurna Labs designs silicon and software that accelerates innovation. Our custom...  ...seeking a Senior Software Engineer to join our ML Distributed...  ..., compiler engineers, runtime engineers and AWS solution...  ...torchtitan and Hugging Face libraries for the Neuron ecosystem.A... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ...build great products that accelerate next-generation computing experiences—from AI and data centers, to...  ...strategically critical to AMD’s AI software roadmap.AMD GPUs are an...  ..., and performance engineering. You are comfortable...  ...deployment, compiler, runtime, and hardware teams.... 

    AMD

    San Jose, CA
    2 days ago
  • Job DescriptionThe AV Runtime team within GM’s Autonomous...  ...organization builds software and data systems that...  ...vehicle data into the AI development pipeline....  ...investigation, and engineering automation.You will work...  ...language models (LLMs) to accelerate issue discovery,... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 hours ago
  • $145k - $190k

     ...sitePlusAI is a Physical AI company pioneering AI-based virtual driver software for factory-built...  ...are working with Plus to accelerate the deployment of next-...  ...growing teams.As software engineer for mapping &...  ...efficient map query in runtime for other modules- like... 
    Full time

    Plus.ai

    Santa Clara, CA
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Accelerator Software Principal Engineer - Runtime Library. Be the first to apply!