Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer, AI Inference Runtime

$262.7k - $355.4k

ARM

As a Principal Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for running SOTA AI Models.You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence how efficiently new models use available compute. Partnering with our AI Infrastructure, compute, and product teams to enhance the performance and efficiency of Arm’s AI platform.Responsibilities:Define the architecture, interfaces, and roadmap for AI inference runtime capabilities, including abstractions that support evolving models, workloads, and compute platforms.Enable new model architectures end to end through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency.Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.Partner with cloud, framework, compiler, hardware, and research teams; lead technical reviews, mentor engineers, and establish meticulous performance-engineering practices.Required Skills and Experience :8+ years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference.Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior.Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, handling of memory resources, and data movement.Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.“Nice To Have” Skills and Experience :Experience developing or modifying inference schedulers, cache managers, batching systems, disaggregated or distributed execution paths.Experience optimizing kernels using accelerator programming tools, assembly, or intrinsics, including attention, matrix multiplication, operator fusion, and low-precision execution.Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization.Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries.In Return:You will be part of our AI Platforms team - A driven and diverse group passionate about developing foundational production capabilities to support AI inference at Arm. We provide a collaborative setting where your ideas can come to life quickly. The success of our AI projects will be directly influenced by your work, crafting the company’s AI inference capabilities and defining and operating production inference workloads. This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!Salary Range:$262,700-$355,400 per yearWe value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.Accommodations at ArmAt Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email View email address on click.appcast.io. To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.Hybrid Working at ArmArm’s approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team’s needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.Equal Opportunities at ArmArm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don’t discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer, AI Inference Runtime in Seattle, WA vacancy
  •  ...in critical industries through AI transformation. We specialize...  ...two systems. The first is our inference control plane — open-weight models...  ...go home — building the fork engine, the guest agent, and the...  ...hoped for. Oversee the sandbox runtime and its host-side control... 
    Suggested
    Full time
    Remote work
    Work visa
    Flexible hours
    Day shift

    Azx Inc

    Seattle, WA
    24 days ago
  • $262.7k - $355.4k

    As a Principal Engineer on Arm’s AI Inference Cloud team, you will shape the technical direction and develop highly...  ...with AI compute, Inference Runtime, and product teams to enhance the performance...  ...lifecycle management.Strong software and production engineering skills,... 
    Suggested
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    1 day ago
  • $143.7k - $194.4k

    AWS Neuron is the complete software stack for the AWS Inferentia...  ...the Software Development Engineer for the Neuron Runtime Team, you will be responsible...  ...learning applications and AI accelerators. You will work...  ...scale distributed training and inference solutions. This... 
    Suggested
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago
  • $168.1k - $227.4k

     ...Neuron is the complete software stack for AWS Inferentia...  .... This senior software engineering role is part of the Machine Learning Inference Applications team and...  ...compiler engineers, and runtime engineers to ensure end...  ...Neuron, TPUs, or other AI accelerator hardware- Experience... 
    Suggested
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $151.8k

    What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role...  ...throughput of ASR service. Profiling and debugging ASR runtime performance bottlenecks across different deployment hardware... 
    Suggested
    Work at office
    Remote work

    Zoom Corporation

    Seattle, WA
    6 days ago
  • $163.9k - $270.88k

     ..., your work matters—and so do you. Principal Software Engineer-ENG: This Principal Software Engineer...  ...secure, scalable, observable, and AI-ready services. You will be responsible...  ...anomaly detection, rule augmentation, model inference pipelines, or decision-support systems... 
    Full time
    Immediate start

    Ukg

    Seattle, WA
    a month ago
  • $165.2k - $223.6k

     ...passionate about building infrastructure that powers AI at scale? Do you want to work on systems that serve millions...  ..., Cost, and Efficiency (ALICE) team is looking for Software Development Engineers to join our Inferences Services Team.We build products and solutions that... 
    Internship
    Local area
    Flexible hours

    Amazon

    Bellevue, WA
    3 days ago
  • $239.7k - $321.4k

     ...global organization of engineers, product developers,...  ...ship, and operate great software without having to...  ...engineering, observability, and AI operations. This role...  ...Summary:As a Senior Principal Software Engineer,...  ...including support for:Inference routing and traffic splittingModel... 

    Disney Interactive

    Seattle, WA
    3 days ago
  • $197.3k - $313.7k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ...person and virtually. As the Principal Engineer focused on architecture priorities...  ...a deep understanding of software development, architecture principles...  ...frameworks such as React, runtimes including Node.js, and CSS... 
    Full time

    Salesforce

    Seattle, WA
    4 hours ago
  •  ..., you’ve come to the right place.As a Principal Software Engineer at JPMorganChase within the Core Foundational...  ...(HPC) practices that intersect with AI/ML. Thus, you are collaborative—...  ...usable patterns to optimize training and inference of ML models on various... 

    JP Morgan Chase

    Seattle, WA
    3 days ago
  • $160k - $250k

     ...Principal Software Engineer Seattle, WA Gradial is the marketing operations system of work that helps...  ..., and help define the future of AI-native content operations, you'll do your...  ...and scaling of LLM agents for real-time inference, dynamic prompting, memory management,... 
    Work at office
    Flexible hours
    3 days per week

    Gradial

    Seattle, WA
    4 days ago
  • $262.7k - $355.4k

    As a Principal Engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure...  ..., fine-tuning, evaluation, and inference. You will guide work across...  ...in developing reliable infrastructure software.Practical knowledge of Kubernetes, containers... 
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    1 day ago
  • $135.2k - $306.4k

     ...production at global scale.Foundational Frameworks: Spearhead the engineering of new container runtimes and distributed frameworks to power OCI’s highest-...  ...from industry innovations to life-saving care. And with AI embedded across our products and services, we help... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    3 days ago
  • $190k - $225k

     ...AssemblyAI builds the best-in-class Voice AI models powering the next generation of...  ...applications. Our models serve 600M+ inference calls monthly, process 1M+ hours of audio...  ...! About the role: We're hiring a Software Engineer to help turn cutting-edge AI research into... 

    AssemblyAI

    Seattle, WA
    8 days ago
  •  ...through innovation and AI-powered automation....  ...developing solutions with engineering excellence, and...  ...globe, our Spectrum-NET software solution performs automated...  ...We are seeking a Principal Software Engineer with...  ...systems, including ML inference  Experience with distributed... 
    Full time
    Relocation
    Flexible hours
    Shift work

    Spectrum Effect

    Bellevue, WA
    7 days ago
  • $188k - $275k

    Staff Software Engineer, Inference CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    2 days ago
  • $262.7k - $355.4k

    As a Principal Software Engineer on the AI Compute Platform team, you will design and build a secure, reliable, and easy-to-use platform for running...  ...systems, and developer tools for distributed training and inference, working closely with compute infrastructure teams, AI... 
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    1 day ago
  • $104.5k - $234.6k

    Responsibilites: Lead engineering and operational management for highly...  ...multiple teams, evolving runtimes or middleware patterns for interoperability...  ...life-saving care. And with AI embedded across our products...  ...Key ResponsibilitiesPlatform Software Development:Lead cross-team... 
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Seattle, WA
    3 days ago
  • $104.5k - $234.6k

    The Software Developer 4 will work on high-impact AI solution designs that enable agents to plan, reason, call tools...  ...Build integrations across agent runtimes, MCP servers, tools gateway, identity...  ...Work with architects and senior engineers to define service contracts,... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    4 hours ago
  •  ...notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate...  ...scalable cloud platforms optimized for AI/ML workloads.Partner with AI teams to...  ...architecture, ML training, and inference.Experience with Infrastructure as Code.... 
    For contractors

    JP Morgan Chase

    Seattle, WA
    4 hours ago
  •  ...ve come to the right place. As a Principal Engineer at JPMorgan Chase on the Core AI Infrastructure Platform team, you...  ...ecosystem that unifies our training and inference pipelines across hybrid-cloud and...  ...training or certification on software engineering concepts and 10+... 

    JP Morgan Chase

    Seattle, WA
    4 days ago
  • $320k

    Staff + Senior Software Engineer, Inference San Francisco, CA | New York City, NY | Seattle, WA About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole... 
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    4 days ago
  •  ...community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the...  ...Kubernetes for inference routing and orchestration. Ensure software solutions are optimized for peak performance during traffic... 
    Full time
    Local area
    Immediate start

    F5 Networks

    Seattle, WA
    4 days ago
  • $172.6k - $259k

     ...games — for players, retailers, and the studios that develop our worlds. We are seeking a Principal Software Engineer who can guide product ideas through every stage with agentic AI — from defining the problem with business partners, to developing the system, to... 
    Full time
    Work experience placement

    Hasbro, Inc.

    Renton, WA
    more than 2 months ago
  •  ...Model Optimization & Deployment Engineer, you will focus on bringing...  ..., and build highly concurrent inference code to ensure real-time, deterministic...  ...maximize memory bandwidth on AI accelerators. Write...  ..., DeepSpeed, Megatron-LM) and runtime efficiency optimization for GPU... 
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    19 days ago
  • $188.7k - $258.39k

     ...enablement, Highspot is committed to building breakthrough software with a spark of magic. We believe a great place to...  .... About the Role We are looking for a Principal Software Engineer to join our Applied AI team. This is a rare opportunity to work at the... 
    Full time
    Immediate start
    Flexible hours

    Highspot

    Seattle, WA
    more than 2 months ago
  • $193.3k - $261.5k

     ...closely with the hardware and software teams to ensure the right...  ...provides ability for performance engineers to develop and improve...  ...other teams including training, inference and runtime.* Collaborate with the...  ...growing suite of generative AI services and other cloud computing... 
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $143.7k - $194.4k

    We are looking for a Software Development Engineer II (SDE-2) to join the EKS Node Runtime team. In this role, you will design, build...  ...work will focus on enabling AI and ML workloads by implementing...  ...to cutting edge AI training and inference.Mentorship: Mentor junior engineers... 
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago
  • $168.1k - $227.4k

    Shape the Future of AI Accelerators at AWS NeuronWe...  ...Amazon Neuron, the software development kit used to...  ...Trainium.As a Senior Software Engineer on our Machine Learning...  ...building distributed inference support for Pytorch in...  ...across compiler, runtime, framework, and hardware... 
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  •  ...the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by...  ...(Snowgrid), data sharing, and data marketplace. AS A PRINCIPAL SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large... 
    Full time

    Snowflake

    Bellevue, WA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer, AI Inference Runtime. Be the first to apply!