Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer, AI Inference Runtime

$209.1k - $282.9k

ARM

As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for running SOTA AI Models.You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence how efficiently new models use available compute. Partnering with our AI Infrastructure, compute, and product teams to enhance the performance and efficiency of Arm’s AI platform.Responsibilities:Define the architecture, interfaces, and roadmap for AI inference runtime capabilities, including abstractions that support evolving models, workloads, and compute platforms.Enable new model architectures end to end through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency.Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.Partner with cloud, framework, compiler, hardware, and research teams; lead technical reviews, mentor engineers, and establish meticulous performance-engineering practices.Required Skills and Experience :5+ years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference.Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior.Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, handling of memory resources, and data movement.Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.“Nice To Have” Skills and Experience :Experience developing or modifying inference schedulers, cache managers, batching systems, disaggregated or distributed execution paths.Experience optimizing kernels using accelerator programming tools, assembly, or intrinsics, including attention, matrix multiplication, operator fusion, and low-precision execution.Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization.Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries.In Return:You will be part of our AI Platforms team - A driven and diverse group passionate about developing foundational production capabilities to support AI inference at Arm. We provide a collaborative setting where your ideas can come to life quickly. The success of our AI projects will be directly influenced by your work, crafting the company’s AI inference capabilities and defining and operating production inference workloads. This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!Salary Range:$209,100-$282,900 per yearWe value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.Accommodations at ArmAt Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email View email address on click.appcast.io. To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.Hybrid Working at ArmArm’s approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team’s needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.Equal Opportunities at ArmArm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don’t discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer, AI Inference Runtime in Seattle, WA vacancy
  • $236k - $330k

     ...this new era, we seek AI-native thinkers across...  ...state of the art in LLM inference systems and...  ...distributed serving and runtime systems to GPU kernels...  ...We embrace AI-native engineering, using AI not only as...  ...approaches to accelerate software development, experimentation... 
    Suggested
    Shift work

    Snowflake

    Bellevue, WA
    2 days ago
  • $209.1k - $282.9k

    As an engineer on Arm’s AI Inference Cloud team, you will shape the technical direction and develop highly...  ...with AI compute, Inference Runtime, and product teams to enhance the performance...  ...workload lifecycle management.Strong software and production engineering skills, including... 
    Suggested
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    3 days ago
  • $143.7k - $194.4k

    AWS Neuron is the complete software stack for the AWS Inferentia...  ...the Software Development Engineer for the Neuron Runtime Team, you will be responsible...  ...learning applications and AI accelerators. You will work...  ...scale distributed training and inference solutions. This... 
    Suggested
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $168.1k - $227.4k

    AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud...  ...This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is...  ...architects, compiler engineers and runtime engineers to deliver performance and... 
    Suggested
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $168.1k - $227.4k

     ...Neuron is the complete software stack for AWS...  ...This senior software engineering role is part of...  ...Machine Learning Inference Applications team...  ...compiler engineers, and runtime engineers to...  ...Neuron, TPUs, or other AI accelerator...  ...supervisors, and staff; adhere to standards... 
    Suggested
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $209.1k - $282.9k

    We seek a Software Engineer to contribute to the next generation of Physical AI platforms on Arm. Your work will focus on building...  ...services, developer tooling, and runtime infrastructure that boost...  ...time systems.Exposure to AI/ML inference systems, even from a platform... 
    Work at office
    Local area

    ARM

    Seattle, WA
    3 days ago
  • $173.5k - $331.05k

     ...develops industry‑leading software products including...  ...mission to build the modern, AI‑powered video rendering...  ...for a senior, hands‑on engineer to own and evolve the cross...  .../ hardware vendors ML inference integration (e.g., TensorRT, ONNX Runtime) within a render... 
    Temporary work

    Adobe

    Seattle, WA
    1 day ago
  • $229k - $343k

     ...digital services.We’re looking for a Staff Software Engineer to join Snap Inc on our Feature Store...  ...features reliably for large-scale batch inference and low-latency online inference, maintaining...  ...with responsible use of emerging AI technologies to improve engineering... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Seattle, WA
    3 days ago
  • $254k - $350k

     ..., WA / Remote (United States)Software – Software Systems /Full-time...  ...alongside a team of strong software engineers and act as a force multiplier...  ...cutting-edge ML Training OR Inference performance optimization...  ...use artificial intelligence (AI) tools to support parts of... 
    Full time
    Remote work

    Zoox

    Seattle, WA
    1 day ago
  • $201k - $315k

     ...MA / Seattle, WASoftware - Software Systems /Full-time /HybridZoox...  ...looking for an experienced Staff Software Engineer to build, scale, and operate...  ...engineering to training our AI models in Perception, Planner...  ...workloads (training, inference, data generation)Experience... 
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    3 days ago
  •  ...Description:DataRobot delivers AI that maximizes impact and...  ...DataRobot’s Fleet team is the engine behind how our platform runs...  ...That’s where you come in.As a Staff Software Engineer, you’ll be responsible...  ...for training and inference.Why Join the Fleet Management... 
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    DataRobot

    Seattle, WA
    3 days ago
  • $209.1k - $282.9k

    As an engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure...  ..., fine-tuning, evaluation, and inference. You will work across Kubernetes...  ...in developing reliable infrastructure software.Practical knowledge of Kubernetes, containers... 
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    3 days ago
  • $188k - $275k

     ...Staff Software Engineer, InferenceCoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams...  ...traded company (Nasdaq: CRWV) in March 2025.Inference Platform Team The Inference team builds and... 
    Permanent employment
    Full time
    Casual work
    Work at office

    CoreWeave

    Bellevue, WA
    2 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving...  ...the future. We are seeking a Staff Engineer to help our development of...  ...of AI training and inference at scale. As a Staff Engineer...  ...~10+ years of experience in software engineering, platform engineering... 
    Work at office
    Local area
    Immediate start
    Work from home
    Flexible hours

    Socket

    Bellevue, WA
    2 days ago
  • $254k - $350k

     ...Seattle, WASoftware - Autonomy Software /Full-time /HybridThe...  ...Learning and System Optimization Engineer, you will orchestrate and allocate...  ...allow for more efficient inference by sharing various parts of the...  ...in low-level programming for AI accelerators, specifically... 
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    4 days ago
  •  ...notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate...  ...scalable cloud platforms optimized for AI/ML workloads.Partner with AI teams to...  ...architecture, ML training, and inference.Experience with Infrastructure as Code.... 
    For contractors

    JP Morgan Chase

    Seattle, WA
    2 days ago
  • $209.1k - $282.9k

    As a software Engineer on the AI Compute Platform team, you will design and build a secure, reliable, and easy-to-use platform for running AI workloads...  ...systems, and developer tools for distributed training and inference, working closely with compute infrastructure teams, AI... 
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    3 days ago
  •  ...services that enable ML engineers and data scientists...  ...monitoring, and agentic AI capabilities. We work closely...  ...-driven products. As a Staff Engineer, you'll make...  ..., low-latency model inference, large-scale feature stores...  ..., and operating great software systems.Who you areWe'... 
    Flexible hours

    Stripe

    Seattle, WA
    3 days ago
  •  ...community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the...  ...Kubernetes for inference routing and orchestration. Ensure software solutions are optimized for peak performance during traffic... 
    Full time
    Local area
    Immediate start

    F5 Networks

    Seattle, WA
    2 days ago
  • $182.4k - $247k

     ...and running the world's best data and AI infrastructure platform so our...  ...to improve their business. Founded by engineers — and customer obsessed — we leap at...  ...traditional SQL query engines. As a software engineer on the Runtime team at Databricks, you will be building... 
    Local area
    Worldwide

    DataBricks

    Bellevue, WA
    1 day ago
  •  ...reporting, and intelligent product workflows.We’re hiring a Staff Software Engineer to build the AI platforms within Data Cloud. At Rippling, you aren’t...  ...team to define shared primitives for model tuning, inference, evaluation, and development of task specific agent harnesses... 
    Work at office
    3 days per week

    Rippling

    Seattle, WA
    3 days ago
  • $168.1k - $227.4k

    Shape the Future of AI Accelerators at AWS NeuronWe...  ...Amazon Neuron, the software development kit used to...  ...Trainium.As a Senior Software Engineer on our Machine Learning...  ...building distributed inference support for Pytorch in...  ...across compiler, runtime, framework, and hardware... 
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    22 hours ago
  • $143.7k - $194.4k

    We are looking for a Software Development Engineer II (SDE-2) to join the EKS Node Runtime team. In this role, you will design, build...  ...work will focus on enabling AI and ML workloads by implementing...  ...to cutting edge AI training and inference.Mentorship: Mentor junior engineers... 
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $193.3k - $261.5k

     ...with the hardware and software teams to ensure the right...  ...for performance engineers to develop and improve...  ...teams including training, inference and runtime.* Collaborate with the...  ...suite of generative AI services and other cloud...  ..., supervisors, and staff; adhere to standards of... 
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  •  ...management (CLM).What you'll doAs a Software Engineer on the AI Platform team, you will architect and...  ...systems to support large-scale model inference and data processingDesign and build resilient...  ...for model serving and inference runtimes, focusing on maximizing resource... 
    Contract work
    Work at office
    Local area
    Remote work
    2 days per week

    DocuSign

    Seattle, WA
    1 day ago
  • $143.7k - $194.4k

     ...builds AWS Neuron, the software development kit used to...  ...includes an ML compiler, runtime, and application...  ...enabling unparalleled ML inference and training performance...  ...software boundary, our engineers build systematic infrastructure...  ...of what's possible in AI acceleration.As part of... 
    Work experience placement
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $143.7k - $194.4k

    Build the AI system that decides how Amazon plans...  ...metrics- Tight engineering team (8 SDEs), high autonomy...  ...— including agent runtime, caching/cost-latency...  ...augmented generation, causal inference, and automated...  ...internship professional software development experience... 
    Internship
    Flexible hours

    Amazon

    Bellevue, WA
    3 days ago
  • $189k - $303k

     ...accessible for all. We are searching for an exceptional Staff-level Backend Software Engineer to join the Aurora Services Engineering team and take on...  ...along with other stakeholder teams within Aurora. Embrace AI tools to add new features which delight our users, along... 
    Full time
    Remote work

    Aurora Innovation

    Seattle, WA
    22 hours ago
  •  ...UKG is unable to offer sponsorship for this position.***   Staff Software Engineer - Agentic Acceleration Group We are seeking a highly...  ...experienced Staff Software Engineer to build and scale our agentic AI capabilities on Google's Agent Development Kit (ADK) and... 
    Full time
    Worldwide

    Ukg

    Seattle, WA
    22 hours ago
  • $210k - $250k

     ...lineage, and developer-first workflows so engineering and data teams can ship with more...  ...Role Gable is growing and hiring a new Staff Software Engineer - Front End, who will own front...  ...infrastructure. Help evaluate and integrate AI-assisted development tools and AI-driven... 
    Full time
    Local area
    3 days per week

    Gable

    Seattle, WA
    22 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer, AI Inference Runtime. Be the first to apply!