Staff Software Engineer, AI Inference Runtime
$209.1k - $282.9kARM
As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for running SOTA AI Models.You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence how efficiently new models use available compute. Partnering with our AI Infrastructure, compute, and product teams to enhance the performance and efficiency of Arm’s AI platform.Responsibilities:Define the architecture, interfaces, and roadmap for AI inference runtime capabilities, including abstractions that support evolving models, workloads, and compute platforms.Enable new model architectures end to end through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency.Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.Partner with cloud, framework, compiler, hardware, and research teams; lead technical reviews, mentor engineers, and establish meticulous performance-engineering practices.Required Skills and Experience :5+ years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference.Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior.Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, handling of memory resources, and data movement.Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.“Nice To Have” Skills and Experience :Experience developing or modifying inference schedulers, cache managers, batching systems, disaggregated or distributed execution paths.Experience optimizing kernels using accelerator programming tools, assembly, or intrinsics, including attention, matrix multiplication, operator fusion, and low-precision execution.Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization.Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries.In Return:You will be part of our AI Platforms team - A driven and diverse group passionate about developing foundational production capabilities to support AI inference at Arm. We provide a collaborative setting where your ideas can come to life quickly. The success of our AI projects will be directly influenced by your work, crafting the company’s AI inference capabilities and defining and operating production inference workloads. This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!Salary Range:$209,100-$282,900 per yearWe value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.Accommodations at ArmAt Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email View email address on click.appcast.io. To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.Hybrid Working at ArmArm’s approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team’s needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.Equal Opportunities at ArmArm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don’t discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
$236k - $330k
...this new era, we seek AI-native thinkers across... ...state of the art in LLM inference systems and... ...distributed serving and runtime systems to GPU kernels... ...We embrace AI-native engineering, using AI not only as... ...approaches to accelerate software development, experimentation...SuggestedShift work$209.1k - $282.9k
As an engineer on Arm’s AI Inference Cloud team, you will shape the technical direction and develop highly... ...with AI compute, Inference Runtime, and product teams to enhance the performance... ...workload lifecycle management.Strong software and production engineering skills, including...SuggestedWork at officeLocal areaVisa sponsorshipRelocation package$143.7k - $194.4k
AWS Neuron is the complete software stack for the AWS Inferentia... ...the Software Development Engineer for the Neuron Runtime Team, you will be responsible... ...learning applications and AI accelerators. You will work... ...scale distributed training and inference solutions. This...SuggestedInternshipWork from homeFlexible hours$168.1k - $227.4k
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud... ...This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is... ...architects, compiler engineers and runtime engineers to deliver performance and...SuggestedInternshipFlexible hours$168.1k - $227.4k
...Neuron is the complete software stack for AWS... ...This senior software engineering role is part of... ...Machine Learning Inference Applications team... ...compiler engineers, and runtime engineers to... ...Neuron, TPUs, or other AI accelerator... ...supervisors, and staff; adhere to standards...SuggestedWork experience placementInternshipLocal areaFlexible hours$209.1k - $282.9k
We seek a Software Engineer to contribute to the next generation of Physical AI platforms on Arm. Your work will focus on building... ...services, developer tooling, and runtime infrastructure that boost... ...time systems.Exposure to AI/ML inference systems, even from a platform...Work at officeLocal area$173.5k - $331.05k
...develops industry‑leading software products including... ...mission to build the modern, AI‑powered video rendering... ...for a senior, hands‑on engineer to own and evolve the cross... .../ hardware vendors ML inference integration (e.g., TensorRT, ONNX Runtime) within a render...Temporary work$229k - $343k
...digital services.We’re looking for a Staff Software Engineer to join Snap Inc on our Feature Store... ...features reliably for large-scale batch inference and low-latency online inference, maintaining... ...with responsible use of emerging AI technologies to improve engineering...Full timeLive inWork at officeLocal area$254k - $350k
..., WA / Remote (United States)Software – Software Systems /Full-time... ...alongside a team of strong software engineers and act as a force multiplier... ...cutting-edge ML Training OR Inference performance optimization... ...use artificial intelligence (AI) tools to support parts of...Full timeRemote work$201k - $315k
...MA / Seattle, WASoftware - Software Systems /Full-time /HybridZoox... ...looking for an experienced Staff Software Engineer to build, scale, and operate... ...engineering to training our AI models in Perception, Planner... ...workloads (training, inference, data generation)Experience...Full timeTemporary workRelocation package- ...Description:DataRobot delivers AI that maximizes impact and... ...DataRobot’s Fleet team is the engine behind how our platform runs... ...That’s where you come in.As a Staff Software Engineer, you’ll be responsible... ...for training and inference.Why Join the Fleet Management...Full timeLocal areaRemote workWorldwideFlexible hours
$209.1k - $282.9k
As an engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure... ..., fine-tuning, evaluation, and inference. You will work across Kubernetes... ...in developing reliable infrastructure software.Practical knowledge of Kubernetes, containers...Work at officeLocal areaVisa sponsorshipRelocation package$188k - $275k
...Staff Software Engineer, InferenceCoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams... ...traded company (Nasdaq: CRWV) in March 2025.Inference Platform Team The Inference team builds and...Permanent employmentFull timeCasual workWork at office- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...the future. We are seeking a Staff Engineer to help our development of... ...of AI training and inference at scale. As a Staff Engineer... ...~10+ years of experience in software engineering, platform engineering...Work at officeLocal areaImmediate startWork from homeFlexible hours
$254k - $350k
...Seattle, WASoftware - Autonomy Software /Full-time /HybridThe... ...Learning and System Optimization Engineer, you will orchestrate and allocate... ...allow for more efficient inference by sharing various parts of the... ...in low-level programming for AI accelerators, specifically...Full timeTemporary workRelocation package- ...notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate... ...scalable cloud platforms optimized for AI/ML workloads.Partner with AI teams to... ...architecture, ML training, and inference.Experience with Infrastructure as Code....For contractors
$209.1k - $282.9k
As a software Engineer on the AI Compute Platform team, you will design and build a secure, reliable, and easy-to-use platform for running AI workloads... ...systems, and developer tools for distributed training and inference, working closely with compute infrastructure teams, AI...Work at officeLocal areaVisa sponsorshipRelocation package- ...services that enable ML engineers and data scientists... ...monitoring, and agentic AI capabilities. We work closely... ...-driven products. As a Staff Engineer, you'll make... ..., low-latency model inference, large-scale feature stores... ..., and operating great software systems.Who you areWe'...Flexible hours
- ...community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the... ...Kubernetes for inference routing and orchestration. Ensure software solutions are optimized for peak performance during traffic...Full timeLocal areaImmediate start
$182.4k - $247k
...and running the world's best data and AI infrastructure platform so our... ...to improve their business. Founded by engineers — and customer obsessed — we leap at... ...traditional SQL query engines. As a software engineer on the Runtime team at Databricks, you will be building...Local areaWorldwide- ...reporting, and intelligent product workflows.We’re hiring a Staff Software Engineer to build the AI platforms within Data Cloud. At Rippling, you aren’t... ...team to define shared primitives for model tuning, inference, evaluation, and development of task specific agent harnesses...Work at office3 days per week
$168.1k - $227.4k
Shape the Future of AI Accelerators at AWS NeuronWe... ...Amazon Neuron, the software development kit used to... ...Trainium.As a Senior Software Engineer on our Machine Learning... ...building distributed inference support for Pytorch in... ...across compiler, runtime, framework, and hardware...InternshipWork from homeFlexible hours$143.7k - $194.4k
We are looking for a Software Development Engineer II (SDE-2) to join the EKS Node Runtime team. In this role, you will design, build... ...work will focus on enabling AI and ML workloads by implementing... ...to cutting edge AI training and inference.Mentorship: Mentor junior engineers...InternshipWork from homeFlexible hours$193.3k - $261.5k
...with the hardware and software teams to ensure the right... ...for performance engineers to develop and improve... ...teams including training, inference and runtime.* Collaborate with the... ...suite of generative AI services and other cloud... ..., supervisors, and staff; adhere to standards of...InternshipLocal areaFlexible hours- ...management (CLM).What you'll doAs a Software Engineer on the AI Platform team, you will architect and... ...systems to support large-scale model inference and data processingDesign and build resilient... ...for model serving and inference runtimes, focusing on maximizing resource...Contract workWork at officeLocal areaRemote work2 days per week
$143.7k - $194.4k
...builds AWS Neuron, the software development kit used to... ...includes an ML compiler, runtime, and application... ...enabling unparalleled ML inference and training performance... ...software boundary, our engineers build systematic infrastructure... ...of what's possible in AI acceleration.As part of...Work experience placementInternshipFlexible hours$143.7k - $194.4k
Build the AI system that decides how Amazon plans... ...metrics- Tight engineering team (8 SDEs), high autonomy... ...— including agent runtime, caching/cost-latency... ...augmented generation, causal inference, and automated... ...internship professional software development experience...InternshipFlexible hours$189k - $303k
...accessible for all. We are searching for an exceptional Staff-level Backend Software Engineer to join the Aurora Services Engineering team and take on... ...along with other stakeholder teams within Aurora. Embrace AI tools to add new features which delight our users, along...Full timeRemote work- ...UKG is unable to offer sponsorship for this position.*** Staff Software Engineer - Agentic Acceleration Group We are seeking a highly... ...experienced Staff Software Engineer to build and scale our agentic AI capabilities on Google's Agent Development Kit (ADK) and...Full timeWorldwide
$210k - $250k
...lineage, and developer-first workflows so engineering and data teams can ship with more... ...Role Gable is growing and hiring a new Staff Software Engineer - Front End, who will own front... ...infrastructure. Help evaluate and integrate AI-assisted development tools and AI-driven...Full timeLocal area3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Software Engineer, AI Inference Runtime. Be the first to apply!
- javascript software engineer Seattle, WA
- internship software Seattle, WA
- software Seattle, WA
- software intern Seattle, WA
- id software Seattle, WA
- healthcare software sales Seattle, WA
- software trainer Seattle, WA
- software sales executive Seattle, WA
- entry level software sales Seattle, WA
- software sales Seattle, WA


