Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer, AI Inference Runtime

$209.1k - $282.9k

ARM

As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for running SOTA AI Models.You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence how efficiently new models use available compute. Partnering with our AI Infrastructure, compute, and product teams to enhance the performance and efficiency of Arm’s AI platform.Responsibilities:Define the architecture, interfaces, and roadmap for AI inference runtime capabilities, including abstractions that support evolving models, workloads, and compute platforms.Enable new model architectures end to end through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency.Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.Partner with cloud, framework, compiler, hardware, and research teams; lead technical reviews, mentor engineers, and establish meticulous performance-engineering practices.Required Skills and Experience :5+ years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference.Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior.Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, handling of memory resources, and data movement.Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.“Nice To Have” Skills and Experience :Experience developing or modifying inference schedulers, cache managers, batching systems, disaggregated or distributed execution paths.Experience optimizing kernels using accelerator programming tools, assembly, or intrinsics, including attention, matrix multiplication, operator fusion, and low-precision execution.Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization.Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries.In Return:You will be part of our AI Platforms team - A driven and diverse group passionate about developing foundational production capabilities to support AI inference at Arm. We provide a collaborative setting where your ideas can come to life quickly. The success of our AI projects will be directly influenced by your work, crafting the company’s AI inference capabilities and defining and operating production inference workloads. This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!Salary Range:$209,100-$282,900 per yearWe value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.Accommodations at ArmAt Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email View email address on us.fitly.work. To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.Hybrid Working at ArmArm’s approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team’s needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.Equal Opportunities at ArmArm is an equal opportunity employer, committed to providing an environment of mutual respect where equal opportunities are available to all applicants and colleagues. We are a diverse organization of dedicated and innovative individuals, and don’t discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer, AI Inference Runtime in Washington DC vacancy
  • $236k - $330k

     ...this new era, we seek AI-native thinkers across...  ...state of the art in LLM inference systems and...  ...distributed serving and runtime systems to GPU kernels...  ...We embrace AI-native engineering, using AI not only as...  ...approaches to accelerate software development, experimentation... 
    Suggested
    Shift work

    Snowflake

    Washington DC
    6 days ago
  • $209.1k - $282.9k

    As an engineer on Arm’s AI Inference Cloud team, you will shape the technical direction and develop highly...  ...with AI compute, Inference Runtime, and product teams to enhance the performance...  ...workload lifecycle management.Strong software and production engineering skills, including... 
    Suggested
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Washington DC
    6 days ago
  •  ...Model Optimization & Deployment Engineer, you will focus on bringing...  ..., and build highly concurrent inference code to ensure real-time, deterministic...  ...maximize memory bandwidth on AI accelerators. Write...  ..., DeepSpeed, Megatron-LM) and runtime efficiency optimization for GPU... 
    Suggested
    Temporary work
    Relocation package

    Zoox

    Washington DC
    24 days ago
  •  ...About the Team Join the engineering teams that bring OpenAI...  ...the benefits of AI, while ensuring that this...  ...Role We’re seeking Software Engineers who can solve...  ...optimizing how we serve inference in unique, high-stakes...  ...title Member of Technical Staff . We use Senior Staff... 
    Suggested
    Full time

    OpenAI

    Washington DC
    11 hours ago
  • $92k - $135k

     ...Description CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,...  ...more at  What You'll Do: Join the Inference team to ship production features that improve...  ...quickly with mentorship from experienced engineers. About the role: Implement well-... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Flexible hours

    CoreWeave

    Washington DC
    17 days ago
  •  ...Provide high-performance inference infrastructure for machine-learning...  ...new model architectures and AI accelerator platforms....  ...Requirements Significant software engineering experience, particularly with...  ...Hybrid work policy requiring staff to work from an Anthropic office... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    Washington DC
    26 days ago
  •  ...You will work alongside a team of strong software engineers and act as a force multiplier for our...  ...operation of cutting-edge ML Training OR Inference performance optimization techniques to...  ...We may use artificial intelligence (AI) tools to support parts of the hiring process... 

    Zoox

    Washington DC
    23 days ago
  • $210k - $250k

     ...lineage, and developer-first workflows so engineering and data teams can ship with more...  ...Role Gable is growing and hiring a new Staff Software Engineer - Front End, who will own front...  ...infrastructure. Help evaluate and integrate AI-assisted development tools and AI-driven... 
    Full time
    Local area
    3 days per week

    Gable

    Washington DC
    11 hours ago
  • $189k - $303k

     ...accessible for all. We are searching for an exceptional Staff-level Backend Software Engineer to join the Aurora Services Engineering team and take on...  ...along with other stakeholder teams within Aurora. Embrace AI tools to add new features which delight our users, along... 
    Full time
    Remote work

    Aurora Innovation

    Washington DC
    11 hours ago
  • $187.53k - $281.3k

     ...Founded in 2015, Shield AI is a venture-backed deep-tech company...  ...a diverse team of experts in software, robotics, control systems, optimization...  ..., hardware, and test engineering to solve some of the hardest...  ...customers. About the Job: Staff Software Engineers on the... 
    Full time
    Temporary work
    Part time
    Worldwide

    Shield Ai

    Washington DC
    11 hours ago
  •  ...enterprise. To usher in this new era, we seek AI-native thinkers across every function...  ...users. But it didn’t stop there. They engineered Snowflake to power the Data Cloud, where...  ...from Principal engineers. AS A STAFF SOFTWARE ENGINEER - IDENTITY & ACCESS MANAGEMENT,... 
    Full time

    Snowflake

    Washington DC
    11 hours ago
  • $187.53k - $281.3k

     ...Shield AI is a venture-backed defense-tech company with the mission of protecting...  ...Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon...  ...is both exciting and crucial.   As a Staff engineer, you will lead the technical delivery of... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Work at office
    Worldwide

    Shield Ai

    Washington DC
    11 hours ago
  • $200k - $275k

     ...with unprecedented speed and accuracy. Our AI-enabled platform turns siloed and...  ...internationally.  Our Team As an engineering team, we believe strongly that empathy improves...  ...serves. We’re looking for a Staff Software Engineer to join our core engineering teams... 
    Full time
    Work at office
    Local area

    Peregrine Technologies, Llc

    Washington DC
    11 hours ago
  •  ...Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers...  ...home day is currently Tuesday.   About the Role As a Staff Software Engineer for the Compute pillar, you will play a critical role in... 
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    Washington DC
    11 hours ago
  • $220k - $275k

     ...patient's request for their medical records to powering the AI revolution in healthcare, Datavanters are building the future...  ...about creating transformative change in healthcare. Staff Software Engineer The Role As a Staff Software Engineer, you will shape... 

    Datavant

    Washington DC
    1 day ago
  • $230k - $280k

     ...Staff Software Applied AI Engineer At HackerOne, we're advancing a new era of AI-powered offensive security. As a Staff AI Engineer, you'll help shape the evolution of our autonomous HAI platform, driving the integration of advanced AI and agentic frameworks into HackerOne... 
    Apprenticeship
    Work at office
    Local area
    Remote work
    Flexible hours
    Shift work
    1 day per week

    HackerOne

    Washington DC
    4 days ago
  • $150k - $180k

     ...Role Title: Staff Software Engineer Pearson's PSG-CP (Pearson Software Group - Content Platform) team is looking for a forward deployed engineer...  ...with our most strategic internal lines of business to drive AI adoption and ship working software against real problems.... 
    Full time
    Remote work
    Shift work

    Pearson

    Washington DC
    4 days ago
  • $150k - $203k

     ...The Nuclear Company is the fastest growing AI tech-enabled startup in the nuclear and energy space, pioneering a fleet-scale...  ...commitment to our mission. About the role As a Staff Software Engineer, you will serve as a technical leader responsible for architecting... 
    Bi-weekly pay

    The Nuclear Company

    Washington DC
    4 days ago
  • $260.1k

     ...surges.”learn more about working at Coinbase. As a Senior Staff Software Engineer on thePlatform Payments team, you'll define the engineering...  ...on high-impact technical decisions. ~ Utilizes generative AI responsibly, maintaining human oversight to deliver business... 
    Local area

    Coinbase

    Washington DC
    3 days ago
  • $100 per hour

     ...Staff Software Engineer Washington, D.C We're on a mission to bring joy to the business of events. Weddings, galas, festivals... the moments...  ...the velocity of the entire Engineering Team. Raise the AI ceiling. AI is changing daily, and we are changing just as fast... 
    Work at office
    Work from home

    Goodshuffle Pro

    Washington DC
    1 day ago
  • $115k - $230k

     ...External Job Posting Description We are seeking an accomplished Staff Software Engineer with a proven track record in Java development, extensive...  ..., Agent Skills, and the end-to-end delivery of Generative AI applications. We require experience in leveraging AI-... 
    Hourly pay
    Work experience placement
    Local area

    GEICO

    Bethesda, MD
    4 days ago
  •  ...Overview: The Staff Software Engineer L5 works across all service aspects of high-throughput and multi-tenant systems, with the ability to design...  ...; Experience with or strong familiarity using AI-assisted development tools such as GitHub Copilot or Claude... 
    H1b

    Inovalon

    Bowie, MD
    1 day ago
  • $143.7k - $194.4k

    We are looking for a Software Development Engineer II (SDE-2) to join the EKS Node Runtime team. In this role, you will design, build...  ...work will focus on enabling AI and ML workloads by implementing...  ...to cutting edge AI training and inference.Mentorship: Mentor junior engineers... 
    Internship
    Work from home
    Flexible hours

    Amazon

    Washington DC
    a month ago
  • $143.7k - $194.4k

     ...Neuron is the complete software stack for the AWS Inferentia...  ...Software Development Engineer for the Neuron...  ...learning applications and AI accelerators. You will...  ...the our C++ compiler and runtime generates key information...  ...distributed training and inference solutions. This... 
    Internship
    Work from home
    Flexible hours

    Amazon

    Washington DC
    a month ago
  • $71.5k - $190k

     ...AI Engineering And Delivery Lead Systems Planning and Analysis, Inc...  ...model endpoints and optimize inference workloads across local, hybrid...  ..., and local/on-prem serving runtimes (e.g., ONNX Runtime, vLLM,...  ...~5-7 years in enterprise software or systems engineering, with... 
    Contract work
    Work at office
    Local area
    Remote work

    Systems Planning and Analysis, Inc

    Washington DC
    5 days ago
  •  ...AI/ML Software Engineer At Gallatin, we are rebuilding logistics infrastructure for the national...  ...large scale ML pipelines and real-time inference systems—while collaborating with cross...  ...streamline model training, deployment, runtime, and monitoring. Own high model... 
    Local area

    Gallatin AI, Inc.

    Washington DC
    1 day ago
  • $190k - $225k

     ...AssemblyAI builds the best-in-class Voice AI models powering the next generation of...  ...applications. Our models serve 600M+ inference calls monthly, process 1M+ hours of audio...  ...! About the role: We're hiring a Software Engineer to help turn cutting-edge AI research into... 

    AssemblyAI

    Washington DC
    19 days ago
  •  ...what turn a massive, multi-tenant data and AI platform into one customers can trust at...  ...across all these surfaces, raising the engineering bar of the combined team, and shaping the...  ...governance.Champion reliable, high-quality software and the operational practices that let a... 
    Worldwide

    DataBricks

    Washington DC
    a month ago
  • $177.19k - $364.8k

     ...work. Creating a career you love? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful partner that augments...  ...conversions in a privacy-preserving way. We are hiring a Staff Software Engineer to lead the backend architecture and implementation of a GenAI... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    Washington DC
    a month ago
  • $250k - $300k

     ...energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and...  ...build with us at Crusoe. About This Role As a Senior Staff Software Engineer for SDN Architecture, you will be a driving technical force... 
    Temporary work

    Crusoe

    Washington DC
    25 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer, AI Inference Runtime. Be the first to apply!