Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff AI Product Engineer

$220k - $293.33k
Full-time

Nscale

About Nscale

Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.

About the Role

Nscale is looking for a Staff AI Engineer (Specialized) to set technical direction for the inference and reinforcement learning systems at the core of our AI services platform — and for the APIs through which other engineers consume them.

This role owns the architecture of how models are served on Nscale’s GPU cloud, how RL and post-training workloads run on it, and how both are exposed to customers and internal teams as clean, reliable, high-performance interfaces. You’ll work across 2–4 teams spanning serving, post-training, and platform, defining how these systems are built and establishing the standards that create engineering leverage across the organization.

As a Staff engineer, you are the technical authority for this area of the AI stack. Your decisions determine the latency, throughput, and cost profile of every token Nscale serves, and the correctness and efficiency of every RL run on our platform. You resolve ambiguous architectural questions where the answer space is genuinely open — disaggregated versus co-located serving, on-policy versus off-policy RL infrastructure, where the API boundary should sit — and your solutions become the standards others build on.

Responsibilities

Inference

  • Set technical direction for Nscale’s inference serving architecture: request routing, scheduling, continuous batching, KV cache management, prefix caching, and speculative decoding
  • Drive the strategy for model efficiency in production — quantization (FP8, INT8/4), sparsity, pruning, distillation, and MoE serving — and the trade-offs between cost, latency, throughput, and model quality
  • Lead the resolution of systemic performance and reliability challenges across the serving stack, from kernel-level bottlenecks to fleet-level capacity and multi-tenant isolation

Reinforcement learning & post-training

  • Own the architecture of Nscale’s RL and post-training systems: RLHF, DPO/GRPO-style methods, reward modelling, and agentic RL with tool calling, off-policy training, and decoupled sampling and policy updates
  • Define how inference and training share infrastructure in RL loops — rollout generation, sample buffering, weight synchronization, and the serving engine’s role inside the training system
  • Establish standards for fine-tuning services (LoRA, QLoRA, adapters, full fine-tuning) and the data curation and processing workflows that feed them

Developer APIs & platform

  • Set the design standards for Nscale’s developer-facing APIs, SDKs, and tooling — OpenAI-compatible and native interfaces, OpenAPI 3.x specifications, versioning, and rate limiting — so that inference and RL capabilities are consumable by engineers who never see the underlying systems
  • Create reusable frameworks and tooling that multiply the effectiveness of other AI engineers across Nscale

Leadership & direction

  • Coach and grow more junior engineers across teams; raise technical capability and engineering quality broadly
  • Collaborate with research, product, and infrastructure leadership to align inference and RL platform strategy with customer demand and business direction
  • Evaluate emerging serving engines, RL frameworks, and accelerator technologies; make decisive build/adopt/contribute recommendations
  • Represent the inference and RL domain in cross-organizational architecture reviews and strategic planning

Requirements

  • 8–12 years of engineering experience, with significant depth in production AI systems at scale (AI labs, hyperscalers, or leading ML infrastructure companies)
  • Demonstrated ability to set technical direction for an AI systems domain at scale
  • Deep expertise in production LLM inference: serving architectures, KV cache and memory management, batching and scheduling strategies, speculative decoding, and low-precision inference
  • Strong hands-on expertise in RL for LLMs — RLHF, DPO and related preference-optimization methods, reward modelling, or agentic/multi-turn RL — including the systems that make them run efficiently on GPU clusters
  • Proven ability to design developer-facing APIs and SDKs that are clean, versioned, and adopted by other engineers and external customers
  • Strong cross-team influence: track record of creating standards and practices adopted by multiple teams
  • Experience designing AI infrastructure with clear control plane / data plane separation and cell-based architecture patterns for scale-out and blast-radius isolation
  • Experience with large-scale GPU/accelerator workloads: CUDA or ROCm, memory bandwidth optimization, and distributed compute paradigms (data/model/tensor parallelism, sharding, scheduling)
  • Strong proficiency in Python and PyTorch, with a track record of building maintainable, well-tested, production-grade ML systems
  • Deep understanding of transformer and LLM architectures and their behavior under production load
  • Ability to design architecture that is both technically excellent and practically adoptable across a diverse engineering organization

Preferred

  • Contributions to widely-used open-source inference or RL frameworks (vLLM, SGLang, TensorRT-LLM, verl, OpenRLHF, TRL, DeepSpeed, etc.)
  • Experience at an AI lab (OpenAI, DeepMind, Anthropic, Meta AI, etc.), hyperscaler AI team, or leading ML infrastructure company
  • Published research or technical writing on inference systems or RL infrastructure (MLSys, NeurIPS, ICLR systems tracks, blog posts)
  • Experience with hardware-software co-design for AI accelerators (custom CUDA kernels, Triton, etc.)
  • Experience operating inference platforms in containerized, distributed environments (Kubernetes, large-scale clusters)
  • Experience building evaluation and benchmarking systems for model quality, safety, and system performance

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$220,000—$293,333 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff AI Product Engineer in Seattle, WA vacancy
  • $148.1k - $282.1k

    The Opportunity We're looking for a Principal Product Manager with deep fluency in the AI ecosystem to lead the evolution of our engineering platform. The rapid adoption of AI is fundamentally reshaping how we build, ship, and scale software at Adobe. To keep pace, we... 
    Suggested
    Full time
    Temporary work
    Local area
    Immediate start
    Worldwide

    Adobe Systems

    Seattle, WA
    3 days ago
  • $153.71k - $267.03k

     ...new areas of inspiration and expand your capabilities, then consider a career in Advisory. KPMG is currently seeking a Manager, AI Engineer to join our Advisory Services practice. Responsibilities: End-to-end design and development of AI/ML solutions,... 
    Suggested
    Full time
    H1b
    Local area

    KPMG

    Seattle, WA
    6 days ago
  • We Are:Accenture's Oracle Business Group AI Center of Excellence is one of the most active...  ...in the Oracle ecosystem.We build real products that close deals and run in production —...  ...client conversation.You might come from engineering, consulting, product, or pre-sales — what... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Seattle, WA
    1 day ago
  • Job Title: Applied AI EngineerLocation: 100% RemoteFull-Time Role Role As an Applied AI Engineer, you will turn model capabilities into real product behavior.You will own problems end-to-end, from shaping model behavior, to building the systems around it, to ensuring it... 
    Suggested

    Spectraforce Technologies

    Seattle, WA
    4 days ago
  • $157.9k - $213.6k

     ...seeking an experienced Senior Compliance Engineer — Environmental to drive global product compliance for hazardous substances...  ..., including hardware supporting AI/ML infrastructure. This role sits...  ...other employees, supervisors, and staff; adhere to standards of excellence... 
    Suggested
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $91.1k - $179.5k

    Position Summary Agentic AI is moving from experimentation to production, and organizations everywhere are racing to figure out how to build it responsibly and at scale. We're growing a team of engineers who want to work at the center of that shift: designing and... 
    Work at office
    Local area
    Visa sponsorship
    Shift work

    Deloitte

    Seattle, WA
    1 day ago
  •  ...everything across a business—until generative AI. Today, AI is the number one driver of...  ....You Are:As a Snowflake Advanced AI Engineer, you will design, build, and operationalize...  ...cloud and third-party AI services to deliver production-ready outcomes. Your role spans the full... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area
    Shift work

    Accenture

    Seattle, WA
    1 day ago
  • $197.3k - $313.7k

     ...DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive...  ...listed, this is a combination UX/AI/Agentic Engineering role*The Experience:We’re hiring a...  ...uncertainty, and reversibility.UX Prototyping to Production: Rapidly prototype new interaction... 
    Full time
    Remote work
    Shift work

    Salesforce

    Seattle, WA
    13 hours ago
  • $124k - $155k

     ...lives with trusted financial services that transcend borders.About the Role:As an AI Native Software Engineer, you orchestrate a fleet of AI agents to architect, build, and deploy production-grade systems, and near-100% of your production code is agent-generated. You... 
    Full time
    Work at office
    Worldwide
    Flexible hours
    3 days per week

    Remitly

    Seattle, WA
    1 day ago
  • $140k - $180k

     ...WashingtonValorem Reply Azure SI, US - Data + AI /Full Time /HybridValorem Reply is an...  ..., IT modernization, customer experience, product transformation and digital workplace....  ...their clients do business. As a Senior AI Engineer, you will understand how AI is positioned... 
    Full time

    Valorem Reply

    Seattle, WA
    3 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will...  ...ML models more efficient leading to significant productivity improvements and cost savingsBuild tools, frameworks,... 
    Full time
    Remote work

    Nvidia

    Seattle, WA
    3 days ago
  • $137.4k - $161.7k

    Are you ready to make an impact?Senior AI Software Engineer - Technology & Experience (TechEx) West Monroe is seeking a Senior AI Software...  ...a hands-on engineering role focused on building scalable, production-grade software applications enhanced by modern AI capabilities... 
    Local area
    Immediate start
    Flexible hours

    West Monroe Partners

    Seattle, WA
    21 hours ago
  •  ...building a new portfolio of agentic software products in the industries we have spent decades...  ...bring product management, design, and engineering together with the autonomy to move quickly...  ...engineering role. You Are You are an exceptional AI-native software engineer who has built... 
    Permanent employment
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Seattle, WA
    1 day ago
  • $152k - $241.5k

     ...recently, GPU deep learning ignited modern AI — the next era of computing — with the...  ...company”.NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep...  ...to thrive in a fast-moving, dynamic, product-oriented team.Ways to stand out from the... 
    Full time
    Remote work

    Nvidia

    Seattle, WA
    1 day ago
  •  ...Vamsi SattaruCompany: SRI Tech SolutionsPosition Title: Agentic AI Software EngineerLocation: Hybrid, Seattle, WAEmployment Type:...  ...Minimum Qualifications: Bachelor’s degree in Computer Science, Engineering, or a closely related discipline. Demonstrated experience in software... 
    Full time

    SRI Tech

    Seattle, WA
    21 hours ago
  •  ...That work changes the operating model an engineering organization runs on, the ways of working...  ...sets the target. What we design from it is AI-native by construction. We redesign...  ...We treat the lifecycle as a system and a product itself - built to operate at a speed required... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Seattle, WA
    1 day ago
  •  ...Building the Future in the Age of AgentsJoin the elite technical and product engine of the Accenture Google Business Group (AGBG). We are not a...  ...the most significant shift in technology: the move to Agentic AI and Product-Led Operating Models. As a Google Cloud Platform (... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area
    Shift work

    Accenture

    Seattle, WA
    21 hours ago
  • $110.7k - $372.9k

     ...afford the care they need. Deloitte has a new AI-first effort, backed by $1B in committed...  ...end to end — from architecture through production — and your work ships into live clinical...  ...months, not into a lab. As an Agentic AI Engineer, you will design, build, and... 
    Local area
    Visa sponsorship

    Deloitte

    Seattle, WA
    2 days ago
  • $152k - $241.5k

     ...technology powers everything from generative AI to autonomous systems, and we continue to...  ..., and tools that enable researchers and engineers to develop the next generation of AI/ML...  ...solid distributed systems fundamentals, production-grade coding, and a passion for operational... 
    Full time

    Nvidia

    Seattle, WA
    3 days ago
  • $258.8k - $304.2k

    Are you ready to make an impact?Applied AI Engineer/Architect West Monroe is seeking an Applied AI Engineer/Architect to lead the design...  ...augmented generation (RAG), tool integrations, and secure production deployment patterns.Define semantic models, ontologies, business... 
    Local area
    Immediate start
    Flexible hours

    West Monroe Partners

    Seattle, WA
    1 day ago
  •  ...with solutions created by the #1 company in e-signature and contract lifecycle management (CLM). What you'll do As an AI Agentic Engineer on the Platform AI Engineering team, you will drive the next wave of automation in IT operations by designing and deploying agentic... 
    Permanent employment
    Full time
    Contract work
    Work at office
    Local area
    Remote work
    2 days per week

    DocuSign

    Seattle, WA
    3 days ago
  • $342.7k

     ...Intelligence is a foundational pillar in the product strategy for Splunk and Cisco, and we are building best-in-class AI capabilities into our platform, security and observability...  .... We are looking for a Distinguished Engineer with outstanding technical leadership in the... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Seattle, WA
    1 day ago
  • $130k - $170k

    CompanyFounded by CPAs, tax attorneys, and engineers, Taxbit is the leading innovator...  ...reporting for the digital economy. Taxbit's AI-enabled platform streamlines compliance related...  ...- building intelligent, customer-facing products that redefine how users interact with tax... 
    Work at office
    Work from home

    Taxbit

    Seattle, WA
    21 hours ago
  •  ...thinking services company at the forefront of AI-native innovation. We partner with...  ...next-generation, agent-powered workflows engineered to scale in real-world settings. Our engineers...  ...real-world business problems. It is a product-minded, customer-embedded engineering role... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Seattle, WA
    3 days ago
  • $114.9k - $154.1k

    Job Posting Title:Product Software Engineer IIReq ID:10152986Job Description:Disney Entertainment and ESPN Product & TechnologyTechnology is at the...  ...and experiment with emerging technologies and AI-assisted engineering workflows to improve developer productivity... 
    Full time
    Work experience placement

    Hulu

    Seattle, WA
    2 days ago
  • $148.7k - $199.4k

    Job Posting Title:Senior Product Software Engineer - FrontendReq ID:10153777Job Description:Technology is at the heart of Disney’s past, present,...  ...grounded search, and tool-use interactions, partnering with senior AI-application engineers on the team.Help establish and refine... 
    Full time
    Work experience placement

    Hulu

    Seattle, WA
    3 days ago
  • $141.9k - $190.3k

    Job Posting Title:Senior Product Software Engineer (Swift)Req ID:10154354Job Description:Department OverviewTechnology is at the heart of Disney’s past, present, and future. Disney Entertainment and ESPN Product & Technology is a global organization of engineers, product... 
    Full time

    Hulu

    Seattle, WA
    2 days ago
  • $73.5k - $212.28k

     ...Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to...  ...continuous improvement initiatives within teams- Understanding production-grade quality gates and monitoringTravel RequirementsUp to 60... 
    Full time
    H1b

    PwC

    Seattle, WA
    1 day ago
  • $87.62k

     ...an outstanding opportunity for AI Security Engineerto join their...  ...Manager, the AI Security Engineer will support the security, governance...  ...to over 50,000 faculty, staff, and students across three campuses...  ...cloud security benchmark.- Production multi-tenant SaaS or AI-platform... 
    Full time
    Temporary work
    Work at office
    Monday to Friday
    Flexible hours
    Shift work
    2 days per week

    University of Washington

    Seattle, WA
    1 day ago
  • $175k - $205.9k

     ...to make an impact?West Monroe is searching for a Principal AI Software Engineer, to join our team in our Technology & Experience (TechEx) practice...  ...with our clients. You will collaborate closely with product, engineering, and design teams to architect and own AI-driven... 
    Local area
    Immediate start
    Flexible hours

    West Monroe Partners

    Seattle, WA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff AI Product Engineer. Be the first to apply!