Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Engineer, Model Training, Inference & Infra

International Recruiting LLC

Job Description

Job Description

Full-time · On-site · San Jose, CA · Austin, TX or Taiwan

About Our Clients

Our client is building the next generation of design automation for the semiconductor industry.

Our mission is to enable every engineering organization to build its own self-improving agentic design workforce. It combines AI agents, engineering knowledge, agent-native tools, advanced models, and continuous learning to automate complex chip-design workflows.

The team brings deep experience in artificial intelligence, electronic design automation, semiconductor design, GPU-accelerated computing, and production software systems. We work closely with leading semiconductor companies to turn advanced research into technology that improves engineering productivity, design quality, and time to market.

The Role

We are looking for an exceptional AI Engineer to own the model training, inference, and infrastructure that power its agentic design workforce.

You will drive the full model lifecycle: data pipelines, pretraining and post-training, reinforcement learning, evaluation, and high-performance serving. You will build and operate the GPU training and serving stack — on the compute substrate our infrastructure team provides — that keeps large-scale training reliable and low-latency inference efficient at production scale. You will work on evaluation systems, model training, and the self-improving, self-evolving learning loops that let our models and agents get better over time from real execution feedback.

Your work directly determines how capable, fast, and cost-effective our agents are. This role is ideal for someone who combines strong research ability with exceptional systems and performance engineering skills, and who wants the models they train and serve deployed in real semiconductor design environments—not left in notebooks or benchmarks.

What You'll Do
  • Train, post-train, and fine-tune large language models for agentic engineering workflows, including supervised fine-tuning, RLHF/RLAIF, reinforcement learning, and distillation.

  • Build scalable data pipelines for pretraining, post-training, and evaluation, including sparse, private, and domain-specific engineering data.

  • Design and operate distributed training on multi-node GPU clusters, using data, tensor, pipeline, and sequence parallelism (for example FSDP, DeepSpeed, or Megatron-style approaches).

  • Build high-throughput, low-latency inference systems with continuous batching, KV-cache management, paged attention, quantization, and speculative decoding.

  • Write and optimize custom GPU kernels (CUDA, Triton) and profile end-to-end performance across CPUs and GPUs.

  • Build core model infrastructure: cluster orchestration, ML job scheduling, checkpointing, fault tolerance, reproducibility, model and environment management, observability, and cost and utilization tracking.

  • Build automated evaluation systems, benchmarks, and reward models that measure agent capability, reliability, and regression across complex engineering tasks, including problems where design data is private or customer-specific.

  • Design self-improving and self-evolving algorithms and learning loops, where models and agents learn from execution feedback, outcomes, and new data to improve continuously over time.

  • Integrate models with agent runtimes, tool use, retrieval, and the production serving stack.

  • Improve reliability, throughput, and cost efficiency across the training and inference platform.

  • Translate promising research ideas into reliable, scalable product capabilities.

  • Collaborate with research, product, platform, and solutions teams across San Jose, Austin, and Taiwan.

  • Contribute to patents, publications, technical presentations, and the broader development of Agentic Design Automation.

What We're Looking For
  • PhD or master's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent practical experience.

  • Strong programming skills in Python and proficiency in at least one systems language such as C++ or Rust.

  • Deep experience with machine learning frameworks such as PyTorch or JAX.

  • Hands-on experience with one or more of the following:

    • Large-scale or distributed model training

    • High-performance model inference and serving

    • GPU programming and performance optimization

    • ML infrastructure and platform engineering

    • Automated evaluation, reward modeling, or self-improving and continuous-learning systems

  • Ability to take a model from data and problem formulation through training, evaluation, and production deployment.

  • Strong analytical, software engineering, performance-optimization, and debugging skills.

  • High ownership, intellectual curiosity, and willingness to work across research and product boundaries.

  • Clear written and verbal communication skills.

Particularly Valuable Experience
  • Pretraining or post-training large language models at scale.

  • Distributed training frameworks such as Megatron-LM, DeepSpeed, FSDP, or Ray.

  • Production inference engines such as vLLM, TensorRT-LLM, SGLang, or TGI.

  • Custom kernel development with CUDA, Triton, or CUTLASS.

  • Inference optimization techniques such as quantization (FP8, GPTQ, AWQ), speculative decoding, or KV-cache optimization.

  • Reinforcement learning, RLHF, or reward-model training for LLMs.

  • Automated evaluation, benchmarking, or LLM-as-judge systems for agents.

  • Self-improving, self-evolving, or continuous-learning systems, including learning from execution feedback, automated curricula, or synthetic data generation.

  • GPU cluster infrastructure with Kubernetes, Slurm, or Ray, and high-performance networking such as NCCL or InfiniBand.

  • Data pipelines and MLOps for training and continuous learning.

  • Experience deploying AI systems in enterprise or security-sensitive environments.

  • A strong record of implementation through research systems, open-source projects, production software, or technical competitions.

Why Our Client

At this company, you will have the opportunity to:

  • Help define a new category of semiconductor design technology.

  • Build the training and inference stack that powers autonomous engineering agents.

  • Develop GPU-accelerated systems that make large-scale training and low-latency serving practical and cost-effective.

  • Build the evaluation and self-improvement loops that let agents learn and get better from real engineering work.

  • Build AI systems that perform complex, consequential engineering work—not just generate recommendations.

  • Work with real semiconductor workflows, tools, and private engineering knowledge.

  • See your models deployed directly with leading chip-design organizations.

  • Work in a small, highly technical team where individual contributions can shape the product and company.

  • Collaborate with colleagues across San Jose, Austin, and Taiwan.

  • Change how chips are designed, rather than focus on only one design or one point tool.

Our client is an equal opportunity employer. We welcome candidates from diverse backgrounds who are excited to combine ambitious research with meaningful engineering impact.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Engineer, Model Training, Inference & Infra in Bellevue, WA vacancy
  •  ...multi-modality foundation model to drive the next...  ...Optimization & Deployment Engineer, you will focus on bringing...  ...highly concurrent inference code to ensure real-time...  ...maximize memory bandwidth on AI accelerators. Write...  ...with distributed training pipelines and model/tensor... 
    Training
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    25 days ago
  • $160.08k - $240.12k

     ...Description Summary The AI Scientist will work in...  ...a Senior Staff AI Engineer to design, build, deploy...  ...systems using large language models (LLMs), foundation...  ...datasets. Develop robust inference pipelines, model...  ...with large-scale model training or inference infrastructure... 
    Training
    Worldwide
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours

    GE Healthcare

    Bellevue, WA
    4 days ago
  • $149k - $279.8k

     ...research and development of large-scale video world models, including the design and construction of training datasets, foundational model algorithm design,...  ...networks and operators, model tuning for training/inference, CPU/GPU acceleration, and distributed training/inference... 
    Training
    Relocation package

    Lightspeed Studios

    Bellevue, WA
    19 hours ago
  • $80.4k - $266.3k

     ...We Are: The Global AI Infrastructure team is at the center...  ...accelerated workloads, large-scale models, simulations, and emerging...  ..., and CUDA along with LLM inference engines (TensorRT-LLM), production serving...  ...layers including multi-node training and inference workloads... 
    Training
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Kirkland, WA
    2 days ago
  • $121.6k - $243.2k

     ...Responsibilities - Develop key technologies to optimize our AI Infra stack, including training infra, inference infra, and AI agents. - Work with academia and open...  ...a Master's degree in Computer Science, Computer Engineering, or a related discipline. - Experience with at... 
    Training
    Temporary work
    Local area

    ByteDance

    Seattle, WA
    19 hours ago
  • $42.75 per hour

     ...Responsibilities: - Explore and develop key technologies for our AI Infra stack, including AI training and inference platforms, AI Agent infrastructure, etc. - Build...  ...Software Development, Computer Science, Computer Engineering, or a related technical discipline. - Experience in... 
    Training
    Hourly pay
    Summer work
    Internship
    Local area

    ByteDance

    Seattle, WA
    19 hours ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure...  ...and resiliency of AI workloads - pre-training, post-training, inference. Our objective is to deliver a stable...  ...an AI infrastructure software engineer to join our team. You'll be... 
    Training
    Full time
    Remote work

    Nvidia

    Redmond, WA
    4 days ago
  •  ...deployment of next‑generation AI systems for automated code generation...  ...for agentic GenAI, enabling models that reason, plan, and...  ...synthesis, and drive scalable training/inference pipelines while collaborating with product and engineering to deliver high-impact capabilities... 
    Training

    Jobleads-US

    Seattle, WA
    8 hours ago
  • $152k - $241.5k

     ...outstanding Senior High Performance AI Engineers to build the next generation of agentic...  ...the full agentic AI stack—from training and improving models, to designing agent architectures and...  ...and hardware stack, from models and inference through compilers, runtimes, libraries... 
    Training

    NVIDIA

    Redmond, WA
    19 hours ago
  • $154k - $200k

    Corporate AI Engineer The Corporate AI Engineer is a hands-on technical role responsible for...  ...risks, including data security, privacy, model limitations, and cost considerations....  ...limited to, location, education, skills, training, and experience. In addition to an annual... 
    Training
    Full time
    Live in
    Work at office
    Worldwide
    Flexible hours
    3 days per week

    Tanium

    Bellevue, WA
    19 hours ago
  • $176.76k - $232k

     ...company for yoga, running, training, and other athletic...  ...The Enterprise Data & AI team is a strategic and...  ...responsibilities As a Senior AI/ML Engineer, you will lead the...  ...from setting up model training and fine‑...  ...design for serving AI/ML inference solutions in production... 
    Training
    Permanent employment
    Full time
    Contract work
    Part time
    Local area
    Work visa

    Lululemon athletica

    Seattle, WA
    19 hours ago
  • $57 per hour

    Student Researcher (AI Foundation Model Infrastructure - Seed) - 2027 Start (PhD) Location...  ..., and reliability across training platforms, inference systems, compilers, and distributed...  ...in computer science, mathematics, engineering, or a related field. Strong programming... 
    Training
    Hourly pay
    Internship
    Local area

    Pangleglobal

    Seattle, WA
    1 day ago
  • $100k - $150k

     ...AI Pipeline Engineer-Remote Bright Vision Technologies is a technology consulting and software...  ...scale data systems that power modern AI training and evaluation pipelines. The role combines...  ...infrastructure choices propagate into model quality and training efficiency. Key... 
    Training
    Full time
    H1b
    Local area
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Kirkland, WA
    1 day ago
  • Job Title: AI & Data - Data Engineer Location: Seattle, WA 98109 Duration: 07 months Job Description: Ingest and process data...  ...datasets for AI/ML. Design simple, reusable data models to support training and inference. Implement data quality tests, documentation, and... 
    Training

    eTeam

    Seattle, WA
    2 days ago
  •  ...a Senior Lead Software Engineer at JPMorgan Chase within...  ...optimized for AI/ML workloads.Partner with...  ...capabilities, and skillsFormal training or certification on...  ...cloud computing delivery models (IaaS, PaaS, SaaS) and...  ...architecture, ML training, and inference.Experience with... 
    Training
    For contractors

    JP Morgan Chase

    Seattle, WA
    1 day ago
  • $110k - $145k

     ...Description Summary The Senior Data Engineer designs and builds the AWS-...  ...behind our enterprise AI applications — knowledge...  ...regressions down over time. Shape training and inference data contracts with AI...  ...loops from user signals. Data Modeling and Pipelines on AWS... 
    Training
    Permanent employment
    Contract work
    Remote work
    Visa sponsorship
    Work visa
    Relocation package

    GE Aerospace

    Redmond, WA
    4 days ago
  • $80.4k - $293.8k

     ...We are Secure, Responsible AI & Data Protection professionals...  ...threats like prompt injection, model manipulation, and unauthorized...  ...Managers are the hands-on delivery engine of the Secure AI practice....  ...playbooks, reusable tools, and training materials. The Work (Role... 
    Training
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Kirkland, WA
    2 days ago
  • $112.9k - $366.3k

     ...services company at the forefront of AI-native innovation. We partner...  ..., agent-powered workflows engineered to scale in real-world...  ...even better through world-class training, career counseling, mentoring...  ...Anthropic, or similar foundation model providers Workflow orchestration... 
    Training
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Kirkland, WA
    2 days ago
  • $144.7k - $261.3k

     ...cloud infrastructure, and ML/AI GPU platforms for AV...  ...looking for a Senior Performance Engineer to join the AV Capacity and...  ...Development: Adopt and run AV models to support GM’s long-term GPU...  ...reliability of large-scale ML training and inference environments. Your skills &... 
    Training
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours
    3 days per week

    General Motors Ventures

    Seattle, WA
    19 hours ago
  • Software Engineer III We have an exciting and rewarding...  ...enterprise-authorized AI coding assist tools within...  ..., and skills: Formal training or certification on software...  ...Exposure to ML model serving frameworks (e.g...  ...TensorFlow Serving, Triton Inference Server) Familiarity... 
    Training

    Chase

    Seattle, WA
    19 hours ago
  •  ...Roles & Responsibilities: An ML Engineer designs, builds, deploys, and maintains...  ...learning solutions across data ingestion, model training, and production deployment systems....  ...identity based access. Good to have AI engineering skills as team is building custom... 
    Training
    Full time

    Purple Drive

    Bellevue, WA
    18 days ago
  • $135.2k - $306.4k

     ...Masters, or Ph.D. in Computer Science, AI/ML, Engineering, or a related discipline, or...  ...includes optimizing large-scale GPU inference or training workloads for latency, throughput, utilization...  ...includes building or operating model serving, inference gateways, agent runtimes... 
    Training
    Full time
    Flexible hours

    Oracle

    Seattle, WA
    2 days ago
  • $290k

     ...applications that power today's AI stack using sustainable technology...  ...is looking for a Principal AI Engineer (Specialised) to lead the inference and post-training pillar of our AI systems...  ...-year technical roadmap for how models are served, evaluated, and post-... 
    Training
    Full time
    Flexible hours

    Nscale

    Seattle, WA
    2 days ago
  •  ...run it. That work changes the operating model an engineering organization runs on, the ways of...  ...the target. What we design from it is AI-native by construction. We redesign intake...  ...engineering playbooks, usage guidelines, and training materials to scale AI literacy across... 
    Training
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Seattle, WA
    3 days ago
  • $320k

    Staff + Senior Software Engineer, Inference San Francisco, CA | New York City...  ..., and steerable AI systems. We want AI to be safe...  ...Claude to life by serving our models via the industry's largest compute...  ...equivalent combination of education, training, and/or experience Required... 
    Training
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    19 hours ago
  •  ...the world?We believe building engineering is more than systems and structures...  ...centers driving the future of AI to dynamic commercial...  ...world. In the role of Senior BIM Model Manager - Digital Delivery Lead...  ...software, versions, access, and training needed to execute the project... 
    Training
    Work at office

    HDR

    Bellevue, WA
    2 days ago
  • $188k - $275k

    CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers...  ...2025. Learn more at What You'll Do: Inference Platform Team The Inference team builds...  .... About the role: As a Staff Software Engineer (IC5) on the Inference team, you will act... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    19 hours ago
  • $61k - $101k

     ...Requirements: We require formal training or certification in software engineering concepts, along with 5+...  ...of cloud delivery models such as IaaS, PaaS, and...  ...architecture, training, and inference. We require experience...  ...of enterprise-approved AI-assisted software development... 
    Training
    Full time
    For contractors

    J.P. Morgan

    Seattle, WA
    9 days ago
  •  ...our PremierUnified customers and Customer Service and Support engineers to manage complex reactive support scenarios and provide insights...  ...job-related skills, experience, and relevant education or training. Depending on the position offered, other forms of compensation... 
    Training
    Temporary work
    Work at office
    Local area

    LTM

    Bellevue, WA
    1 day ago
  • $115k - $231k

     ...AI Engineer Location US-WA-Bothell ID 2026-1717 Category Information Technology Position Type Full Time Work Model Hybrid Company Overview Verathon is a global medical device company focused...  ...functional teams through adoption, training, and scaling of deployed AI tools.... 
    Training
    Full time

    Verathon

    Bothell, WA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Engineer, Model Training, Inference & Infra. Be the first to apply!