Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Engineer, Model Training, Inference & Infra

International Recruiting LLC

Job Description

Job Description

Full-time · On-site · San Jose, CA · Austin, TX or Taiwan

About Agentrys

Agentrys is building the next generation of design automation for the semiconductor industry.

Our mission is to enable every engineering organization to build its own self-improving agentic design workforce. Agentrys Studio combines AI agents, engineering knowledge, agent-native tools, advanced models, and continuous learning to automate complex chip-design workflows.

Our team brings deep experience in artificial intelligence, electronic design automation, semiconductor design, GPU-accelerated computing, and production software systems. We work closely with leading semiconductor companies to turn advanced research into technology that improves engineering productivity, design quality, and time to market.

The Role

We are looking for an exceptional AI Engineer to own the model training, inference, and infrastructure that power Agentrys' agentic design workforce.

You will drive the full model lifecycle: data pipelines, pretraining and post-training, reinforcement learning, evaluation, and high-performance serving. You will build and operate the GPU training and serving stack — on the compute substrate our infrastructure team provides — that keeps large-scale training reliable and low-latency inference efficient at production scale. You will work on evaluation systems, model training, and the self-improving, self-evolving learning loops that let our models and agents get better over time from real execution feedback.

Your work directly determines how capable, fast, and cost-effective our agents are. This role is ideal for someone who combines strong research ability with exceptional systems and performance engineering skills, and who wants the models they train and serve deployed in real semiconductor design environments—not left in notebooks or benchmarks.

What You'll Do
  • Train, post-train, and fine-tune large language models for agentic engineering workflows, including supervised fine-tuning, RLHF/RLAIF, reinforcement learning, and distillation.

  • Build scalable data pipelines for pretraining, post-training, and evaluation, including sparse, private, and domain-specific engineering data.

  • Design and operate distributed training on multi-node GPU clusters, using data, tensor, pipeline, and sequence parallelism (for example FSDP, DeepSpeed, or Megatron-style approaches).

  • Build high-throughput, low-latency inference systems with continuous batching, KV-cache management, paged attention, quantization, and speculative decoding.

  • Write and optimize custom GPU kernels (CUDA, Triton) and profile end-to-end performance across CPUs and GPUs.

  • Build core model infrastructure: cluster orchestration, ML job scheduling, checkpointing, fault tolerance, reproducibility, model and environment management, observability, and cost and utilization tracking.

  • Build automated evaluation systems, benchmarks, and reward models that measure agent capability, reliability, and regression across complex engineering tasks, including problems where design data is private or customer-specific.

  • Design self-improving and self-evolving algorithms and learning loops, where models and agents learn from execution feedback, outcomes, and new data to improve continuously over time.

  • Integrate models with agent runtimes, tool use, retrieval, and the production serving stack.

  • Improve reliability, throughput, and cost efficiency across the training and inference platform.

  • Translate promising research ideas into reliable, scalable product capabilities.

  • Collaborate with research, product, platform, and solutions teams across San Jose, Austin, and Taiwan.

  • Contribute to patents, publications, technical presentations, and the broader development of Agentic Design Automation.

What We're Looking For
  • PhD or master's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent practical experience.

  • Strong programming skills in Python and proficiency in at least one systems language such as C++ or Rust.

  • Deep experience with machine learning frameworks such as PyTorch or JAX.

  • Hands-on experience with one or more of the following:

    • Large-scale or distributed model training

    • High-performance model inference and serving

    • GPU programming and performance optimization

    • ML infrastructure and platform engineering

    • Automated evaluation, reward modeling, or self-improving and continuous-learning systems

  • Ability to take a model from data and problem formulation through training, evaluation, and production deployment.

  • Strong analytical, software engineering, performance-optimization, and debugging skills.

  • High ownership, intellectual curiosity, and willingness to work across research and product boundaries.

  • Clear written and verbal communication skills.

Particularly Valuable Experience
  • Pretraining or post-training large language models at scale.

  • Distributed training frameworks such as Megatron-LM, DeepSpeed, FSDP, or Ray.

  • Production inference engines such as vLLM, TensorRT-LLM, SGLang, or TGI.

  • Custom kernel development with CUDA, Triton, or CUTLASS.

  • Inference optimization techniques such as quantization (FP8, GPTQ, AWQ), speculative decoding, or KV-cache optimization.

  • Reinforcement learning, RLHF, or reward-model training for LLMs.

  • Automated evaluation, benchmarking, or LLM-as-judge systems for agents.

  • Self-improving, self-evolving, or continuous-learning systems, including learning from execution feedback, automated curricula, or synthetic data generation.

  • GPU cluster infrastructure with Kubernetes, Slurm, or Ray, and high-performance networking such as NCCL or InfiniBand.

  • Data pipelines and MLOps for training and continuous learning.

  • Experience deploying AI systems in enterprise or security-sensitive environments.

  • A strong record of implementation through research systems, open-source projects, production software, or technical competitions.

Why Agentrys

At Agentrys, you will have the opportunity to:

  • Help define a new category of semiconductor design technology.

  • Build the training and inference stack that powers autonomous engineering agents.

  • Develop GPU-accelerated systems that make large-scale training and low-latency serving practical and cost-effective.

  • Build the evaluation and self-improvement loops that let agents learn and get better from real engineering work.

  • Build AI systems that perform complex, consequential engineering work—not just generate recommendations.

  • Work with real semiconductor workflows, tools, and private engineering knowledge.

  • See your models deployed directly with leading chip-design organizations.

  • Work in a small, highly technical team where individual contributions can shape the product and company.

  • Collaborate with colleagues across San Jose, Austin, and Taiwan.

  • Change how chips are designed, rather than focus on only one design or one point tool.

Agentrys is an equal opportunity employer. We welcome candidates from diverse backgrounds who are excited to combine ambitious research with meaningful engineering impact.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Engineer, Model Training, Inference & Infra in Washington DC vacancy
  • $197.3k - $225.1k

     ...Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable...  ...to millions of customers. Our AI models and platforms empower teams across...  ...including foundation model training, large language model inference, similarity... 
    Training
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    3 days ago
  • $229.9k - $262.4k

    ## Senior Lead AI Engineer (FM Hosting, LLM Inference)Applylocations: McLean, VA: New York, NY: San Jose, CAtime...  ...to millions of customers. Our AI models and platforms empower teams across...  ...components including foundation model training, large language model inference,... 
    Training
    Full time
    Part time
    Local area

    Capital One Group

    McLean, VA
    2 days ago
  •  ...individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI...  ...bridging the gap between high-performance model development and optimized deployment environments...  ..., compensation, promotion, benefits, training, discipline, and termination. F5... 
    Training
    Full time
    Local area
    Immediate start

    F5 Networks

    Washington DC
    15 days ago
  • $77.6k - $176k

    AI Model EngineerThe Opportunity: As an AI Model Engineer, you’ll advance organizational strategy to deliver leading edge...  ...years of experience developing, training, and deploying ML models into production...  ...toolsExperience designing inference pipelines optimized for... 
    Training
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Springfield, VA
    6 days ago
  •  ...what is delivered. Our AI-native platform, Air Enterprise...  ...experienced Senior AI Engineer to join our Agentic AI...  ...agent architectures, model integrations, tools,...  .... Develop model and inference infrastructure...  ...with fine‑tuning, post‑training, reinforcement learning... 
    Training
    Full time
    Work at office
    Remote work

    artificial intelligence and robotics laboratory (itu air lab...

    Arlington, VA
    2 days ago
  • $75 - $80 per hour

     ...Matlen Silver Job Title: Senior AI/LLM Engineer Duration: 12+ Months...  ...strong focus on Large Language Models (LLMs). The ideal candidate...  ...databases, and distributed training. Proven experience deploying...  .../NLP. Experience optimizing inference on GPUs, TPUs, or other accelerators... 
    Training
    Full time

    Matlen Silver

    Washington DC
    1 day ago
  • $131.3k - $237.35k

     ...you ready to thrive? Through training, teamwork, and exposure to challenging...  ...deploying enterprise-scale AI, data, and mission platform...  ...next level. As a Senior AI Engineer, you will : Support the...  ...APIs Python development GPU inference optimization Preferred Qualifications... 
    Training
    Local area
    Immediate start
    Remote work
    Flexible hours

    Leidos

    Alexandria, VA
    2 days ago
  • $100k - $110k

     ...AI Engineer Triunity Software is looking for a talented and passionate AI Engineer...  ...and deploy machine learning and AI models into production environments...  ...maintain data pipelines to support model training, evaluation, and inference Research and implement state-of-... 
    Training

    Triunity Software

    Washington DC
    3 days ago
  •  ...AI Engineer Location: Arlington, VA. Clearance Required: Secret. Employment...  ...infrastructure supporting the training, validation, deployment, and operation of AI/ML models. Support model-serving platforms and efficient model inference. Design and implement automated... 
    Training
    Full time
    Work at office

    The JAAW Group

    Arlington, VA
    1 day ago
  •  ...Senior Forward Deployed AI Engineer We are seeking a Senior Forward...  ...responsible for transforming prototype models into scalable, efficient, and...  ...across the stack—from model inference on consumer hardware to...  ...recruitment, hiring, training, compensation, promotion, benefits... 
    Training
    Casual work
    Live out
    Work at office
    Local area
    Remote work
    Flexible hours

    webAI

    Washington DC
    3 days ago
  • $229.9k - $262.4k

     ...Overview AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable...  ...to millions of customers.  Our AI models and platforms empower teams across...  ...components including foundation model training, large language model inference,... 
    Training
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    4 days ago
  • Job Title: Applied AI EngineerLocation: 100%...  ...Role As an Applied AI Engineer, you will turn model capabilities into real...  ..., orchestration, infra, UX)Optimize for latency...  ...APIs, LLaMA, Qwen, etc.)Inference / serving (e.g. vLLM)...  ...-on experience with training, fine-tuning, or... 
    Training

    Spectraforce Technologies

    Washington DC
    a month ago
  • $115.7k - $119.5k

     ...ventures—and business purpose. We work in a uniquely collaborative model across the firm and throughout all levels of the client...  ...assets for the business and progressively assist in onboarding and training junior colleagues based on your topic/sector expertise. You'... 
    Training
    Work at office
    Local area

    Boston Consulting Group

    Washington DC
    1 day ago
  • $120k - $130k

     ...world-class management of global logistics, training and procurement services for U.S....  ...Institute (PRI). About this position: AI Engineer Location - Washington, DC The Essential...  ...incorporating workflow orchestration, model and tool integration, state management,... 
    Training
    Full time
    Contract work
    For contractors

    Bering Straits Native Corporation

    Washington DC
    3 days ago
  •  ...AI Engineer Location: McLean, VA Job Type: Full-Time Experience: 5...  ...and optimize machine learning models using Python, PyTorch, TensorFlow...  ...evaluation, fine-tuning, and inference optimization techniques....  ...data ingestion, transformation, training, and model deployment.... 
    Training
    Full time

    AceStack LLC

    McLean, VA
    2 days ago
  •  ...Overview: AI Engineer Position Responsibilities 2-4 DAYS ON SITE AI Engineer...  ...Minimum 3 years of experience in AI/ML model development and deployment. - Experience...  ...initiatives. - Support user adoption through training and documentation. - Support existing... 
    Training

    Guru Schools

    Washington DC
    21 hours ago
  • $197.3k - $225.1k

    AI Engineer 4 At Capital One, we are creating responsible and reliable AI systems,...  ...value to millions of customers. Our AI models and platforms empower teams across...  ...including foundation model training, large language model inference, agents and multi-agent workflows,... 
    Training
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    3 days ago
  •  ...regression testing. Identify and document defects, track their resolution, and participate in defect triage meetings. User Support and Training - Provide user support and troubleshooting assistance for IT systems, responding to inquiries and issues promptly and... 
    Training
    Work at office
    Local area
    Immediate start
    3 days per week

    Saliense Consulting

    Arlington, VA
    1 day ago
  •  ...Standardize processes, draft standard operating procedures (SOPs), and build process workflows. Produce executive briefings, training decks, and multimedia content to help staff adopt knowledge-sharing protocols. Help bridge the gap between users and emerging platforms... 
    Training

    Crowned Grace Inc

    Washington DC
    2 days ago
  • AI & NLP Fellowship: Data Engineering for Social Impact 2 days ago Be among the first 25 applicants Institute...  ...fine-tuning domain-specific language models. This is not a theoretical exercise....  ...contributing to the quality and integrity of training data that powers an open-access AI... 
    Training
    Summer work
    Internship
    Remote work
    Worldwide
    Flexible hours

    Institute for Development Impact - I4DI

    Washington DC
    1 day ago
  • $179.4k - $204.7k

     ...Lead AI Engineer (AI Foundations) Overview At Capital One, we are creating responsible...  ...to millions of customers. Our AI models and platforms empower teams across...  ...including foundation model training, large language model inference, similarity search, guardrails, model... 
    Training
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    4 days ago
  • $197.3k - $225.1k

     ...Lead AI Engineer Overview: At Capital One, we are creating responsible and...  ...value to millions of customers. Our AI models and platforms empower teams across...  ...including foundation model training, large language model inference, similarity search, guardrails, model... 
    Training
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    1 day ago
  •  ...Senior Life Sciences Knowledge Engineer Company: Norstella...  ...domain expertise and applied AI development. This role will be...  ...end-to-end behavior we want a model to internalize. The datasets and...  ...production requirements across NPD. Train and enable subject matter... 
    Training
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Flexible hours
    Shift work

    Norstella

    Washington DC
    21 hours ago
  • $197.3k - $225.1k

     ...Lead AI Engineer (AI Foundations, LLM Core and Agentic AI) Overview At Capital One...  ...value to millions of customers. Our AI models and platforms empower teams across...  ...including foundation model training, large language model inference, similarity search, guardrails, model... 
    Training
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    3 days ago
  • $75 - $80 per hour

     ...AI Solutions Engineer Hybrid - Washington D.C Innodata is a global data engineering company...  ...to turn raw corpora into high-quality, model-ready data. This role suits an engineer...  ...Evaluation design for AI/ML training data: IAA methodology, drift detection... 
    Training
    Hourly pay

    Innodata Inc.

    Washington DC
    1 day ago
  •  ...interview. What we offer A high-performing team of developers, engineers, data scientists, architects, and strategists solving complex,...  ...roles and clients, with extensive onboarding and sponsored training and professional development. And Peregrine has been a Benefit... 
    Training
    Temporary work
    Work at office
    Immediate start

    OpenDataJobs

    Washington DC
    2 days ago
  • $100k

     ...highly motivated, hands-on systems engineer with exceptional technical and...  ...pursuit of the advancement of AI and autonomy to revolutionize...  ...developing simulation models using C/C++/JAVA or other suitable...  ...market conditions, education/training and skill level with consideration... 
    Training
    Temporary work
    For contractors
    Work experience placement
    Relocation package
    Flexible hours

    The Johns Hopkins University Applied Physics Laboratory

    Laurel, MD
    3 days ago
  • $157k - $224k

     ...Lead Model-Based Systems Engineer STR is hiring a Lead Model-Based Systems Engineer in our Woburn, MA office to work across a broad portfolio...  ...but not limited to, the candidate's experience, education, training, key skills/critical skills, security clearances, and prevailing... 
    Training
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Flexible hours
    Night shift

    Science & Technology Research (STR)

    Arlington, VA
    21 hours ago
  •  ...building machine learning models and systems to protect...  ...MoE architecture training and routing optimization...  ...stability3. Context engineering and tool collaboration...  ...represents the frontier of AI today. This topic...  ...for training/finetuning/inference; Being familiar with... 
    Training
    Flexible hours
    Shift work

    TikTok

    Washington DC
    a month ago
  • $184k - $259.44k

     ...Scale AI is seeking a highly skilled and motivated Software Engineer, Frontier AI Infrastructure to join our...  ..., you will own the model inference layer - enabling state...  ...break, moving us from "infra-only debugging" to proactive...  ...relevant education or training. Scale employees in... 
    Training
    Full time
    Work at office
    3 days per week
    Early shift

    Scale AI

    Washington DC
    21 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Engineer, Model Training, Inference & Infra. Be the first to apply!