AI Engineer, Model Training, Inference & Infra
International Recruiting LLC
Job Description
Job Description
Full-time · On-site · San Jose, CA · Austin, TX or Taiwan
About Our ClientsOur client is building the next generation of design automation for the semiconductor industry.
Our mission is to enable every engineering organization to build its own self-improving agentic design workforce. It combines AI agents, engineering knowledge, agent-native tools, advanced models, and continuous learning to automate complex chip-design workflows.
The team brings deep experience in artificial intelligence, electronic design automation, semiconductor design, GPU-accelerated computing, and production software systems. We work closely with leading semiconductor companies to turn advanced research into technology that improves engineering productivity, design quality, and time to market.
The RoleWe are looking for an exceptional AI Engineer to own the model training, inference, and infrastructure that power its agentic design workforce.
You will drive the full model lifecycle: data pipelines, pretraining and post-training, reinforcement learning, evaluation, and high-performance serving. You will build and operate the GPU training and serving stack — on the compute substrate our infrastructure team provides — that keeps large-scale training reliable and low-latency inference efficient at production scale. You will work on evaluation systems, model training, and the self-improving, self-evolving learning loops that let our models and agents get better over time from real execution feedback.
Your work directly determines how capable, fast, and cost-effective our agents are. This role is ideal for someone who combines strong research ability with exceptional systems and performance engineering skills, and who wants the models they train and serve deployed in real semiconductor design environments—not left in notebooks or benchmarks.
What You'll DoTrain, post-train, and fine-tune large language models for agentic engineering workflows, including supervised fine-tuning, RLHF/RLAIF, reinforcement learning, and distillation.
Build scalable data pipelines for pretraining, post-training, and evaluation, including sparse, private, and domain-specific engineering data.
Design and operate distributed training on multi-node GPU clusters, using data, tensor, pipeline, and sequence parallelism (for example FSDP, DeepSpeed, or Megatron-style approaches).
Build high-throughput, low-latency inference systems with continuous batching, KV-cache management, paged attention, quantization, and speculative decoding.
Write and optimize custom GPU kernels (CUDA, Triton) and profile end-to-end performance across CPUs and GPUs.
Build core model infrastructure: cluster orchestration, ML job scheduling, checkpointing, fault tolerance, reproducibility, model and environment management, observability, and cost and utilization tracking.
Build automated evaluation systems, benchmarks, and reward models that measure agent capability, reliability, and regression across complex engineering tasks, including problems where design data is private or customer-specific.
Design self-improving and self-evolving algorithms and learning loops, where models and agents learn from execution feedback, outcomes, and new data to improve continuously over time.
Integrate models with agent runtimes, tool use, retrieval, and the production serving stack.
Improve reliability, throughput, and cost efficiency across the training and inference platform.
Translate promising research ideas into reliable, scalable product capabilities.
Collaborate with research, product, platform, and solutions teams across San Jose, Austin, and Taiwan.
Contribute to patents, publications, technical presentations, and the broader development of Agentic Design Automation.
PhD or master's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent practical experience.
Strong programming skills in Python and proficiency in at least one systems language such as C++ or Rust.
Deep experience with machine learning frameworks such as PyTorch or JAX.
Hands-on experience with one or more of the following:
Large-scale or distributed model training
High-performance model inference and serving
GPU programming and performance optimization
ML infrastructure and platform engineering
Automated evaluation, reward modeling, or self-improving and continuous-learning systems
Ability to take a model from data and problem formulation through training, evaluation, and production deployment.
Strong analytical, software engineering, performance-optimization, and debugging skills.
High ownership, intellectual curiosity, and willingness to work across research and product boundaries.
Clear written and verbal communication skills.
Pretraining or post-training large language models at scale.
Distributed training frameworks such as Megatron-LM, DeepSpeed, FSDP, or Ray.
Production inference engines such as vLLM, TensorRT-LLM, SGLang, or TGI.
Custom kernel development with CUDA, Triton, or CUTLASS.
Inference optimization techniques such as quantization (FP8, GPTQ, AWQ), speculative decoding, or KV-cache optimization.
Reinforcement learning, RLHF, or reward-model training for LLMs.
Automated evaluation, benchmarking, or LLM-as-judge systems for agents.
Self-improving, self-evolving, or continuous-learning systems, including learning from execution feedback, automated curricula, or synthetic data generation.
GPU cluster infrastructure with Kubernetes, Slurm, or Ray, and high-performance networking such as NCCL or InfiniBand.
Data pipelines and MLOps for training and continuous learning.
Experience deploying AI systems in enterprise or security-sensitive environments.
A strong record of implementation through research systems, open-source projects, production software, or technical competitions.
At this company, you will have the opportunity to:
Help define a new category of semiconductor design technology.
Build the training and inference stack that powers autonomous engineering agents.
Develop GPU-accelerated systems that make large-scale training and low-latency serving practical and cost-effective.
Build the evaluation and self-improvement loops that let agents learn and get better from real engineering work.
Build AI systems that perform complex, consequential engineering work—not just generate recommendations.
Work with real semiconductor workflows, tools, and private engineering knowledge.
See your models deployed directly with leading chip-design organizations.
Work in a small, highly technical team where individual contributions can shape the product and company.
Collaborate with colleagues across San Jose, Austin, and Taiwan.
Change how chips are designed, rather than focus on only one design or one point tool.
Our client is an equal opportunity employer. We welcome candidates from diverse backgrounds who are excited to combine ambitious research with meaningful engineering impact.
$229.9k - $262.4k
...AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems... ...value to millions of customers. Our AI models and platforms empower teams across... ...including foundation model training, large language model inference, agents...TrainingFull timePart timeLocal area$229.9k - $262.4k
## Senior Lead AI Engineer (FM Hosting, LLM Inference)Applylocations: McLean, VA: New York, NY: San Jose, CAtime... ...to millions of customers. Our AI models and platforms empower teams across... ...components including foundation model training, large language model inference,...TrainingFull timePart timeLocal area$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and... ...to millions of customers. Our AI models and platforms empower teams across... ...components including foundation model training, large language model inference,...TrainingFull timePart timeLocal area- ...Capital One is seeking an AI Engineer 4 to help build responsible AI systems and scalable AI infrastructure. You will design and deploy AI components, training pipelines, inference engines, and governance mechanisms while partnering with cross-functional teams across...Training
- ...Bitwise in Laurel, MD is seeking a Senior Software Engineer – Inference to lead evaluation, configuration, and deployment of inference models, ensuring access to high-quality LLMs within our AI infrastructure. You will mentor engineers, shape best practices, and collaborate...Suggested
- ...performance Health insurance Training & development About the... ...a talented and passionate AI Engineer to join our growing team in DC... ...machine learning and AI models into production environments... ...model training, evaluation, and inference Research and implement state...Training
- ...AI Engineer Location: Arlington, VA. Clearance Required: Secret. Employment... ...infrastructure supporting the training, validation, deployment, and operation of AI/ML models. Support model-serving platforms and efficient model inference. Design and implement...TrainingFull timeWork at office
$160k - $190k
...Economic Consulting is seeking an AI Engineer to design, develop, and... ...emerging AI systems, including model selection, prompt design,... ...stores, and tradeoffs among inference, retrieval, and fine-tuning to... ...program. Career support: training programs, an assigned mentor...TrainingZero hours contractWork at office3 days per week- ...adversaries, we're hiring an AI engineer that lives and breathes the frontier... ...beyond what off-the-shelf models can robustly accomplish. You'... ...the cutting-edge in post-training to take them in directions that... ..., GPU memory management, inference optimization. Bonus Points...TrainingPermanent employmentFlexible hours
- ...what is delivered. Our AI-native platform, Air Enterprise... ...experienced Senior AI Engineer to join our Agentic AI... ...agent architectures, model integrations, tools,... .... Develop model and inference infrastructure... ...with fine-tuning, post-training, reinforcement learning...TrainingFull timeWork at officeRemote work
- ...Senior Forward Deployed AI EngineerWe are seeking... ...Senior Forward Deployed AI Engineer to support our Public... ...transforming prototype models into scalable,... ...across the stack—from model inference on consumer hardware to... ...including recruitment, hiring, training, compensation,...TrainingCasual workLive outWork at officeLocal areaRemote work
$282k - $332k
...Bitwise hires talented engineers who are driven by purpose and... ...building the next generation of AI infrastructure to power... ...highest quality large language models across our inference software stack. As a... ...annual education assistance for training, certifications, tuition,...TrainingExtra incomeContract workTemporary workFlexible hoursShift workNight shift$197.5k - $232.5k
...Bitwise hires talented engineers who are driven by purpose... ...the next generation of AI infrastructure that... ...to users throughout the inference software stack. Requirements... ...and test new inference models, preparing them for... ...assistance for training, certificates, tuition,...TrainingExtra incomeContract workTemporary workFlexible hoursShift workNight shift$165k - $180k
...seeking a highly skilled and motivated Sr. AI Data Engineer with a proven track record in... ...data pipelines, automated testing, and model deployment strategies. Open-Source Integration... ...high-quality datasets for model training and inference. Required Qualifications: BA or BS...TrainingWork experience placementH1bWork at officeLocal area$75 - $80 per hour
...Matlen Silver Job Title: Senior AI/LLM Engineer Duration: 12+ Months... ...strong focus on Large Language Models (LLMs). The ideal candidate... ...databases, and distributed training. Proven experience deploying... .../NLP. Experience optimizing inference on GPUs, TPUs, or other accelerators...TrainingFull time$197.5k - $232.5k
...Bitwise hires talented engineers who are driven by purpose... ...the next generation of AI infrastructure that... ...to users throughout the inference software stack. Requirements... ...and test new inference models, preparing them for... ...assistance for training, certificates, tuition,...TrainingExtra incomeContract workTemporary workFlexible hoursShift workNight shift- ...Central Strategies is seeking an AI/ML Data Engineer to support the U.S. Coast Guard. This... ...warehouses. Develop scalable data models and medallion architectures that... ...analytics, feature engineering, model training, and model inference. Optimize Spark workloads, data...Training
$150k - $200k
...We’re looking for an Applied AI / Machine Learning Engineer to design, build, and deploy practical... ...data Build and evaluate models for tasks such as classification,... ...engineering workflows to support model training and inference Evaluate model performance, bias...TrainingContract workFor contractorsWork at officeRemote workFlexible hours$135k - $150k
...Generative AI Application Engineer Location: United States (Remote) Employment... ...SaaS Reports To: AI/Modeling Team Leadership About... ...and with internally hosted inference engines, including configuration... ...feature engineering, model training/evaluation using Pandas,...TrainingFull timeTemporary workImmediate startRemote workFlexible hoursWeekend work$135.2k - $306.4k
...Masters, or Ph.D. in Computer Science, AI/ML, Engineering, or a related discipline, or... ...includes optimizing large-scale GPU inference or training workloads for latency, throughput, utilization... ...includes building or operating model serving, inference gateways, agent runtimes...TrainingFull timeFlexible hours- ...Capital One is seeking an AI Engineer 5 to help build and scale responsible AI systems. You will collaborate with cross-functional teams on foundation models, LLM inference, and multi-model orchestration within a scalable AI infrastructure. The role emphasizes performance...
$110.7k - $218.3k
Position Summary Our Deloitte AI & Engineering team works to transform technology platforms, drive innovation, and help make a significant... ...and client use casesDevelop and maintain data pipelines, model training workflows, and production-grade application components that...TrainingLocal area$99k - $225k
Agentic AI Forward-Deployed EngineerThe Opportunity:As an AI-forward engineer, you know that intelligence has become a commodity.... ...has access to the same frontier models, and the advantage now belongs... ...trust, managing the change, training the workforce, and leaving each...TrainingFull timeContract workPart timeLocal areaRemote work$114.6k - $252.1k
Job Title: Senior AI EngineerJob Category: Information TechnologyTime... ...an accomplished Senior AI Engineer to support the Department of... ...implementation of OIGChat—a large language model-based assistant—build AI... ..., troubleshooting guides, and training content to support user...TrainingContract workWork experience placementWork at officeFlexible hours$229.9k - $262.4k
...AI Engineer 5 (AI Foundations) At Capital One, we are creating responsible and... ...value to millions of customers. Our AI models and platforms empower teams across... ...including foundation model training, large language model inference, agents and multi-agent workflows,...TrainingFull timePart time$61k - $101k
...Requirements: We require formal training or certification in software engineering concepts, along with 5+... ...of cloud delivery models such as IaaS, PaaS, and... ...architecture, training, and inference. We require experience... ...of enterprise-approved AI-assisted software development...TrainingFull timeFor contractors$229.9k - $262.4k
...AI Engineer 5 (AI Foundations, VLM Customization) At Capital One, we are creating... ...to millions of customers. Our AI models and platforms empower teams across... ...components including foundation model training, large language model inference, agents and multi‑agent workflows,...TrainingFull timePart timeLocal area$197.3k - $225.1k
...AI Engineer 4 Team Description: At Capital One, we are creating responsible and reliable AI systems, changing... ...support AI software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search...TrainingFull timePart timeLocal area$120k - $130k
...world-class management of global logistics, training and procurement services for U.S.... ...Institute (PRI). About this position: AI Engineer Location – Washington, DC The... ...systems incorporating workflow orchestration, model and tool integration, state management,...TrainingFull timeContract workFor contractors$102k - $170k
...SecretWhat You Will Do:Implement AI strategy to increase adoption... ...data scientists, platform engineers, and product teams to iterate... ...caching, context window management, model routing).What Would Be Nice To... ...to skill sets, experience and training, security clearances,...TrainingFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Engineer, Model Training, Inference & Infra. Be the first to apply!




