Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Lab Tech Engineer

Momento USA

AI Lab Tech Engineer

Location: Fremont, CA [Ideally 2 days Hybrid Onsite in Fremont, CA/Remote]
Job Type: Long-Term Contract

Please ensure that the candidates submitted closely match the technical requirements and are available to proceed immediately.

Overview:

We're looking for an Infrastructure Engineer to own the execution layer beneath our RL environments: the systems that let an agent operate inside a realistic, multi-tool world coherently for hours or days.

This is a hard systems problem disguised as an AI job.
As the tasks agents can complete keep lengthening, the environments that train them have to stay coherent across far longer horizons than anything that exists today.
That means sandboxing and isolation you can trust, execution that's fast and cheap enough to run at training scale, and the ability to snapshot, restore, inspect, and branch a running environment instead of treating every rollout as one-shot. You'll build the platform that makes all of this possible.

You'll work closely with our research and data teams, and directly with frontier labs and enterprise customers, to turn environment designs into infrastructure that runs reliably in production.

What You'll Do

  1. Environment Execution & Sandboxing:
    • Design and own the sandboxing and execution layer that environments run inside. Build systems to snapshot and restore environment state (disk, process, and where relevant memory and accelerator state) so runs can be paused, resumed, inspected, and branched rather than executed once.
    • Develop the machinery to detect failure modes early in a rollout (reward hacks, infra faults, fairness issues) and to revert to a known-good state, patch, and continue.
    • Extend execution to long-horizon and multi-node environments, where an agent operates across many tools and services over hours or days.
  2. Performance & Scale
    • Own the performance characteristics of the platform: throughput, latency, and cost-per-rollout at scale.
    • Drive utilization and scheduling so we can run far more environment rollouts per dollar without sacrificing reliability.
    • Profile and remove bottlenecks across the stack, from container startup to environment teardown.
    • Build the observability that lets us understand what's happening inside thousands of concurrent, long-running rollouts.
  3. Environment Platform
    • Build and maintain the framework for specifying, packaging, and deploying RL environments which is used by both humans and agents authoring environments internally.
    • Create the tooling that lets researchers and environment authors debug a specific failure across hundreds of long agent traces.
    • Deploy large / small models on on-prem hardware
  4. Collaboration & Production Excellence
    • Scale prototypes into production systems with reproducible workflows and high engineering standards.
    • Write the documentation and tools that let internal teams and external users build on the platform.

What We're Looking For

  1. Systems & Infrastructure
    • Strong track record building production systems or research infrastructure at scale: distributed systems, execution engines, container/sandboxing infrastructure, or similar.
    • Deep comfort with the systems layer: containers and isolation (e.g. namespaces, cgroups, VMs, gVisor/Firecracker-style sandboxing), filesystems, process and state management.
    • Experience making systems fast and cheap - profiling, scheduling, resource utilization, and cost optimization at scale.
    • Proficiency with cloud platforms (GCP, AWS) and distributed computing.
    • Strong engineering fundamentals and a systematic approach to testing, validation, and reliability.
    • Experience of deploying large / small models on on-prem hardware
  2. Execution & Ownership
    • Comfort operating in ambiguity.
    • Strong Python skills; comfort in a systems language (Rust, Go, or C++) is a plus.
    • Ability to use modern tools such as Claude Code effectively.
  3. Collaboration & Communication
    • Excellent communication skills for working with research teams and enterprise customers.
    • Ability to translate between research needs and infrastructure requirements.
    • Comfortable presenting technical work to diverse audiences.

Thanks,

Amaresh

Momento USA | Exceeding Customer Expectations

440 Benigno Blvd, Unit#A 2nd Floor. Bellmawr, NJ 08031

Interstate Business Park

Direct : View phone number on us.fitly.work Ext (1022)

Email: View email address on us.fitly.work

Are You LinkedIn? : linkedin.com/in/gandi-amareshwar-74a710144

Minority Certified by SWAM
National Minority Certified by NMSDC

One of the fastest growing company in NJ
Awarded fastest growing Asian American business by Diversitybusiness.com
E-verified Company

Information transmitted by this e-mail is proprietary to Momento USA and/ or its Customers and is intended for use only by the individual or entity to which it is addressed, and may contain information that is privileged, confidential or exempt from disclosure under applicable law. If you are not the intended recipient or it appears that this mail has been forwarded to you without proper authority, you are notified that any use or dissemination of this information in any manner is strictly prohibited. In such cases, please notify us immediately at View email address on us.fitly.work and delete this mail from your records.

Note: Momento USA is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status, or disability status.

Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Lab Tech Engineer. Be the first to apply!