Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer

$100k - $160k

Bright Vision Technologies

AI Infrastructure Engineer

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: AI Infrastructure Engineer

Location: 100% Remote (U.S.)

Position Type: Full-time, Direct W2

Salary Range: $100,000–$160,000 Annually

Experience Required: 10+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

We are seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control.

Key Responsibilities
  • Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations.
  • Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.
  • Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering.
  • Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate.
  • Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
  • Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics.
  • Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale.
  • Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing.
  • Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently.
  • Partner with research and applied ML teams to plan capacity for upcoming training runs.
  • Implement security controls, isolation, and access management for multi-tenant AI infrastructure.
  • Drive automation across cluster provisioning, lifecycle management, and configuration enforcement.
  • Maintain runbooks, capacity dashboards, and operational documentation for the AI platform.
  • Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science or a related field.
  • Ten or more years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
  • Strong understanding of Linux internals, networking, and high-performance storage.
  • Experience with at least one major cloud provider's ML infrastructure offerings.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.
Preferred Qualifications
  • Experience operating InfiniBand or RDMA networking at scale.
  • Contributions to open-source ML infrastructure projects.
  • Familiarity with custom orchestrators or research-grade training stacks.
  • Exposure to frontier model training operations.
  • Experience with FinOps for AI workloads.

How to Apply

Would you like to know more about this opportunity? For immediate consideration, please send your resume to View email address on click.appcast.io.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in United States vacancy
  •  ...AI Infrastructure Engineer Redstone Arsenal/Huntsville, AL At IPT Associates (IPTA), we enjoy solving real-world problems with technology. We work closely with our customers, teammates, experts, and partners to come up with practical solutions that actually make... 
    Suggested
    Full time

    Interactive Process Technology Llc

    Remote
    1 day ago
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered...  ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Dallas, TX
    6 days ago
  • $154.39k - $247.02k

     ...ownership and drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter.AI Infrastructure Engineer, Corporate AI TeamTeam & Role OverviewAxon’s Corporate AI Team sits within Business Technology and builds internal-facing... 
    Suggested
    Work experience placement

    Axon

    Boston, MA
    2 days ago
  • $170.5k - $315.49k

    Job Details:Job Description: We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures.In this role, you will dive deep into the inference stack and redefine peak performance... 
    Suggested
    Full time
    Local area
    Immediate start
    Shift work

    Intel

    Hillsboro, OR
    4 days ago
  • $184k - $287.5k

    AI Infrastructure Engineers at NVIDIA build the systems, tooling, and data infrastructure that enable operation of our GPU cloud services. We are enabling engineering teams to innovate while proactively identifying, tracking, and mitigating risks across the entire technical... 
    Suggested
    Full time
    Remote work

    Nvidia

    Westford, MA
    5 days ago
  • $150k - $170k

     ...affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure... 
    Work experience placement
    Local area
    Remote work
    Worldwide

    Analysis Group

    Boston, MA
    4 days ago
  • $215k - $350k

    We are seeking an AI Infrastructure Engineer to build, operate, and continuously enhance the Linux and GPU-based infrastructure that powers our AI platforms and performance testing environments. This is a highly hands-on infrastructure, automation, and performance engineering... 
    Worldwide
    Home office

    Fortinet

    New York, NY
    3 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded...  ...THE ROLE: We are seeking a DevOps / Platform Engineer to join our team building and operating large-scale GPU compute infrastructure that powers AI and ML workloads. THE PERSON:... 

    AMD

    San Jose, CA
    6 days ago
  • $151.8k - $332.2k

    What you can expectWe are seeking an experienced AI Infrastructure Engineer to join our AI Incubation team. You will be focused on building and optimizing large-scale training infrastructure for Large Language Models (LLMs). The ideal candidate will combine engineering... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    5 days ago
  •  ...AI Infrastructure Engineer This role focuses on managing and optimizing our AI infrastructure, ensuring seamless operations, and providing guidance and training to our team members. The ideal candidate will have hands-on experience with AI operations, infrastructure... 
    Remote work

    United IT

    United States
    3 days ago
  •  ...AI Engineer The AI Engineer will design, develop, and deploy scalable machine learning and AI-driven analytics capabilities. Responsibilities include multi-source data fusion, entity resolution and behavioral modeling, predictive and prescriptive intelligence analytics... 
    Remote work

    10x National Security

    United States
    8 hours ago
  • $191k - $315k

     ...Overview:  The Network Growth and Relationship AI team is at the forefront of creating...  ...close collaboration with the product, engineering and data science team and has a very...  ...Prior experience with large scale ML data infrastructure ~ Experience with developing and designing... 
    Full time
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Remote
    1 day ago
  •  ...The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to automate the lifecycle of our AI workflow, including model development, evaluation, optimization, deployment, and monitoring across thousands of video... 
    Full time

    Zensors

    San Francisco, CA
    1 day ago
  • $105k - $115k

     ...Description Cloud AI Engineer (Mid-Level) Location: Various U.S. Federal Client Sites (Onsite) Travel: Travel required...  ...trusted provider of enterprise cloud modernization, AI, and infrastructure solutions supporting U.S. Federal Government agencies. As we... 
    Full time
    Temporary work
    Immediate start

    Pgtek

    New York, NY
    1 day ago
  • ​GTSC seeks a   Cloud Architect Artificial Intelligence (AI) Engineer , Level IV, to support our customer in the   Annapolis Junction, MD   area. Location:   Annapolis Junction, Maryland All work is on-site. This is not a hybrid or remote position.   Mission... 
    Full time
    Temporary work

    Gtsc-Talent Solutions

    Remote
    1 day ago
  •  ...Strategic Innovation Group (SIG) is seeking a Cloud/AI Engineer to design, develop, and deploy secure, scalable cloud-native and artificial...  ...Expertise implementing CI/CD pipelines, DevOps practices, infrastructure automation (e.g., Terraform, CloudFormation), and platform... 
    Full time
    Temporary work
    For contractors
    Work at office
    Remote work

    Strategic Innovation Group, Llc

    Arlington, TX
    1 day ago
  • $175k - $275k

     ...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in...  ...would be possible. The Role As an AI Engineer - Cloud Infrastructure on Traversal’s Infrastructure team, you’ll design,... 
    Full time
    Work at office
    Flexible hours

    Traversal

    New York, NY
    1 day ago
  •  ...are poised to disrupt the billion-dollar engineering simulation industry with our fast-...  ...You will help develop and improve core infrastructure and user-facing capabilities across our...  ...generation of engineering software and AI-enabled workflows. What You’ll Work... 
    Full time

    Flexcompute

    Remote
    1 day ago
  • At the National Robotics Engineering Center (NREC), it is our engineers and technicians who drive the breakthroughs that define our...  ...career development.We are seeking a dynamic Senior Applied AI Infrastructure Engineer to lead and contribute to the evaluation, deployment... 
    Full time
    Work experience placement
    Flexible hours

    Carnegie Mellon University

    Pittsburgh, PA
    4 days ago
  • $190k - $260k

     ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...large-scale world models - depends on infrastructure that turns thousands of hours of multimodal...  ...training throughput. We are looking for engineers who make model training fast: streaming... 
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    4 days ago
  •  ...Job Description Job Description We are hiring Sr. AI Infrastructure Engineer- Hybrid for a Full Time position in costa mesa, CA Senior AI Infrastructure Engineer, Physical Infrastructure Costa Mesa, California, United States ABOUT THE TEAM CorpTech Infrastructure... 
    Full time

    Calance US

    Costa Mesa, CA
    3 days ago
  • $220k - $292k

     ...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands...  ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems...  ...We are looking for a founding Staff AI Infrastructure Engineer to architect, build, and scale... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    3 days ago
  •  ...AI Infrastructure EngineerWe're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure. You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.Essential Duties... 
    Work at office
    Immediate start
    Relocation package

    Netpreme

    Boston, MA
    2 days ago
  • Job Description: RADIOLOGY PARTNERS OVERVIEW Radiology Partners, through its owned and affiliated practices, is a leading radiology practice in the U.S., serving hospitals and other healthcare facilities across the nation. As a physician-led and physician-owned...
    Remote work

    Radiology Partners

    United States
    8 hours ago
  •  ...A tech company specializing in AI infrastructure is seeking a Software Engineer to build a scalable compute platform for its generative video models. The ideal candidate will have over 5 years of experience in MLOps or AI infrastructure management, along with strong Python... 

    HeyGen

    Los Angeles, CA
    8 hours ago
  •  ...AI Infrastructure Engineer (Forward Deployed AI Engineer)Location: Fort Mill, SC (Onsite – 5 Days/Week)Tax Term (W2, C2C): W2 / C2CJob Type (Permanent/Contract): PermanentDuration: Full-TimeDescriptionWe are building a Forward Deployed AI Engineering team to support the... 
    Permanent employment
    Contract work

    Apolis

    Fort Mill, York County, SC
    8 hours ago
  • $170k - $210k

     ...AI Infrastructure Engineer Utilidata is a fast-growing NVIDIA-backed AI company enabling AI data centers to dynamically orchestrate power and unlock more compute capacity from existing energy infrastructure. For over a decade, we have applied AI to the electric grid... 
    Local area
    Remote work
    Flexible hours

    Utilidata

    United States
    4 days ago
  •  ...AI Infrastructure Engineer Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source... 
    Casual work
    Remote work

    Mirantis

    United States
    1 day ago
  •  ...Senior AI Infrastructure EngineerAustin, Texas, United States; Reston, Virginia, United StatesSeekr is building the infrastructure that powers...  ...generation of enterprise AI. As a Senior AI Infrastructure Engineer, you will design, build, and operate the platforms that... 
    Permanent employment
    Flexible hours

    Seekr

    Austin, TX
    4 days ago
  •  ...What is an AI Infrastructure Engineer? An AI infrastructure engineer builds and manages the technology systems that support artificial intelligence applications. Rather than creating AI models themselves, they focus on the computing power, storage, networks, and cloud... 
    Remote work

    Valence

    United States
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!