AI Infrastructure Engineer
$100k - $160kBright Vision Technologies
AI Infrastructure Engineer
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: AI Infrastructure Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$160,000 Annually
Experience Required: 10+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control.
Key Responsibilities
- Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations.
- Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.
- Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering.
- Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate.
- Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
- Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics.
- Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale.
- Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing.
- Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently.
- Partner with research and applied ML teams to plan capacity for upcoming training runs.
- Implement security controls, isolation, and access management for multi-tenant AI infrastructure.
- Drive automation across cluster provisioning, lifecycle management, and configuration enforcement.
- Maintain runbooks, capacity dashboards, and operational documentation for the AI platform.
- Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling.
Required Qualifications
- Bachelor's or Master's degree in Computer Science or a related field.
- Ten or more years of experience in infrastructure, platform, or HPC engineering.
- Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
- Strong proficiency in Python and at least one systems language such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, and collective communication.
- Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
- Strong understanding of Linux internals, networking, and high-performance storage.
- Experience with at least one major cloud provider's ML infrastructure offerings.
- Strong software engineering practices including testing, CI/CD, and code review.
- Excellent communication and cross-functional collaboration skills.
Preferred Qualifications
- Experience operating InfiniBand or RDMA networking at scale.
- Contributions to open-source ML infrastructure projects.
- Familiarity with custom orchestrators or research-grade training stacks.
- Exposure to frontier model training operations.
- Experience with FinOps for AI workloads.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to View email address on click.appcast.io.
- ...AI Infrastructure Engineer Redstone Arsenal/Huntsville, AL At IPT Associates (IPTA), we enjoy solving real-world problems with technology. We work closely with our customers, teammates, experts, and partners to come up with practical solutions that actually make...SuggestedFull time
- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered... ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),...SuggestedFull timeWork experience placementLive inWork at officeLocal area
$154.39k - $247.02k
...ownership and drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter.AI Infrastructure Engineer, Corporate AI TeamTeam & Role OverviewAxon’s Corporate AI Team sits within Business Technology and builds internal-facing...SuggestedWork experience placement$170.5k - $315.49k
Job Details:Job Description: We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures.In this role, you will dive deep into the inference stack and redefine peak performance...SuggestedFull timeLocal areaImmediate startShift work$184k - $287.5k
AI Infrastructure Engineers at NVIDIA build the systems, tooling, and data infrastructure that enable operation of our GPU cloud services. We are enabling engineering teams to innovate while proactively identifying, tracking, and mitigating risks across the entire technical...SuggestedFull timeRemote work$150k - $170k
...affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure...Work experience placementLocal areaRemote workWorldwide$215k - $350k
We are seeking an AI Infrastructure Engineer to build, operate, and continuously enhance the Linux and GPU-based infrastructure that powers our AI platforms and performance testing environments. This is a highly hands-on infrastructure, automation, and performance engineering...WorldwideHome office- ...next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded... ...THE ROLE: We are seeking a DevOps / Platform Engineer to join our team building and operating large-scale GPU compute infrastructure that powers AI and ML workloads. THE PERSON:...
$151.8k - $332.2k
What you can expectWe are seeking an experienced AI Infrastructure Engineer to join our AI Incubation team. You will be focused on building and optimizing large-scale training infrastructure for Large Language Models (LLMs). The ideal candidate will combine engineering...Full timeWork at officeRemote work- ...AI Infrastructure Engineer This role focuses on managing and optimizing our AI infrastructure, ensuring seamless operations, and providing guidance and training to our team members. The ideal candidate will have hands-on experience with AI operations, infrastructure...Remote work
- ...AI Engineer The AI Engineer will design, develop, and deploy scalable machine learning and AI-driven analytics capabilities. Responsibilities include multi-source data fusion, entity resolution and behavioral modeling, predictive and prescriptive intelligence analytics...Remote work
$191k - $315k
...Overview: The Network Growth and Relationship AI team is at the forefront of creating... ...close collaboration with the product, engineering and data science team and has a very... ...Prior experience with large scale ML data infrastructure ~ Experience with developing and designing...Full timeFor contractorsWork at officeFlexible hours- ...The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to automate the lifecycle of our AI workflow, including model development, evaluation, optimization, deployment, and monitoring across thousands of video...Full time
$105k - $115k
...Description Cloud AI Engineer (Mid-Level) Location: Various U.S. Federal Client Sites (Onsite) Travel: Travel required... ...trusted provider of enterprise cloud modernization, AI, and infrastructure solutions supporting U.S. Federal Government agencies. As we...Full timeTemporary workImmediate start- GTSC seeks a Cloud Architect Artificial Intelligence (AI) Engineer , Level IV, to support our customer in the Annapolis Junction, MD area. Location: Annapolis Junction, Maryland All work is on-site. This is not a hybrid or remote position. Mission...Full timeTemporary work
- ...Strategic Innovation Group (SIG) is seeking a Cloud/AI Engineer to design, develop, and deploy secure, scalable cloud-native and artificial... ...Expertise implementing CI/CD pipelines, DevOps practices, infrastructure automation (e.g., Terraform, CloudFormation), and platform...Full timeTemporary workFor contractorsWork at officeRemote work
$175k - $275k
...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in... ...would be possible. The Role As an AI Engineer - Cloud Infrastructure on Traversal’s Infrastructure team, you’ll design,...Full timeWork at officeFlexible hours- ...are poised to disrupt the billion-dollar engineering simulation industry with our fast-... ...You will help develop and improve core infrastructure and user-facing capabilities across our... ...generation of engineering software and AI-enabled workflows. What You’ll Work...Full time
- At the National Robotics Engineering Center (NREC), it is our engineers and technicians who drive the breakthroughs that define our... ...career development.We are seeking a dynamic Senior Applied AI Infrastructure Engineer to lead and contribute to the evaluation, deployment...Full timeWork experience placementFlexible hours
$190k - $260k
...has developed an artificial intelligence (AI) powered technology stack purpose-built... ...large-scale world models - depends on infrastructure that turns thousands of hours of multimodal... ...training throughput. We are looking for engineers who make model training fast: streaming...Temporary workWork at officeVisa sponsorship- ...Job Description Job Description We are hiring Sr. AI Infrastructure Engineer- Hybrid for a Full Time position in costa mesa, CA Senior AI Infrastructure Engineer, Physical Infrastructure Costa Mesa, California, United States ABOUT THE TEAM CorpTech Infrastructure...Full time
$220k - $292k
...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands... ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems... ...We are looking for a founding Staff AI Infrastructure Engineer to architect, build, and scale...Full timeWork experience placementImmediate start- ...AI Infrastructure EngineerWe're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure. You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.Essential Duties...Work at officeImmediate startRelocation package
- Job Description: RADIOLOGY PARTNERS OVERVIEW Radiology Partners, through its owned and affiliated practices, is a leading radiology practice in the U.S., serving hospitals and other healthcare facilities across the nation. As a physician-led and physician-owned...Remote work
- ...A tech company specializing in AI infrastructure is seeking a Software Engineer to build a scalable compute platform for its generative video models. The ideal candidate will have over 5 years of experience in MLOps or AI infrastructure management, along with strong Python...
- ...AI Infrastructure Engineer (Forward Deployed AI Engineer)Location: Fort Mill, SC (Onsite – 5 Days/Week)Tax Term (W2, C2C): W2 / C2CJob Type (Permanent/Contract): PermanentDuration: Full-TimeDescriptionWe are building a Forward Deployed AI Engineering team to support the...Permanent employmentContract work
$170k - $210k
...AI Infrastructure Engineer Utilidata is a fast-growing NVIDIA-backed AI company enabling AI data centers to dynamically orchestrate power and unlock more compute capacity from existing energy infrastructure. For over a decade, we have applied AI to the electric grid...Local areaRemote workFlexible hours- ...AI Infrastructure Engineer Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source...Casual workRemote work
- ...Senior AI Infrastructure EngineerAustin, Texas, United States; Reston, Virginia, United StatesSeekr is building the infrastructure that powers... ...generation of enterprise AI. As a Senior AI Infrastructure Engineer, you will design, build, and operate the platforms that...Permanent employmentFlexible hours
- ...What is an AI Infrastructure Engineer? An AI infrastructure engineer builds and manages the technology systems that support artificial intelligence applications. Rather than creating AI models themselves, they focus on the computing power, storage, networks, and cloud...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!
- ai research engineer United States
- senior ai engineer United States
- ai ml engineer United States
- ai engineer remote United States
- ai developer United States
- ai prompt engineer United States
- machine learning ai engineer United States
- ai engineer United States
- infrastructure developer United States
- principal infrastructure engineer United States



