Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer

$100k - $160k
Full-time

Bright Vision Technologies

AI Infrastructure Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: AI Infrastructure Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 Salary Range: $100,000–$160,000 Annually Experience Required: 10+ years Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. Job Summary: We are seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control. The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work. Key Responsibilities Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations. Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams. Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering. Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate. Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication. Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics. Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale. Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing. Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently. Partner with research and applied ML teams to plan capacity for upcoming training runs. Implement security controls, isolation, and access management for multi-tenant AI infrastructure. Drive automation across cluster provisioning, lifecycle management, and configuration enforcement. Maintain runbooks, capacity dashboards, and operational documentation for the AI platform. Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling. Required Qualifications Bachelor’s or Master’s degree in Computer Science or a related field. Ten or more years of experience in infrastructure, platform, or HPC engineering. Hands-on experience operating GPU clusters or large-scale ML training infrastructure. Strong proficiency in Python and at least one systems language such as Go or C++. Deep understanding of distributed training, accelerator architectures, and collective communication. Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads. Strong understanding of Linux internals, networking, and high-performance storage. Experience with at least one major cloud provider’s ML infrastructure offerings. Strong software engineering practices including testing, CI/CD, and code review. Excellent communication and cross-functional collaboration skills. Preferred Qualifications Experience operating InfiniBand or RDMA networking at scale. Contributions to open-source ML infrastructure projects. Familiarity with custom orchestrators or research-grade training stacks. Exposure to frontier model training operations. Experience with FinOps for AI workloads. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to View email address on click.appcast.io or contact us at View phone number on click.appcast.io. Learn more about Bright Vision Technologies at Bright Vision Technologies is an Equal Opportunity Employer. Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in United States vacancy
  • $151.8k

    What you can expect We are seeking an experienced AI Infrastructure Engineer to join our AI Incubation team. You will be focused on building and optimizing large-scale training infrastructure for Large Language Models (LLMs). The ideal candidate will combine engineering... 
    Suggested
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    1 day ago
  • Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform...  ...direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering... 
    Suggested
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    1 day ago
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered...  ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Atlanta, GA
    5 days ago
  • $200k - $322k

     ...recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a...  ...choice to join us today.Design-for-X Engineering at NVIDIA works on groundbreaking innovations...  ...deployment cycles as part of the AI Infrastructure requirements at an org-wide level.For... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    21 hours ago
  • $150k - $170k

     ...affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure... 
    Suggested
    Work experience placement
    Local area
    Remote work
    Worldwide

    Analysis Group

    Boston, MA
    2 days ago
  • AI Infrastructure EngineerAt BNY, our culture allows us to run our company better and enables employees’ growth and success. As a leading...  ...seeking a future team member for the role of AI Infrastructure Engineer to join our Technology team. This role is located in Lake... 
    Work experience placement
    Worldwide
    Flexible hours

    The Bank of New York Mellon

    Lake Mary, FL
    5 days ago
  • $144k - $198k

     ...globally, ADI ensures today's innovators stay Ahead of What's Possible. Learn more at and on LinkedIn and Twitter (X).Senior AI Infrastructure Engineer, Developer Experience Analog Devices, Inc. (NASDAQ: ADI) is a global semiconductor leader that bridges the physical and... 
    Permanent employment
    Full time
    Work at office
    Day shift

    Analog Devices

    Wilmington, MA
    2 days ago
  • $160k - $200k

    Job DescriptionWe're looking for a senior Platform engineer for GenAI Infrastructure to help build and operate the platform behind our agentic systems...  ...IAM, networking, and the supporting data stores that our AI agents rely on. This role partners with our existing DevOps... 
    Temporary work
    Work at office
    Remote work

    Bessemer Trust

    Woodbridge, NJ
    2 days ago
  • $190k - $310k

    The role As an AI platform engineer, you'll build the products, interfaces, and tools that define how people interact with Applied Compute...  ...for the enterprise. We provide the continual learning infrastructure for companies to build agent workforces trained on proprietary... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Applied Compute

    San Francisco, CA
    1 day ago
  • $190k - $260k

     ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...large-scale world models - depends on infrastructure that turns thousands of hours of multimodal...  ...training throughput. We are looking for engineers who make model training fast: streaming... 
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    3 days ago
  • $116.4k - $194k

    Role Summary The Lead AI Platform Engineer is a hands‑on technical contributor responsible for designing, developing, and integrating Generative AI-enabled capabilities within M&T Bank’s secure, governed enterprise environment. This role combines strong software engineering... 
    Full time
    Work experience placement

    M&T Bank

    Buffalo, NY
    2 days ago
  •  ...Description About us New Co is a new AI-native product organization within...  ...a deliberately small, senior team. Our engineering model is agentic: engineers author the specifications...  ...gateway, agent runtime, evaluation infrastructure, guardrails) and the shared product... 
    Full time
    Local area

    Capgemini

    Atlanta, GA
    2 days ago
  • $200k - $275k

     ...globally, ADI ensures today's innovators stay Ahead of What's Possible. Learn more at and on LinkedIn and Twitter (X).Principal AI Infrastructure Engineer, Developer Experience Analog Devices, Inc. (NASDAQ: ADI) is a global semiconductor leader that bridges the physical and... 
    Permanent employment
    Full time
    Work at office
    Day shift

    Analog Devices

    Wilmington, MA
    2 days ago
  • Role Description We’re building the next generation of infrastructure for AI-driven scientific discovery, and we need someone who can help own...  ..., reduce operational toil, support researchers and engineers, and make practical decisions about infrastructure. ~Design... 
    Full time
    Remote work

    FirstPrinciples

    Remote
    2 days ago
  • Role Description Radiology Partners, through its owned and affiliated practices, is a leading radiology practice in the U.S., serving hospitals and other healthcare facilities across the nation. As a physician-led and physician-owned practice, we advance our bold mission...
    Full time

    Radiology Partners

    Remote
    5 days ago
  • $155k - $180k

    Role Description The Senior AI Infrastructure Engineer will serve as a technical cornerstone of VideoAmp's AI Infra team, driving the design and execution of agentic workflow systems that bridge VideoAmp's platform APIs and AI-powered customer experiences. This is a high... 
    Full time
    Summer holiday
    Work at office
    Flexible hours

    VideoAmp

    Remote
    21 hours ago
  • $100k - $150k

    Role Description We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level... 
    Full time
    Local area
    Immediate start

    Bright Vision Technologies

    Remote
    21 hours ago
  •  ...AI Engineer – Routing & Network Optimizatio About Gallatin At Gallatin, we are rebuilding defense logistics for the warfighters of the United States and allied forces. We take an AI-first approach to modernizing how materiel, fuel, and equipment move from factory... 
    Full time

    Gallatin

    El Segundo, CA
    21 hours ago
  • Job PostingAI Cloud Engineer - Duties and ResponsibilitiesDevelop, deploy, and operate enterprise AI systems, APIs, and model-serving infrastructure that support retail operations at scale, including personalization, forecasting, pricing intelligence, and real-time decisioning... 

    Murphy USA

    El Dorado, AR
    4 days ago
  •  ...aspects of assigned platforms. Typically, coordinates with other engineers and teams for support and maintenance of platforms and services....  ...Microsoft Azure: Administrator Exper/Security/DevOps/AI/Data Science, and those related to Solution Architect or equivalent... 
    Work at office

    Synovus

    Columbus, GA
    2 days ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools...  .... We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in... 
    Full time
    Remote work

    Nvidia

    Austin, TX
    3 days ago
  • $86.8k - $198k

    AWS AI Cloud EngineerThe Opportunity: As an AWS AI Cloud Engineer, you can resolve a problem with a complete end-to-end GenAI solution in a fast, agile environment. If you’re looking for the chance to not just develop software, but to create a system that will make a difference... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Reston, VA
    5 days ago
  • $121k - $159k

     ...approximately $3.2 trillion. We create the intelligent infrastructure that powers global commerce, seamlessly...  ...what comes next.Job Title:Senior Cloud & AI EngineerCompany:PrologisA day in the lifeWe are seeking a senior cloud engineer who combines deep infrastructure... 
    Full time

    Prologis

    Denver, CO
    2 days ago
  • $191k - $315k

     ...Sunnyvale, CA.Team Overview: The Network Growth and Relationship AI team is at the forefront of creating cutting-edge, AI-powered...  ...network. The team works in close collaboration with the product, engineering and data science team and has a very exciting roadmap ahead. If... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    1 day ago
  • $77k - $100k

    Why This Role Is CriticalThe Azure Infrastructure & Cloud AI Engineer role is critical to Framatome’s digital transformation by designing, implementing, and maintaining secure, scalable cloud infrastructure that supports AI platforms, enterprise applications, and data-... 
    Temporary work
    Work experience placement

    Framatome

    Lynchburg, VA
    4 days ago
  • $105.5k - $213.5k

    Cloud and AI Ops EngineerThis role has been designed as ‘’Onsite’ with an expectation...  ...Edge (SASE) business. The Cloud and AI Engineer builds from the ground up to meet the...  ...Job Responsibilities):Deployment of Cloud infrastructure using Docker Containerization.Automate Cloud... 
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start
    2 days per week

    Hewlett Packard Enterprise

    San Jose, CA
    1 day ago
  • $151.3k - $283.8k

     ...technological advancements such as cloud, AI, and network security. While driving the digitalization...  ...of emerging technologies within cloud infrastructure.Who We Look For1.Education: Master’s or Ph.D. degree in Computer Engineering, Electronic Engineering, Microelectronics,... 
    Full time
    Relocation package

    Tencent

    Palo Alto, CA
    3 days ago
  •  ...sell easily and a better life within reach.The Global E-commerce Engineering Efficiency team focuses on building SDLC platforms, CI/CD...  ...performance through continuous system optimization.- Design and deliver AI-powered developer productivity solutions, including AI... 

    TikTok

    Seattle, WA
    19 hours ago
  • $55 - $65 per hour

    DescriptionKforce has a client that is seeking a Senior AWS Cloud & Agentic AI Engineer in Jupiter, FL.Summary:We are seeking a Senior AWS Cloud &...  ...with AI-driven and agentic workflows* Build and manage infrastructure using AWS CloudFormation and AWS CDK* Design and maintain CI... 

    KForce

    Jupiter, FL
    3 days ago
  • $126.2k - $264.1k

     ...distributed systems, networking, multi-tenant Infrastructure-as-a-Service (IaaS), and Software...  ...innovations to life-saving care. And with AI embedded across our products and...  ...in Computer Science, Electrical/Hardware Engineering or related field.Ability to work with minimal... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Austin, TX
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!