AI Infrastructure Engineer
Bright Vision Technologies
AI Infrastructure Engineer
Job Title: AI Infrastructure EngineerLocation: 100% Remote (Continental United States)
Position Type: In-house Bright Vision Technologies SOW engagement (no third-party client or vendor)
Experience: 6+ years
Salary - 100 K - 150 K
Sponsorship: No new H1B sponsorship available. H1B transfers welcomed for qualified candidates.
Employment Type: Full-time, direct W2 with Bright Vision Technologies (no C2C, no 1099, no third-party)
Engagement: Long-term, multi-year, aligned to the Bright Vision SOW delivery roadmap
Compensation: Competitive base salary commensurate with experience, plus benefits. Employment Terms & Visa Policy
This is a 100% remote, full-time, direct W2 position with Bright Vision Technologies.
This role is part of Bright Vision Technologies’ in-house Statement of Work (SOW) engagement. The client, end customer, and employer for this position is Bright Vision Technologies — there is no third-party client, vendor, or implementation partner involved.
We do not engage in C2C, 1099, or third-party arrangements for this role. BUT STRICTLY NO C2C/1099/3RD PARTY COMPANIES. ALL OUR ROLES ARE W2 AND NO 3RD PARTY BROKERING PLEASE.
Candidates must be willing to work directly as a full-time W2 employee of Bright Vision Technologies and contribute to our in-house SOW deliverables.
No new H1B sponsorship is available for this role.
However, candidates who are currently on a valid H1B visa and require a transfer are welcome to apply. We will support H1B transfers for qualified candidates.
For every role, a technical coding assessment is mandatory. Please apply only if you are confident in your technical abilities and hands-on experience. Job Summary
We are seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control. The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work. Key Responsibilities
- Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations.
- Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.
- Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering.
- Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate.
- Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
- Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics.
- Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale.
- Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing.
- Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently.
- Partner with research and applied ML teams to plan capacity for upcoming training runs.
- Implement security controls, isolation, and access management for multi-tenant AI infrastructure.
- Drive automation across cluster provisioning, lifecycle management, and configuration enforcement.
- Maintain runbooks, capacity dashboards, and operational documentation for the AI platform.
- Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling.
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Six or more years of experience in infrastructure, platform, or HPC engineering.
- Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
- Strong proficiency in Python and at least one systems language such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, and collective communication.
- Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
- Strong understanding of Linux internals, networking, and high-performance storage.
- Experience with at least one major cloud provider’s ML infrastructure offerings.
- Strong software engineering practices including testing, CI/CD, and code review.
- Excellent communication and cross-functional collaboration skills.
- Experience operating InfiniBand or RDMA networking at scale.
- Contributions to open-source ML infrastructure projects.
- Familiarity with custom orchestrators or research-grade training stacks.
- Exposure to frontier model training operations.
- Experience with FinOps for AI workloads.
Would you like to know more about this opportunity?
For immediate consideration, please send your resume to View email address on click.appcast.io or contact us at View phone number on click.appcast.io. Learn more about Bright Vision Technologies at
We recognize that our people are our strength, and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company.
We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs.
Bright Vision Technologies is an Equal Opportunity Employer, including Disability/Veterans.
Position offered by “No Fee Agency.”
Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
$120k - $190k
...Job Responsibilities Lead execution of AI-assisted development and AI-enabled... ...Architecture, AI Architecture, Cybersecurity, Infrastructure, and Business teams to ensure aligned... ...in developer platforms, internal engineering tooling, or DevEx initiatives Experience...SuggestedLocal areaImmediate startRemote workFlexible hours- Electrical Hardware Engineer - HPC/AI Platform Engineering (Early Career) This role is onsite, with an expectation that you will primarily work from an HPE office. Key Responsibilities Design portions of electrical and electronic parts, subsystems, integrated circuitry,...SuggestedWork at officeLocal area
- ...technologies to create scalable, secure, and user-friendly applications. As we continue to grow, we’re looking for a skilled AI Data Infrastructure Engineer to join our dynamic team and contribute to our mission of transforming business processes through technology. This...SuggestedFull timeH1bLocal areaImmediate startRemote workVisa sponsorshipWork visa
$120k - $190k
...Job Responsibilities Lead execution of AI-assisted development and AI-enabled... ...Architecture, AI Architecture, Cybersecurity, Infrastructure, and Business teams to ensure aligned... ...in developer platforms, internal engineering tooling, or DevEx initiatives Experience...Suggested$119.5k - $275k
...looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE. Job Description: HPC & AI Senior Performance Engineer HPE is seeking an experienced HPC performance engineer who is excited to help drive performance of systems used to...SuggestedWork experience placementWork at officeLocal areaImmediate startWorldwide2 days per week- ...AI Engineer - Developer Productivity This role has been designed as ‘Hybrid’ with an expectation that you will work on average 2 days per week from an HPE office. Who We Are: Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people...Work experience placementWork at officeLocal areaImmediate start2 days per week
$60 - $80 per hour
...technology, partnering closely with stakeholders across the organization to identify, prioritize, and deliver high-value AI solutions. The AI Engineer will design, build, test, and deploy Copilot agents using tools such as Agent Builder and Copilot Studio, grounding...Hourly payWork at office$105.05k - $161.8k
...Sr. Gaming AI Engineer Description - We are seeking a highly skilled Gaming AI Engineer with strong system engineering and hands-on coding experience to design, develop, and integrate AI-driven capabilities across HP's Gaming Solutions portfolio. This role bridges...Full timeTemporary workLocal areaRelocationFlexible hoursShift work- ...AI Engineer — AI Efficiency and Economics This role has been designed as "Onsite" with an expectation that you will primarily work from an HPE office. Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies...Work at office
- ...Job Description The Senior AI Agentic Engineer designs, builds, and operationalizes intelligent agent systems that automate complex... ...Service, AWS Bedrock, Google Vertex AI) with containerized infrastructure (Docker, Kubernetes). ~10. Experience implementing...Permanent employmentWork experience placement
- ...AI Engineer — AI Efficiency and Economics This role has been designed as ‘’Onsite’ with an expectation that you will primarily work from an HPE office. Who We Are: Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and...Work experience placementWork at officeLocal areaImmediate start
- ...Description • Design, build, and deploy end to end, multi step agentic AI systems that replace or reimagine DBS business workflows. •... ...systems. Strong proficiency in Python, REST APIs, prompt engineering, RAG pipelines, and agent orchestration, with experience using...
- ...to join our team. If you're excited to be part of a winning team, CirrusLabs () is a great place to grow your career. Role: AI Data Engineer- Expert Level Location: Spring, Texas (5 Days Onsite Each Week) Duration: Long term contact The ideal candidate...
$119.5k - $275k
Hewlett Packard Enterprise Development LP is looking for an HPC & AI Senior Performance Engineer with expertise in high-end HPC systems and performance tuning. This hybrid role requires collaboration on technical benchmarking and supporting customer opportunities. The...$119.5k - $275k
HPC & AI Senior Performance Engineer Hybrid role: Expected to work on average 2 days per week from an HPE office. Location: Bloomington, MN or Spring, TX Responsibilities Provide pre‑ and post‑sales technical and benchmarking support to enable HPC or AI customer opportunities...Work at office2 days per week$120.5k - $276.5k
.... Typically, 8+ years of experience Product/Solutions Engineering or development experience Very good analytical/problem solving... ...Experience of server/storage/network management(GPU infrastructure) In-depth understanding and hands-on experience in Kubernetes...Work experience placementWork at office- ...of market data, trading systems, and financial instruments related to oil and gas. ~ Bachelor's degree in Computer Science, Engineering, or a related field. ~3+ years of professional software development experience ~ Certifications in relevant technologies...Long term contractLocal area
- ...Overview: Platform Engineer with GenAI experience + Full Stack skillset Position Responsibilities... ...proactive engineering Author infrastructure-as-code using Terraform for cloud... ...of RAG, model orchestration, and AI application patterns Soft Skills:...Long term contractLocal areaRelocation
- ...Functions: You will be responsible for virtualization infrastructure software development. You will design, implement, monitor,... ..., and develop test plans and documentation with senior engineers' guidance. You will partner with architects and engineering...Work experience placementWork at officeShift work2 days per week
- ...and cross-cloud authorization patterns that let workloads and engineers move between providers without compromising on least-... ...connectivity, and zero-trust patterns. Establish reusable infrastructure-as-code patterns that abstract cloud-specific implementations...Full timeH1bLocal areaImmediate startRemote workVisa sponsorshipWork visa
- ...development of storage and networking capabilities for virtualization infrastructure. You will design, implement, monitor, and troubleshoot... ..., and develop test plans and documentation with senior engineers' guidance. You will partner with architects and engineering...Work experience placementWork at officeShift work2 days per week
- ...Cloud Infrastructure Engineer (AWS & Azure) Location: The Woodlands, TX Department: Information Technology / Infrastructure Position Summary The Cloud Infrastructure Engineer is responsible for building, supporting, and maintaining cloud infrastructure across...
- ...1 | Cloud Developer - Advanced Level Job Title:- AWS Cloud Engineer Location:- Spring Texas (On-Site) Job Type:- Long Term Contract... ..., and reliable AWS cloud platform. Develop and maintain Infrastructure as Code (IaC) using Terraform to automate the provisioning...Long term contractLocal areaRelocation
- ...Title: Sr. Cloud Engineer (Azure & AWS) Location: Hybrid in The Woodlands, TX - 773... ...Overview As a member of our client's Infrastructure Team, this position will report to the... ...Solutions' Privacy Policy and INSPYR Solutions' AI and Automated Employment Decision Tool...Contract workWork at officeLocal areaRemote workFlexible hours
- ...customers Debug and fix software issues in highly scalable and performance intensive deployments Collaborate with cross-functional engineering teams (Configuration and Control path) to ensure seamless integration Optimize performance of SDN in truly distributed, cloud...Work experience placementWork at office2 days per week
$155.66k - $225.16k
...with one place to chat, explore and build with a wide variety of AI language models (bots), including o3, o4-mini, Claude 3.7 Sonnet... ...the Team and Role: We’re hiring our first AI Automation Engineer to lead how we apply AI internally across the company. This is...Remote jobFull timeShift work- ...continue to grow, we’re looking for a skilled VMware Platform Engineer to join our dynamic team and contribute to our mission of... ...Recovery Manager and replication. Build automation for VMware infrastructure using PowerCLI, Terraform, Ansible, and vRO. Design...Full timeH1bLocal areaImmediate startRemote workVisa sponsorshipWork visa
$90 - $100 per hour
...client in Spring, TX is looking for a Cloud Platform Standards Engineer to design, author, and operationalize Enterprise Architecture... ...standards through reference ('golden') architectures and Infrastructure-as-Code (IaC). This person will be defining enterprise standards...- ...Cloud Platform Engineer (AWS) We are CirrusLabs. Our vision is to become the world's most sought-after niche digital transformation... ...more platform governance & security than app development. Infrastructure as Code Strong Terraform background (writing modules,...Contract work
- ...to grow, we’re looking for a skilled SAP Basis / SAP Platform Engineer to join our dynamic team and contribute to our mission of... ...system copies, upgrades, and migrations, and will partner with infrastructure, security, and SAP functional teams to deliver a stable, well...Full timeH1bLocal areaImmediate startRemote workVisa sponsorshipWork visa
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!



