Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer

Bright Vision Technologies

AI Infrastructure Engineer Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: AI Infrastructure Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 Salary Range: $100,000–$160,000 Annually Experience Required: 10+ years Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. Job Summary: We are seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control. The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work. Key Responsibilities Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations. Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams. Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering. Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate. Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication. Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics. Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale. Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing. Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently. Partner with research and applied ML teams to plan capacity for upcoming training runs. Implement security controls, isolation, and access management for multi-tenant AI infrastructure. Drive automation across cluster provisioning, lifecycle management, and configuration enforcement. Maintain runbooks, capacity dashboards, and operational documentation for the AI platform. Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling. Required Qualifications Bachelor's or Master's degree in Computer Science or a related field. Ten or more years of experience in infrastructure, platform, or HPC engineering. Hands-on experience operating GPU clusters or large-scale ML training infrastructure. Strong proficiency in Python and at least one systems language such as Go or C++. Deep understanding of distributed training, accelerator architectures, and collective communication. Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads. Strong understanding of Linux internals, networking, and high-performance storage. Experience with at least one major cloud provider's ML infrastructure offerings. Strong software engineering practices including testing, CI/CD, and code review. Excellent communication and cross-functional collaboration skills. Preferred Qualifications Experience operating InfiniBand or RDMA networking at scale. Contributions to open-source ML infrastructure projects. Familiarity with custom orchestrators or research-grade training stacks. Exposure to frontier model training operations. Experience with FinOps for AI workloads. Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment. Bright Vision Technologies

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in New York, NY vacancy
  • $175k - $275k

     ...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in...  ...would be possible. The Role As an AI Engineer - Cloud Infrastructure on Traversal’s Infrastructure team, you’ll design,... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Traversal

    New York, NY
    more than 2 months ago
  • $215k - $350k

    We are seeking an AI Infrastructure Engineer to build, operate, and continuously enhance the Linux and GPU-based infrastructure that powers our AI platforms and performance testing environments. This is a highly hands-on infrastructure, automation, and performance engineering... 
    Suggested
    Worldwide
    Home office

    Fortinet

    New York, NY
    15 days ago
  • $160k - $220k

     ...of:     The role   We are looking for an experienced AI Engineer to lead the implementation of Azure AI Foundry within an established...  ...our existing data platform, governance model, and analytics infrastructure.   You will work closely with data engineering,... 
    Suggested
    Full time
    Remote work

    Valtech Se

    New York, NY
    more than 2 months ago
  •  ...About Build AI for the Built World: Build has created the agentic AI stack for institutional...  ...most important built projects - digital infrastructure, energy, industrial - from concept to...  ...create and configure workflows without engineering. The interfaces through which clients... 
    Suggested
    Full time
    Live in

    Build Technologies

    New York, NY
    28 days ago
  • $150k - $300k

     ...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in...  ...constantly. This is a place to grow your career, make a real impact, and help define a new category of infrastructure software.... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Traversal

    New York, NY
    more than 2 months ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer, Gen AI Platform Overview: At Capital One, we are creating responsible and reliable...  ...personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in... 
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    New York, NY
    more than 2 months ago
  • $160k - $220k

     .... As a global leader in residential wellness and healthcare infrastructure, we create vibrant, purpose-driven communities where housing...  ...we want you on our best-in-class team. SUMMARY The AI Engineer, Platform and Data builds the core that everything else runs... 
    Full time

    Welltowercareers

    New York, NY
    1 day ago
  • Overview Our client is a large global integrator. They are seeking an AI Infrastructure Engineer. In this role you will support the modernization of applications within the client's program by designing, deploying, and managing secure AWS infrastructure and AI/ML integrations... 
    Hourly pay
    Contract work
    Remote work

    The Squires Group

    New York, NY
    1 day ago
  • Palona’s AI agents operate continuously in production, handle real-time guest interactions...  ...systems, and face sharp traffic peaks. Infrastructure is therefore part of the product:...  ...experience. We are looking for an Infrastructure Engineer who combines cloud and reliability depth... 
    Temporary work

    Palona AI

    New York, NY
    4 days ago
  • AI Platform Engineer About Tessera Labs Tessera Labs is redefining how enterprises adopt and operationalize Artificial Intelligence. Backed...  ...team builds and operates the foundational AI agent infrastructure that lets Tessera run reliably, securely, and consistently... 

    Tessera Labs

    New York, NY
    4 days ago
  • $100k - $150k

    AI Infrastructure Engineer- Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    New York, NY
    4 days ago
  • Job Title: AI Infrastructure Engineer Job Summary We are seeking an AI Infrastructure Engineer to design, implement, and manage the infrastructure that powers AI and machine learning workloads. In this role, you will build scalable, secure, and high-performance environments... 
    Full time
    Remote work

    Ova Technologies

    New York, NY
    4 days ago
  • $162k - $215k

    AI-Enabled Infrastructure And Systems Engineer Aladdin Platform Engineering powers the technology foundation behind BlackRock's Aladdin platform - a unified system that connects risk, portfolio management, trading, and operations for investors globally. We build the mission... 
    Apprenticeship
    Work at office
    Work from home
    Worldwide
    Flexible hours
    1 day per week

    Hackajob

    New York, NY
    4 days ago
  • $200k

    AI Infrastructure Engineer Location: New York (4 Days Onsite) Base Salary: $200k + 50% bonus This is a rare opportunity to shape the AI foundations of a complex, global organisation at a pivotal moment in its technology journey. You will play a central role in building... 
    Work at office

    Harnham

    New York, NY
    3 days ago
  • Senior AI Storage Infrastructure Engineer Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to... 
    Local area

    Bitdeer Technologies Group

    New York, NY
    4 days ago
  • As an AI Platform Engineer for AI & Emerging Tech, you will drive AI platform enablement across the enterprise. This role sits at the intersection...  ...guardrails (budgets, alerts, quota enforcement) using Infrastructure as Code. Act as a subject... 

    Luxoft

    New York, NY
    more than 2 months ago
  •  ...We are hiring a Senior AI Platform Engineer for a 12-month contract position based in New York City. This role focuses on designing, building, and securing AI platform infrastructure in a regulated environment, bridging technical implementation with security and compliance... 
    Contract work

    Intone Inc

    New York, NY
    1 day ago
  •  ...PRI Technology in New York, NY seeks a Senior Software Engineer to build reliable, compliant AI platform services powering enterprise workflows. You will design Kubernetes-based PaaS, AI/LLM gateways, and self-service developer tooling, with emphasis on reliability... 

    PRI Technology

    New York, NY
    1 hour ago
  • $200k - $230k

    SVP, Lead AI/MLOps Infrastructure Engineer - Full Time - HybridWe’re partnering with our client, a fast-growing fintech firm, on a senior-level hire to lead and own the infrastructure behind their AI and machine learning platforms. This is a highly visible, hands-on leadership... 
    Full time
    Remote work

    Benchmark IT

    New York, NY
    3 days ago
  • $175k - $200k

     ...dedicated owner for OUTFRONT's internal AI platform. Today, our internal AI assistant...  ...into production, and our external agent infrastructure (the Agency Connect MCP server, which...  ...early partner adoption. Both need a senior engineering owner, and both need to grow into the... 
    Full time
    Internship

    OUTFRONT Media

    New York, NY
    a month ago
  • $197.3k - $225.1k

     ...Lead AI Engineer ( MLX, Gen AI Platform Services, Agentic AI) Overview At Capital One, we are creating responsible and reliable...  ...personalized customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    2 days ago
  • Google Cloud AI Engineer POC We are seeking a highly skilled Artificial Intelligence Engineer. This role is pivotal in establishing...  ...preprocessing techniques (scaling, encoding, imputation). • Cloud Infrastructure: Hands-on experience with Google Cloud Storage and Vertex AI... 

    United IT

    New York, NY
    4 days ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (GenAI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    4 days ago
  • $65k - $102k

     ...Requirements: We require 8+ years of software engineering experience, including 3+ years in a...  ...leadership role building multi-agent AI systems, FinOps tools, or LLM-powered...  ...familiarity with Google Cloud Platform infrastructure constraints for selecting optimal AI agentic... 
    Full time

    EPAM Systems

    New York, NY
    6 days ago
  • $130k - $147k

     ...customers at the center of every decision. Our AI-first platform transforms proprietary...  ...We are seeking a Senior ML/AI Platform Engineer to help build and operate Curinos'...  ...responsible for designing and implementing the infrastructure, tooling, automation, observability, and... 
    Part time
    Work at office
    Remote work
    Work from home
    Flexible hours

    Curinos Inc

    New York, NY
    27 days ago
  • $180k - $230k

    About the RoleThe Platform Infrastructure team at iCapital plays a critical role in ensuring that...  ...and intellectually curious MLOps/DevOps Engineers with deep expertise in machine learning...  ...).Enable production workloads for AI/ML and Generative AI systems, including... 
    Full time
    Work at office
    Remote work

    iCapital Network

    New York, NY
    a month ago
  • $162k - $215k

     ...Requirements: We need strong experience building and operating production AI systems, including LLM-based workflows. We need a solid grasp of modern AI methods such as prompt engineering, retrieval-augmented generation, fine-tuning, and evaluation. We need proven... 
    Full time
    Apprenticeship
    Work at office
    Work from home
    Flexible hours

    BlackRock

    New York, NY
    10 days ago
  • $155k - $215k

     ...Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm's Technology principles...  ...of AI across our businesses. This role is for a platform engineering specialist who will help build a firmwide AI Development... 
    Full time
    Temporary work

    Morgan Stanley

    New York, NY
    13 days ago
  •  ...As the Lead AI Engineer for our next-generation TCO Agent Platform, you will serve as the technical lead and multi-agent system architect...  ..., and API/MCP error specifications (RFC 7807/9457) DevOps & Infrastructure: Proficiency with Docker builds, GKE deployment patterns,... 

    EPAM Systems

    New York, NY
    12 days ago
  • The OpportunityJoin a team building the data foundations that support the firm’s AI and analytics capabilities. This role sits within the engineering effort to develop a modern Lakehouse and AI data platform that enables reliable, well-governed and high-performing data... 

    Goldman Sachs

    New York, NY
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!