Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Systems Engineer

$100k - $150k
Full-time

Bright Vision Technologies

AI Systems Engineer – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: AI Systems Engineer
Location: 100% Remote (U.S.)
Position Type:  Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years

Sponsorship:  U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary
We are seeking an AI Systems Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control. The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work.

Key Responsibilities
  • Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations.
  • Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.
  • Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering.
  • Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate.
  • Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
  • Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics.
  • Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale.
  • Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing.
  • Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently.
  • Partner with research and applied ML teams to plan capacity for upcoming training runs.
  • Implement security controls, isolation, and access management for multi-tenant AI infrastructure.
  • Drive automation across cluster provisioning, lifecycle management, and configuration enforcement.
  • Maintain runbooks, capacity dashboards, and operational documentation for the AI platform.
  • Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
  • Strong understanding of Linux internals, networking, and high-performance storage.
  • Experience with at least one major cloud provider’s ML infrastructure offerings.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.
Preferred Qualifications
  • Experience operating InfiniBand or RDMA networking at scale.
  • Contributions to open-source ML infrastructure projects.
  • Familiarity with custom orchestrators or research-grade training stacks.
  • Exposure to frontier model training operations.
  • Experience with FinOps for AI workloads.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to View email address on brightvisiontechnologies.applytojob.com
Bright Vision Technologies is an Equal Opportunity Employer.
​​​​​​​

Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Systems Engineer in San Jose, CA vacancy
  • $160k - $180k

     ...What You’ll Do As a Senior AI Systems Engineer, you will architect, deploy, and manage the critical infrastructure services required for large-scale AI model training and inference. You will ensure our machine learning platforms are robust and efficient, bridging the gap... 
    Suggested
    Local area

    Archer56

    San Jose, CA
    5 days ago
  • $100k - $150k

     ...AI Systems Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected... 
    Suggested
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    San Jose, CA
    14 hours ago
  • NVIDIA is seeking a Senior System Software Engineer to enhance its Dynamo-Triton Inference Server. Your responsibilities will include developing GPU-accelerated AI inference software and facilitating the convergence of AI frameworks. Applicants should possess an MS or... 
    Suggested

    NVIDIA

    Santa Clara, CA
    5 days ago
  •  ...company in California is seeking a skilled engineer to optimize deep learning frameworks for...  ...enhance performance across multi-GPU systems, focus on GPU kernel development, and collaborate...  ...GPGPU C++, Triton, and experience with AI frameworks like PyTorch and SGLang. A... 
    Suggested
    Full time

    AMD

    Santa Clara, CA
    5 days ago
  • $120.5k - $243k

    Hewlett Packard Enterprise is seeking a Software Engineer for their innovative HPE Mist Networking team in Cupertino, California. This...  ...involves developing and testing features across embedded and cloud systems, requiring strong programming skills in Go, C, or Python, along... 
    Suggested

    Hewlett Packard Enterprise

    Cupertino, CA
    5 days ago
  • AMD is seeking a senior software engineer who blends deep systems performance work with AI, spanning GPU kernels to distributed training and inference. You’ll influence ROCm and AMD’s AI software strategy while mentoring others and tackling high-impact problems on AMD GPUs... 

    AMD

    Santa Clara, CA
    5 days ago
  • NVIDIA is seeking a Sr. Software Engineer in Santa Clara specializing in AI products for semiconductor inspection. This role focuses on developing production-ready models and inference pipelines designed to enhance semiconductor manufacturing processes. Ideal candidates... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

    NVIDIA is seeking AI systems engineers to innovate in the inference systems software stack. The role involves designing and optimizing libraries and kernel technologies, significantly impacting AI workloads. Candidates should have a Master's degree (PhD preferred) and... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • NVIDIA AI is seeking a Sr. Software Engineer to develop cutting-edge AI products for semiconductor inspection in Santa Clara. The role focuses on creating production-ready models in computer vision and anomaly detection, addressing challenges with limited data and tight... 

    NVIDIA AI

    Santa Clara, CA
    5 days ago
  • NVIDIA is seeking outstanding AI systems engineers in Santa Clara to advance the inference software stack. You will build libraries, code generators, and GPU kernels for NVIDIA hardware, designing abstractions for LLM serving engines and JIT compilers to accelerate large... 

    Segment (Twilio)

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...seeking an experienced applied ML expert to join the Agentic Engineering team in Santa Clara, California. In this role, you'll develop...  ...Candidates must possess strong Python skills and experience with AI systems. The base salary ranges from $184,000 to $287,500 for Level 4... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $224k - $356.5k

    NVIDIA is seeking a System Software Engineer for Vision AI in Santa Clara, CA. This role involves developing high-performance vision systems and optimizing AI pipelines. Ideal candidates have 12+ years of experience in software development using C++ and Python and a solid... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $152k - $208.5k

     ...Materials is a global leader in materials engineering solutions used to produce virtually every...  ...that literally connect our world - like AI and IoT. If you want to push the boundaries...  ...you’ll contribute expertise in intricate systems, deciphering code, and anticipating... 
    Full time
    Relocation

    Applied Materials

    Santa Clara, CA
    2 days ago
  • NVIDIA Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team... 

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • NVIDIA is seeking a data storage engineer to advance cloud-native storage across AI workloads. You will design storage technologies, client libraries, and filesystem frameworks to support exabyte-scale data management in hybrid environments. You will optimize training... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • NVIDIA is seeking a System Software Engineer to work on Dynamo‑Triton Inference Server in Santa Clara, CA. You will contribute to a high-performance, GPU-accelerated AI inference platform that serves both LLM and non-LLM workloads and participate in a fast-paced open source... 

    2100 NVIDIA USA

    Santa Clara, CA
    3 days ago
  •  ...Corporation is looking for an experienced professional to optimize AI workloads at DGX Station (Galaxy) in Santa Clara,...  ...stack. The ideal candidate has over 12 years of experience in systems software engineering and strong expertise with deep learning frameworks. The... 

    Nvidia Corporation

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

    NVIDIA Gruppe is seeking talented AI systems engineers to advance innovative technologies in AI inference systems software. This role involves developing cutting-edge libraries, code generators, and kernel technologies for NVIDIA's architecture, emphasizing high-impact... 

    NVIDIA Gruppe

    Santa Clara, CA
    5 days ago
  • NVIDIA is looking for a dedicated System Software Engineer to join their team in Santa Clara, California. The role focuses on AI infrastructure, where you will collaborate closely with various engineering and marketing teams to enhance the performance of NVIDIA's products... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • Applied Materials is hiring an Agentic AI Systems Engineer in Santa Clara, CA. This role involves designing the infrastructure for GenAI applications, bridging AI and software needs. Candidates must have 7+ years of experience with a strong proficiency in programming languages... 

    Applied Materials

    Santa Clara, CA
    2 days ago
  • $184k - $356.5k

    NVIDIA is seeking a Senior Software Engineer in Santa Clara, CA to solve exciting problems at the intersection of quantum computing, distributed systems, and data platforms. You will design data interfaces and build orchestration layers connecting various applications.... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • NVIDIA is seeking a Senior Software Engineer to join its Metropolis Synthetic Data Generation team in Santa Clara, California. The role requires building scalable solutions for AI and 3D simulation applications while collaborating with multi-functional teams. The ideal... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  • Siemens in Santa Clara, CA is seeking a Principal Engineer to set the technical direction for systems that blend cloud infrastructure, machine learning, and hardware interaction. You will define architecture that is robust, scalable, and adaptable to rapid research and... 

    Siemens

    Santa Clara, CA
    5 days ago
  • Tenstorrent is building high‑performance AI compute clusters and TT‑Fabric, the low‑level networking layer that connects thousands of RISC‑V and AI processors. We seek engineers who can architect and implement scalable networking for distributed inference and training... 

    Tenstorrent

    Santa Clara, CA
    2 days ago
  • $184k - $356k

    Nvidia Corporation is seeking a Senior Inference Engineer to advance AIConfigurator, optimizing deployment configurations for large-scale LLM inference on NVIDIA platforms. The role involves building production-quality APIs and collaborating with multiple teams to improve... 

    Nvidia Corporation

    Santa Clara, CA
    5 days ago
  • NVIDIA is seeking a Senior System Mechanical Engineer in Santa Clara, CA to lead mechanical system development for GPU and AI infrastructure. You will drive chassis design, cooling, and system integration from concept to production, collaborating with electrical and thermal... 

    NVIDIA

    Santa Clara, CA
    3 days ago
  • NVIDIA is seeking an experienced software engineer for its DGXC Data Services team to build...  ...infrastructure supporting exabyte-scale AI workloads. You will contribute to storage...  ...paths, integrate object stores with file systems, and develop telemetry to diagnose bottlenecks... 

    NVIDIA AI

    Santa Clara, CA
    5 days ago
  • Role, Inc. is seeking a Systems Software Engineer to contribute to HPE's Marvis Minis and edge AI technology. This role will involve developing and testing features across embedded and cloud systems, with a focus on networking applications. Successful candidates will have... 
    2 days per week

    Role, Inc.

    Cupertino, CA
    5 days ago
  • $181.1k - $318.4k

     ...technology company in Cupertino is seeking a Machine Learning Engineer to build infrastructure for product-focused machine learning projects...  .... The ideal candidate will have a strong background in backend systems development and solid knowledge of machine learning... 

    Apple

    Cupertino, CA
    5 days ago
  • Apple Inc. in Cupertino, California, is seeking a senior ML/AI engineer to help turn research into practical features on Apple platforms...  ...team that collaborates across disciplines to build scalable AI systems, with hands‑on prototyping and observable, debuggable... 

    Apple Inc.

    Cupertino, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Systems Engineer. Be the first to apply!