AI Systems Engineer
$100k - $150kBright Vision Technologies
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: AI Systems Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. Job Summary
We are seeking an AI Systems Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control. The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work. Key Responsibilities
- Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations.
- Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.
- Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering.
- Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate.
- Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
- Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics.
- Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale.
- Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing.
- Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently.
- Partner with research and applied ML teams to plan capacity for upcoming training runs.
- Implement security controls, isolation, and access management for multi-tenant AI infrastructure.
- Drive automation across cluster provisioning, lifecycle management, and configuration enforcement.
- Maintain runbooks, capacity dashboards, and operational documentation for the AI platform.
- Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling.
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Six or more years of experience in infrastructure, platform, or HPC engineering.
- Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
- Strong proficiency in Python and at least one systems language such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, and collective communication.
- Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
- Strong understanding of Linux internals, networking, and high-performance storage.
- Experience with at least one major cloud provider’s ML infrastructure offerings.
- Strong software engineering practices including testing, CI/CD, and code review.
- Excellent communication and cross-functional collaboration skills.
- Experience operating InfiniBand or RDMA networking at scale.
- Contributions to open-source ML infrastructure projects.
- Familiarity with custom orchestrators or research-grade training stacks.
- Exposure to frontier model training operations.
- Experience with FinOps for AI workloads.
Would you like to know more about this opportunity? For immediate consideration, please send your resume to View email address on brightvisiontechnologies.applytojob.com
Bright Vision Technologies is an Equal Opportunity Employer.
Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
$160k - $180k
...What You’ll Do As a Senior AI Systems Engineer, you will architect, deploy, and manage the critical infrastructure services required for large-scale AI model training and inference. You will ensure our machine learning platforms are robust and efficient, bridging the gap...SuggestedLocal area$100k - $150k
...AI Systems Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected...SuggestedFull timeH1bLocal areaImmediate startRemote workVisa sponsorship- NVIDIA is seeking a Senior System Software Engineer to enhance its Dynamo-Triton Inference Server. Your responsibilities will include developing GPU-accelerated AI inference software and facilitating the convergence of AI frameworks. Applicants should possess an MS or...Suggested
- ...company in California is seeking a skilled engineer to optimize deep learning frameworks for... ...enhance performance across multi-GPU systems, focus on GPU kernel development, and collaborate... ...GPGPU C++, Triton, and experience with AI frameworks like PyTorch and SGLang. A...SuggestedFull time
$120.5k - $243k
Hewlett Packard Enterprise is seeking a Software Engineer for their innovative HPE Mist Networking team in Cupertino, California. This... ...involves developing and testing features across embedded and cloud systems, requiring strong programming skills in Go, C, or Python, along...Suggested- AMD is seeking a senior software engineer who blends deep systems performance work with AI, spanning GPU kernels to distributed training and inference. You’ll influence ROCm and AMD’s AI software strategy while mentoring others and tackling high-impact problems on AMD GPUs...
- NVIDIA is seeking a Sr. Software Engineer in Santa Clara specializing in AI products for semiconductor inspection. This role focuses on developing production-ready models and inference pipelines designed to enhance semiconductor manufacturing processes. Ideal candidates...
$184k - $287.5k
NVIDIA is seeking AI systems engineers to innovate in the inference systems software stack. The role involves designing and optimizing libraries and kernel technologies, significantly impacting AI workloads. Candidates should have a Master's degree (PhD preferred) and...- NVIDIA AI is seeking a Sr. Software Engineer to develop cutting-edge AI products for semiconductor inspection in Santa Clara. The role focuses on creating production-ready models in computer vision and anomaly detection, addressing challenges with limited data and tight...
- NVIDIA is seeking outstanding AI systems engineers in Santa Clara to advance the inference software stack. You will build libraries, code generators, and GPU kernels for NVIDIA hardware, designing abstractions for LLM serving engines and JIT compilers to accelerate large...
$184k - $287.5k
...seeking an experienced applied ML expert to join the Agentic Engineering team in Santa Clara, California. In this role, you'll develop... ...Candidates must possess strong Python skills and experience with AI systems. The base salary ranges from $184,000 to $287,500 for Level 4...$224k - $356.5k
NVIDIA is seeking a System Software Engineer for Vision AI in Santa Clara, CA. This role involves developing high-performance vision systems and optimizing AI pipelines. Ideal candidates have 12+ years of experience in software development using C++ and Python and a solid...$152k - $208.5k
...Materials is a global leader in materials engineering solutions used to produce virtually every... ...that literally connect our world - like AI and IoT. If you want to push the boundaries... ...you’ll contribute expertise in intricate systems, deciphering code, and anticipating...Full timeRelocation- NVIDIA Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team...
- NVIDIA is seeking a data storage engineer to advance cloud-native storage across AI workloads. You will design storage technologies, client libraries, and filesystem frameworks to support exabyte-scale data management in hybrid environments. You will optimize training...
- NVIDIA is seeking a System Software Engineer to work on Dynamo‑Triton Inference Server in Santa Clara, CA. You will contribute to a high-performance, GPU-accelerated AI inference platform that serves both LLM and non-LLM workloads and participate in a fast-paced open source...
- ...Corporation is looking for an experienced professional to optimize AI workloads at DGX Station (Galaxy) in Santa Clara,... ...stack. The ideal candidate has over 12 years of experience in systems software engineering and strong expertise with deep learning frameworks. The...
$184k - $287.5k
NVIDIA Gruppe is seeking talented AI systems engineers to advance innovative technologies in AI inference systems software. This role involves developing cutting-edge libraries, code generators, and kernel technologies for NVIDIA's architecture, emphasizing high-impact...- NVIDIA is looking for a dedicated System Software Engineer to join their team in Santa Clara, California. The role focuses on AI infrastructure, where you will collaborate closely with various engineering and marketing teams to enhance the performance of NVIDIA's products...
- Applied Materials is hiring an Agentic AI Systems Engineer in Santa Clara, CA. This role involves designing the infrastructure for GenAI applications, bridging AI and software needs. Candidates must have 7+ years of experience with a strong proficiency in programming languages...
$184k - $356.5k
NVIDIA is seeking a Senior Software Engineer in Santa Clara, CA to solve exciting problems at the intersection of quantum computing, distributed systems, and data platforms. You will design data interfaces and build orchestration layers connecting various applications....- NVIDIA is seeking a Senior Software Engineer to join its Metropolis Synthetic Data Generation team in Santa Clara, California. The role requires building scalable solutions for AI and 3D simulation applications while collaborating with multi-functional teams. The ideal...
- Siemens in Santa Clara, CA is seeking a Principal Engineer to set the technical direction for systems that blend cloud infrastructure, machine learning, and hardware interaction. You will define architecture that is robust, scalable, and adaptable to rapid research and...
- Tenstorrent is building high‑performance AI compute clusters and TT‑Fabric, the low‑level networking layer that connects thousands of RISC‑V and AI processors. We seek engineers who can architect and implement scalable networking for distributed inference and training...
$184k - $356k
Nvidia Corporation is seeking a Senior Inference Engineer to advance AIConfigurator, optimizing deployment configurations for large-scale LLM inference on NVIDIA platforms. The role involves building production-quality APIs and collaborating with multiple teams to improve...- NVIDIA is seeking a Senior System Mechanical Engineer in Santa Clara, CA to lead mechanical system development for GPU and AI infrastructure. You will drive chassis design, cooling, and system integration from concept to production, collaborating with electrical and thermal...
- NVIDIA is seeking an experienced software engineer for its DGXC Data Services team to build... ...infrastructure supporting exabyte-scale AI workloads. You will contribute to storage... ...paths, integrate object stores with file systems, and develop telemetry to diagnose bottlenecks...
- Role, Inc. is seeking a Systems Software Engineer to contribute to HPE's Marvis Minis and edge AI technology. This role will involve developing and testing features across embedded and cloud systems, with a focus on networking applications. Successful candidates will have...2 days per week
$181.1k - $318.4k
...technology company in Cupertino is seeking a Machine Learning Engineer to build infrastructure for product-focused machine learning projects... .... The ideal candidate will have a strong background in backend systems development and solid knowledge of machine learning...- Apple Inc. in Cupertino, California, is seeking a senior ML/AI engineer to help turn research into practical features on Apple platforms... ...team that collaborates across disciplines to build scalable AI systems, with hands‑on prototyping and observable, debuggable...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Systems Engineer. Be the first to apply!
- ai engineer San Jose, CA
- senior ai engineer San Jose, CA
- ai developer San Jose, CA
- ai engineer remote San Jose, CA
- advanced systems engineer San Jose, CA
- distributed systems engineer San Jose, CA
- computer system validation engineer San Jose, CA
- mission system engineer San Jose, CA
- senior staff systems engineer San Jose, CA
- software system engineer San Jose, CA

