Sr. AI Infrastructure Engineer- Hybrid
Calance
Senior AI Infrastructure Engineer, Physical Infrastructure Costa Mesa, California, United States ABOUT THE TEAM
CorpTech Infrastructure Engineering builds and operates the foundational infrastructure that powers Company at large. We give engineers, researchers, and product teams across the company a place to deploy fast, scalable infrastructure without having to become infrastructure experts themselves. As Company's AI and autonomy ambitions grow, our team is responsible for delivering the next generation of compute, networking, and storage capabilities that make cutting edge model training and inference possible company wide. ABOUT THE JOB
We re looking for a Senior AI Infrastructure Engineer to lead the vision, execution, and long-term stability of how Company trains with GPUs at scale. In this role, you will take absolute ownership of cluster robustness, ensuring our high-performance GPU systems are highly available, fault-tolerant, and resilient for ML platform and research teams company-wide. This is a highly hands-on role where your primary focus is logical stability and automated resilience building self-healing mechanisms to proactively detect and isolate hardware faults, tuning NCCL and high-speed networking, and optimizing Kubernetes, Run:AI, and Ray scheduling. By replacing manual triage with automated deployment tooling and deep observability, you will ensure our massive-scale training infrastructure runs seamlessly and scales without linear headcount growth. WHAT YOU'LL DO
Rack, stack, cable, and bring up GPU compute (H200/B200/B300, NVL72) including physical topology, power, cooling, firmware/BIOS, and burn in validation.
Build and tune the interconnect fabric (NVLink, InfiniBand, RoCE, Spectrum-X) connecting hundreds of GPUs into low latency training and inference clusters.
Integrate high performance parallel storage (VAST, DDN, Weka) to sustain the throughput demanded by distributed training and terabyte scale multi modal datasets across Company's programs.
Automate cluster deployment and configuration end to end, including infrastructure as code for bring up, firmware/driver management, and fabric config, so new capacity comes online with minimal manual work.
Operate and extend our Kubernetes/Run:AI environment for GPU scheduling, quota management, and multi tenant workload isolation across research and engineering teams company wide.
Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults (bad transceivers, GPU Xid errors, NCCL/collective failures, RoCE congestion).
Onboard engineers and researchers onto the platform and act as their escalation point, working directly alongside them to debug, train, and optimize their workloads whenever infrastructure, not the model, is the bottleneck.
Partner with product facing teams across Company to understand emerging compute needs and translate them into platform capability. REQUIRED QUALIFICATIONS
10+ years in a hands on infrastructure, HPC, or datacenter engineering role supporting GPU compute at scale.
Hands on experience with H200/B200/B300 (or comparable) GPU systems: bring up, cabling, firmware/driver management.
Experience with high performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of hundreds of GPUs.
Experience with high performance parallel storage (VAST, DDN, Weka, Lustre, or similar).
Kubernetes required; Run:ai or similar GPU scheduling/orchestration experience strongly preferred.
Strong automation background. You build repeatable, automated deployment pipelines rather than manual processes.
Able to lift/move 50+ lbs and perform physical datacenter work (rack/stack/cable/troubleshoot).
Eligible to obtain and maintain an active U.S. Top Secret clearance. PREFERRED QUALIFICATIONS
Experience with NVIDIA NVL72 rack scale systems.
Experience supporting LLM token serving/inference infrastructure alongside training clusters.
Network fabric tuning experience (congestion control, adaptive routing, QoS) for RoCE/InfiniBand at scale.
Familiarity with GPU/network observability tooling (DCGM, fabric telemetry) and automated fault detection.
Experience supporting infrastructure as a shared platform serving multiple internal customer teams with differing requirements.
CorpTech Infrastructure Engineering builds and operates the foundational infrastructure that powers Company at large. We give engineers, researchers, and product teams across the company a place to deploy fast, scalable infrastructure without having to become infrastructure experts themselves. As Company's AI and autonomy ambitions grow, our team is responsible for delivering the next generation of compute, networking, and storage capabilities that make cutting edge model training and inference possible company wide. ABOUT THE JOB
We re looking for a Senior AI Infrastructure Engineer to lead the vision, execution, and long-term stability of how Company trains with GPUs at scale. In this role, you will take absolute ownership of cluster robustness, ensuring our high-performance GPU systems are highly available, fault-tolerant, and resilient for ML platform and research teams company-wide. This is a highly hands-on role where your primary focus is logical stability and automated resilience building self-healing mechanisms to proactively detect and isolate hardware faults, tuning NCCL and high-speed networking, and optimizing Kubernetes, Run:AI, and Ray scheduling. By replacing manual triage with automated deployment tooling and deep observability, you will ensure our massive-scale training infrastructure runs seamlessly and scales without linear headcount growth. WHAT YOU'LL DO
Rack, stack, cable, and bring up GPU compute (H200/B200/B300, NVL72) including physical topology, power, cooling, firmware/BIOS, and burn in validation.
Build and tune the interconnect fabric (NVLink, InfiniBand, RoCE, Spectrum-X) connecting hundreds of GPUs into low latency training and inference clusters.
Integrate high performance parallel storage (VAST, DDN, Weka) to sustain the throughput demanded by distributed training and terabyte scale multi modal datasets across Company's programs.
Automate cluster deployment and configuration end to end, including infrastructure as code for bring up, firmware/driver management, and fabric config, so new capacity comes online with minimal manual work.
Operate and extend our Kubernetes/Run:AI environment for GPU scheduling, quota management, and multi tenant workload isolation across research and engineering teams company wide.
Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults (bad transceivers, GPU Xid errors, NCCL/collective failures, RoCE congestion).
Onboard engineers and researchers onto the platform and act as their escalation point, working directly alongside them to debug, train, and optimize their workloads whenever infrastructure, not the model, is the bottleneck.
Partner with product facing teams across Company to understand emerging compute needs and translate them into platform capability. REQUIRED QUALIFICATIONS
10+ years in a hands on infrastructure, HPC, or datacenter engineering role supporting GPU compute at scale.
Hands on experience with H200/B200/B300 (or comparable) GPU systems: bring up, cabling, firmware/driver management.
Experience with high performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of hundreds of GPUs.
Experience with high performance parallel storage (VAST, DDN, Weka, Lustre, or similar).
Kubernetes required; Run:ai or similar GPU scheduling/orchestration experience strongly preferred.
Strong automation background. You build repeatable, automated deployment pipelines rather than manual processes.
Able to lift/move 50+ lbs and perform physical datacenter work (rack/stack/cable/troubleshoot).
Eligible to obtain and maintain an active U.S. Top Secret clearance. PREFERRED QUALIFICATIONS
Experience with NVIDIA NVL72 rack scale systems.
Experience supporting LLM token serving/inference infrastructure alongside training clusters.
Network fabric tuning experience (congestion control, adaptive routing, QoS) for RoCE/InfiniBand at scale.
Familiarity with GPU/network observability tooling (DCGM, fabric telemetry) and automated fault detection.
Experience supporting infrastructure as a shared platform serving multiple internal customer teams with differing requirements.
Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Sr. AI Infrastructure Engineer- Hybrid in Costa Mesa, CA vacancy
- ...Job Description Job Description We are hiring Sr. AI Infrastructure Engineer- Hybrid for a Full Time position in costa mesa, CA Senior AI Infrastructure Engineer, Physical Infrastructure Costa Mesa, California, United States ABOUT THE TEAM CorpTech Infrastructure...SeniorFull time
- Calance seeks a Senior AI Infrastructure Engineer to own the scalable GPU training pipeline, from bring-up to operation. You will design fault-tolerant... ...management, and fabric optimization while maintaining fault isolation and observability in a hybrid #J-18808-Ljbffr CalanceSenior
- Pacific Life is looking for a Senior Full Stack Engineer to join their Journey team in Newport Beach, CA. In this role, you will design... ...emphasis on collaboration and problem-solving. The position offers a hybrid work model and competitive benefits including medical, 401k...SeniorRemote job
$191k - $253k
...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands... ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems... ...JOB We are looking for a Senior AI Infrastructure Engineer to build, scale, and optimize the...SeniorFull timeWork experience placementImmediate start- Description#LI-DNIJob Title: Sr. Cloud Infrastructure ArchitectHyundai Capital America (HCA) helps... ...cloud solutions. This role partners with engineering, security, and product teams to... ...multi-account strategies, landing zones, hybrid cloud architectures, networking,...SeniorFull timeLocal areaImmediate startRemote workFlexible hours1 day per week
- ...Job Title: Sr. AWS Cloud Engineer Location: Irvine, CA (Onsite) Fulltime Roles & Responsibilities: Must Have Technical/Functional... ...with AWS migration tools, automation frameworks, and hybrid networking. Familiarity with integration to Oracle...SeniorFull time
$220k - $292k
...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands... ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems... ...We are looking for a founding Staff AI Infrastructure Engineer to architect, build, and scale...Full timeWork experience placementImmediate start$122k - $240.5k
Position Summary Agentic AI is moving from experimentation to production... ...and at scale. We're growing a team of engineers who want to work at the center of that... ...solutions in software, data, AI, network, and hybrid cloud infrastructure. These solutions are powered by...SeniorWork at officeLocal areaVisa sponsorshipShift work$55k - $151.47k
...Description & Summary At PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to... ...and Analytics Engineering team you will develop and implement AI solutions that enhance product offerings. As a Senior Associate...SeniorH1b- ...Sr. Infrastructure Engineer/DevOps AWS EngineerRemote PST6+ month contractWhat You’ll Be DoingCreate and modify Jenkins pipelines to support CI/CD of Kubernetes microservicesWork with Software Development teams to write and tune their application Helm charts for EKSCreate...Senior
- Sr. Cloud EngineerWho We AreSEW is the # 1 Energy and Water Cloud... ...Engagement, and Smart AI / Machine Learning to the Energy... ...talented team!.SummaryThe Cloud Engineer is responsible to design, install... ...three critical divisions: Infrastructure, Data Center and Security. This...SeniorWork experience placementWork at office
$146k - $194k
...of systems is powered by Lattice OS, an AI-powered operating system that turns... ...THE ROLEWe are seeking a Senior Network Infrastructure Engineer to join our Core Network Engineering team... ...Design, provision, and maintain resilient hybrid-cloud transit networks across AWS,...SeniorFull timeWork experience placementImmediate startRemote work- ...Description:We are seeking a highly motivated and self-driven AI Engineer with a strong background in building scalable Gen AI/Agentic AI... ...AI, Machine Learning, Unstructured Data Extraction using hybrid AI/ML, Prompt Engineering, Multi Agent Orcherstration, RAG, Model...Full timeWork experience placementLocal areaFlexible hours
$117.69k - $155k
...Glidewell Dental Network Engineer Position Position at Glidewell Dental Essential Functions: Analyzes, architects, designs,... ...Border Element (CUBE) routers, analog gateways, and any other infrastructure equipment. Works with carriers and configures/troubleshoots...SeniorRelocation- ...our mission of innovating our business and creating superior customer experience. We’re actively seeking a talented Sr. Infrastructure Platform Engineer - Data Center to join our Data Center Services team in Newport Beach, CA. This role is on-site.As a Sr. Infrastructure...SeniorFull timeWork at officeFlexible hours
$137.61k - $168.19k
...into the future. We are seeking a talented Senior Platform Engineer to join our Enterprise Technology team. This role will be responsible... ..., you will collaborate with application teams, architects, infrastructure engineers, and business stakeholders to deliver scalable,...SeniorWork experience placementWork at officeFlexible hours$172.55k - $249.55k
...Responsibilities:Architect and design hybrid cloud solutions that integrate on-premises... ...the implementation of hybrid cloud infrastructures, including provisioning resources, configuring... ...mentorship to junior architects and engineers on hybrid cloud best...SeniorPermanent employmentFull timeInterim roleLocal areaVisa sponsorshipWork visaRelocation packageFlexible hoursShift work$75 - $100 per hour
...experienced Python Software Engineer to help modernize and enhance... ...modern web technologies, cloud infrastructure, and advanced analytics.What... ...and long-term project roadmap Hybrid work flexibilityIf you're... ...of Artificial Intelligence (AI): We may use Artificial Intelligence...SeniorContract workTemporary work$128.6k - $160.8k
...Sr. Platform Engineer Shaping the Future of Lending With 30 years of fintech leadership, Origence... ...future of lending, we're leveraging AI to enhance the way we work and create... ...tooling applied to platform and infrastructure problems. You will work across CI/CD pipeline...SeniorFull timeLocal areaFlexible hours$131k - $271.6k
...building specialized foundation models and AI agents that accelerate SAP customers'... ...scale multi-agent systems, and raise the engineering bar across a global team. You'll work directly... ...: Professional | Employment Type: Regular Full Time | Additional Locations: #LI-Hybrid...SeniorPermanent employmentFull timeWorldwideFlexible hours- ...financing freedom of movement. Apply today.WHAT YOU WILL DOThe Sr. Application Developer Associate, Insurance Platform supports auto... ...Bachelor’s degree in Computer Science, Information Technology, Engineering or a related field.Relevant certifications preferred, but not...SeniorFull timeContract workWork at officeLocal areaImmediate start
$166k - $220k
...is powered by Lattice OS, an AI-powered operating system that... ...TEAMInfrastructure Reliability Engineering (IRE) is a small but growing team responsible for the infrastructure and operations behind the... ...Grafana)Background in SRE or hybrid SWE/DevOps rolesExperience with...SeniorFull timeWork experience placementImmediate start$150k - $200k
...Design and implement compute and network infrastructure capabilities on AWS, including VPC,... ...Collaborate closely with application engineering, architecture, and platform teams to standardize... ...network platforms in multi-cloud or hybrid-cloud environments. Experience...SeniorLocal areaWorldwide$95k - $115k
...is becoming genuinely intelligent — AI-driven scheduling, computer-vision... ...GovCloud. This role builds and runs that infrastructure. This is a hands-on engineering role, not a systems administration... ..., and real-time monitoring across hybrid cloud and on-premise environments...Permanent employmentFull timeLocal areaRelocationFlexible hours$122k - $240.5k
Position Summary Join our AI & Engineering team in transforming technology platforms, driving innovation, and helping make... ...verticalized sector solutions in software, data, AI, network, and hybrid cloud infrastructure. These solutions are powered by engineering for business...Local areaVisa sponsorship$130k - $178k
...variety of technical tasks for Infrastructure and Operations relating Azure... .... Conducts technical engineering and operational responsibilities... ...scale, production Azure and hybrid environments. ~Experience with... ...Cisco UCS and Private Cloud/AI infrastructure Why work for...SeniorWork at officeLocal area$135k - $150k
...Job Description Job Description 10778 - Sr. Network Engineer Irvine, CA 92614 (5 days on-site) Company Overview Hyundai AutoEver... ...position requires in-depth knowledge of HAEA’s Network Infrastructure, (LAN/Switches, DNS – DHCP, WAN, Load Balancers, VOIP, VPN...SeniorContract workWork experience placementLocal areaRemote work$150k - $180k
...Overview: The Senior Cloud Security Engineer reduces material risk across large-... ...role. You will use Python, APIs, infrastructure as code, and modern AI-assisted engineering tools to... ...across AWS, Azure, Kubernetes, and hybrid infrastructure. Use CSPM/CNAPP telemetry...SeniorLocal areaWorldwide$112.2k - $190.7k
...to drive innovation and efficiency across the IT ecosystem. Our AI Center of Excellence (CoE) is a dynamic and strategic group dedicated... ....Your role:We are seeking a talented and motivated Agentic AI Engineer to join our AI CoE. In this hands-on role, you will be...SeniorFull timeTemporary workWorldwide- ...Job Title: : Senior AI Engineer Lo cation: Irvine, CA Type: Full Time Job Description experience... ...Backend: APIs, microservices (e.g., Spring Boot, Node.js) DevOps: Docker, Kubernetes, CI/CD, Infrastructure as Code...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. AI Infrastructure Engineer- Hybrid. Be the first to apply!
Related searches
- infrastructure engineer Costa Mesa, CA
- infrastructure developer Costa Mesa, CA
- remote infrastructure engineer Costa Mesa, CA
- senior infrastructure engineer Costa Mesa, CA
- senior operations technician Costa Mesa, CA
- senior cloud service delivery manager Costa Mesa, CA
- sr accountant Costa Mesa, CA
- senior financial analyst remote Costa Mesa, CA
- senior designer Costa Mesa, CA
- senior manager accounts payable Costa Mesa, CA



