Network Engineer - ML Infrastructure (High-Speed Interconnects)
$180kSpaceXAI
Job Description
Job Description
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:SpaceXAI is building at a furious pace with the latest compute and switching hardware to help people understand the universe. We are looking for exceptional ML Infrastructure Engineers with deep expertise in high-speed interconnect technologies to design, build, and optimize the network fabric that powers large-scale AI training and inference clusters. This strategic role will drive innovation in high-bandwidth, low-latency, power-efficient interconnects critical for AI/ML clusters based on advanced computing platforms.
You will have the opportunity to work on all modalities of interconnects connecting GPUs and switches both inside and between data centers, including our primary front and backend networks that train Grok and that customers use for inference. Engineers will own all aspects from design and development to build and operations. You will be expected to define and improve team processes and to contribute to scaling and maintenance efforts.
You will focus on the physical layer and system-level integration of copper (ACC, AEC, CPC) and optical (FRO, LRO/TRO, LPO, AOC, CPO) interconnects that directly determine the performance, power efficiency, scale, and cost of next-generation AI/ML clusters. This is a highly technical, hands-on role bridging ML cluster requirements with cutting-edge interconnect hardware — ideal for engineers who love both large-scale AI systems and the physics/engineering of 200G+ SerDes, PAM4, photonics, signal integrity and diagnostics.
RESPONSIBILITIES:- Design, validate, and productize high-speed copper and optical connectivity solutions for AI clusters (100k+ GPU scale).
- Own vendor due diligence and onboarding for new 1.6T products including AEC and pluggable optical transceivers (DR4/8, FR4) including rigorous bring-up & characterization.
- Investigate the opportunity for LPO and LRO in our network.
- Evaluate early co-packaged and near-packaged engines for switches and GPUs.
- Pathfinding for new interconnect modalities including VCSEL, microLED, THz radio-based solutions to improve network economics and reliability.
- Work closely with vendors (transceiver, cable, SerDes, DSP, silicon photonics foundries) to influence roadmaps and ensure timely delivery of next-gen solutions.
- Collaborate with ML training teams to translate workload communication patterns into concrete interconnect topology and optical reconfigurability requirements.
- Perform system-level simulation of end-to-end fabric performance.
- Drive failure analysis, root cause, and corrective actions for interconnect-related issues in production clusters through fleet-level metrics gathering and analysis.
- Contribute to internal tooling and automation for interconnect health monitoring, telemetry, diagnostics, remediation and automated qualification pipelines.
- Stay current with industry standards (OIF CMIS, IEEE) and emerging technologies (multi-core/hollow-core fiber, 448G SerDes, TFLN, ring resonators)
- At least 8+ years of hands-on experience in designing, deploying and operating high-speed copper and optical interconnects, preferably in a module design role or in a hyperscale datacenter environment.
- Master's or PhD degree in Electrical Engineering, Photonics or Physics.
- Deep knowledge of PAM4 SerDes performance, equalization, jitter, crosstalk.
- Solid operational understanding of FEC, Retimers, TIAs and Drivers.
- Deep knowledge of optical link budget analysis and performance metrics including TDECQ, OMA, Tcode, stressed receiver sensitivity and associated diagnostics.
- Expertise in transceiver components including CW lasers, SiPh PICs, EML, DSP, passive subassemblies, their failure modes and characterization.
- Knowledge of thermal, mechanical, power, signal integrity constraints in dense hardware.
- Knowledge of SiPh design process, yield improvement and reliability testing.
- Familiarity with CPO technologies and challenges/risk areas.
- Familiarity with subcomponent supply chains and global manufacturers, ODMs and CMs.
- Strong problem-solving skills and ability to thrive in a fast-paced, ambiguous setting.
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
$262k - $364k
...and coach a distributed engineering team, fostering... ...reliability.Performance and Network Engineering: Architect... ...HPC, RDMA, and ML workloads.Strategic Roadmap... ...products.Experience with C, High Performance Computing,... ...or Machine Learning Infrastructure.Google's software...SuggestedRemote workWorldwide$163k - $237k
...execution to deliver high-performance network design components... ...degree in Electrical Engineering, Computer... ...Experience with high-speed interconnects.Experience developing... ...shape the future of AI/ML hardware acceleration... ....The AI and Infrastructure team is redefining...SuggestedWorldwide$127.1k - $185k
...talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll... ...intersection of high-performance computing... ...machine learning infrastructure - building the... ...networking, or RDMA/high-speed interconnects- Exposure to...SuggestedInternshipLocal areaFlexible hours$190k - $260k
...as good as the speed at which we... ...– depends on infrastructure that turns thousands... ...looking for engineers who make model... ...: Design high-throughput data... ...across storage, network, CPU... ...Partner with ML teams to scale... ...Profiler), and interconnects (NVLink, InfiniBand...SuggestedTemporary workWork at officeVisa sponsorshipFlexible hours$150k - $230k
...researchers and veteran systems engineers who share a vision for... ...complex, traditional infrastructure struggles to meet the... ...for fault-tolerant, high-performance... ...of GPU systems, high-speed networking, and distributed coordination... ...(RDMA, InfiniBand) ML framework or runtime internals...Suggested$159k - $230k
...failure reproduction.Design network, power and cooling infrastructure for end-to-end system... ...designers, test engineers, on project planning within... ...tests of end-to-end highly complex ML system design... ...specialized ML ASICS and high-speed interconnects. You will also be...$193.93k - $291.15k
...approach to autonomous driving, and the ML Infrastructure team builds and operates the... ...streaming ingestion, storage layout, and high-throughput data generation and storage.... ...or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 3+ years...Work experience placementImmediate startFlexible hours$193.3k - $261.5k
We are seeking an experienced engineer to work on distributed AI/ML systems. This role involves... ...valued, and experience with high-speed networking or HPC interconnects is valued highly.If you like... ...critical building blocks for EC2 infrastructure. Every instance in EC2 is...InternshipLocal areaFlexible hours$111.07k - $166.4k
...building blocks of the data infrastructure that connects our... ..., Your ImpactHigh-speed interconnects — from switch and NIC... ...data centers and AI/ML fabrics move data at... .... As a member of the Network Validation / Interoperability... ...-generation, ultra-high-speed networking gear...Permanent employmentFull timeInternshipWork from home$148.2k - $222.2k
Staff Engineer, Infrastructure Platforms Position SummaryThe Staff Engineer, Infrastructure... ..., enterprise storage, High Performance Computing (HPC),... ...storage, virtualization, networking, and High Performance... ...genomics, bioinformatics, AI/ML, or other scientific computing...Full timeWork from homeMonday to Friday$198k - $326k
...largest professional network, built to create... ...signals throughout the ML lifecycle, process high volumes of model and... ...must operate at the speed and scale of LinkedIn... ..., and enforceable engineering capabilities.As a Sr... ...generation of AI Governance infrastructure. You will solve...For contractorsWork at officeFlexible hours- ...ROLE:We are seeking a highly skilled Director of... ...and execution of our network infrastructure products. This role is... ...GPU clusters for AI, ML, and HPC workloads.THE... ....Collaboration with Engineering Teams: Work closely with... ...data transfer speeds, network congestion,...Remote work
- ...training and inference speeds; over 10 times... ...with ultra high-speed inference.About... ...the TeamThe Core Infrastructure team builds the software... ...that power engineering workflows across Cerebras... ..., filesystems, networking, remote machines,... ...release, quality, ML systems, and product...Remote work
- ...experienced and strategic Senior Manager of Network Engineering to lead the team responsible for the... .... This position manages our network infrastructure across multiple global sites, data... ...business deadlines. You will guide a highly capable engineering team while heavily...Temporary workRelocation package
$180k
...of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is... .... ABOUT THE ROLE: As an ML Infrastructure Engineer, you will play a pivotal... ...NVIDIA drivers, CUDA toolkits, and networking Comfortable with Linux...Temporary workWork experience placement- ...AI Infrastructure Engineer at HOAiHOAi is a fast-growing startup revolutionizing... ...serving infrastructure, and ML pipelines that power HOAi's productsPerformance... ...-minded: Balances speed with quality; ships... ...iteratesBar raiser: Sets and maintains high standards; elevates team...
$183k - $275k
...autonomy code change, from ML model updates to... ...road. You will own the infrastructure that makes this possible... ...much much more. Engineers across the company rely... ...onboard logging rate or network contention as more... ...visualization layer. This is a high-ownership, high-...Temporary workImmediate startFlexible hours$188.5k - $255k
...tripling the speed at which developers... ...— infrastructure, operational intelligence... ...developer, or engineering leader gets... ...Consumption, AI/ML Enablement,... ...Abstraction & Network Infrastructure... ...across three interconnected areas: enterprise... ...)Convert high-ambiguity charters...WorldwideShift work$165.2k - $223.6k
We are seeking an experienced engineer to work on distributed AI/ML systems. This role involves... ...valued, and experience with high-speed networking or HPC interconnects is valued highly.If you like... ...critical building blocks for EC2 infrastructure. Every instance in EC2 is...InternshipLocal areaFlexible hours$193.93k - $352.29k
...investors. About the Role The Autonomy ML Infrastructure team is responsible for building &... ...compression. Work with autonomy engineers to optimize, validate, and deploy large... ...compiler framework, FTL. Write robust, high quality software to increase our confidence...Work experience placementImmediate startFlexible hours$207k - $340k
...largest professional network, built to create economic... ...Staff Software Engineer to lead LinkedIn’s GPU... ...Platform, a foundational AI infrastructure stack that powers... ...scalable, observable, and highly available multi-tenant... ...Suggested Skills:AI / ML Infrastructure Technical...For contractorsWork at officeRemote workWork from homeFlexible hours$124k - $250k
...of others.A Day in the LifeAs a member of our software engineering infra team, you'll solve technical challenges,... ...upgrading and implementing state-of-the-art software infrastructure. The team builds a high-performance, high availability, globally distributed ecosystem...$1,000 - $2,030 per month
...Work on the Machine Learning and AI Infrastructure team, which is focused on providing infrastructure... ...for data scientists and other engineers to use ML and AI technologies effectively.... ...and have experience building scalable high-performance systems ~ Familiarity with...Full timeTemporary workWork at officeFlexible hours$193.3k - $261.5k
...seeking an experienced engineer and technical... ...team that owns the network stack for EC2 distributed AI/ML systems. The team develops... ...experience with high-speed networking or HPC/RDMA interconnects is highly valued.If... ...blocks for EC2 infrastructure. Every instance in...Local areaFlexible hours$180k - $250k
...blog posts sharing our high-level results for text... ...research and data engineering necessary to solve this... ...an experienced Cloud Infrastructure Engineer to join our core... ...training large-scale ML models Drive cost-efficiency... ...-level debugging—networking issues, memory leaks,...Work at officeRelocation package$150.32k - $225.48k
...operations, systems and safety engineering - all dedicated to making a... ...automated driving technology at the speed of a technology startup.... ...team: The Onboard Platform - High Level OS team is an embedded... ...of software tools and infrastructure that improve developer experience...Permanent employmentFull timeWork at officeImmediate startVisa sponsorship$207k - $300k
Design infrastructure for scalable usage of collected research... .... Investigate new ML tools and agent platforms... ...’s degree or PhD in Engineering, Computer Science, or... ...-scale system design, networking and data storage, security... ...daily lives.We are a highly collaborative team....$157k - $235k
...other digital services.Snap Engineering teams build fun and technically... ...team, part of the Content ML organization. We develop and... ...you’ll do:Design and optimize infrastructure systems for machine learning... ...training data generationDevelop high-performance inference systems...Full timeLive inWork at officeLocal area$184k - $287.5k
...is seeking elite ASIC Infrastructure engineers to deliver the tooling... ...all HW Design teams Use ML/DL/AI techniques to... ...productivity Improve the speed, flexibility and... ...compute farm, filer, and network topology requirements... ...opportunity employer. As we highly value diversity in our...Full timeWork experience placementWork at officeWorldwide$184k - $287.5k
...Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-... ...-driven validation to high-scale, automated cloud validation... ...function-to-function networking, gRPC bottlenecks, and in-cluster... ..., preferably involving AI/ML inference, computer vision,...Full timeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Network Engineer - ML Infrastructure (High-Speed Interconnects). Be the first to apply!
- senior network engineer remote Palo Alto, CA
- network engineer - transport Palo Alto, CA
- network developer Palo Alto, CA
- remote cisco network engineer Palo Alto, CA
- network software engineer Palo Alto, CA
- network applications engineer Palo Alto, CA
- network infrastructure engineer Palo Alto, CA
- cisco network engineer Palo Alto, CA
- network engineer level Palo Alto, CA
- data center network engineer Palo Alto, CA




