Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Network Engineer - ML Infrastructure (High-Speed Interconnects)

$180k

SpaceXAI

Job Description

Job Description

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

SpaceXAI is building at a furious pace with the latest compute and switching hardware to help people understand the universe. We are looking for exceptional ML Infrastructure Engineers with deep expertise in high-speed interconnect technologies to design, build, and optimize the network fabric that powers large-scale AI training and inference clusters. This strategic role will drive innovation in high-bandwidth, low-latency, power-efficient interconnects critical for AI/ML clusters based on advanced computing platforms.

You will have the opportunity to work on all modalities of interconnects connecting GPUs and switches both inside and between data centers, including our primary front and backend networks that train Grok and that customers use for inference. Engineers will own all aspects from design and development to build and operations. You will be expected to define and improve team processes and to contribute to scaling and maintenance efforts.

You will focus on the physical layer and system-level integration of copper (ACC, AEC, CPC) and optical (FRO, LRO/TRO, LPO, AOC, CPO) interconnects that directly determine the performance, power efficiency, scale, and cost of next-generation AI/ML clusters. This is a highly technical, hands-on role bridging ML cluster requirements with cutting-edge interconnect hardware — ideal for engineers who love both large-scale AI systems and the physics/engineering of 200G+ SerDes, PAM4, photonics, signal integrity and diagnostics.

RESPONSIBILITIES:
  • Design, validate, and productize high-speed copper and optical connectivity solutions for AI clusters (100k+ GPU scale).
  • Own vendor due diligence and onboarding for new 1.6T products including AEC and pluggable optical transceivers (DR4/8, FR4) including rigorous bring-up & characterization.
  • Investigate the opportunity for LPO and LRO in our network.
  • Evaluate early co-packaged and near-packaged engines for switches and GPUs.
  • Pathfinding for new interconnect modalities including VCSEL, microLED, THz radio-based solutions to improve network economics and reliability.
  • Work closely with vendors (transceiver, cable, SerDes, DSP, silicon photonics foundries) to influence roadmaps and ensure timely delivery of next-gen solutions.
  • Collaborate with ML training teams to translate workload communication patterns into concrete interconnect topology and optical reconfigurability requirements.
  • Perform system-level simulation of end-to-end fabric performance.
  • Drive failure analysis, root cause, and corrective actions for interconnect-related issues in production clusters through fleet-level metrics gathering and analysis.
  • Contribute to internal tooling and automation for interconnect health monitoring, telemetry, diagnostics, remediation and automated qualification pipelines.
  • Stay current with industry standards (OIF CMIS, IEEE) and emerging technologies (multi-core/hollow-core fiber, 448G SerDes, TFLN, ring resonators)
BASIC QUALIFICATIONS:
  • At least 8+ years of hands-on experience in designing, deploying and operating high-speed copper and optical interconnects, preferably in a module design role or in a hyperscale datacenter environment.
  • Master's or PhD degree in Electrical Engineering, Photonics or Physics.
  • Deep knowledge of PAM4 SerDes performance, equalization, jitter, crosstalk.
  • Solid operational understanding of FEC, Retimers, TIAs and Drivers.
  • Deep knowledge of optical link budget analysis and performance metrics including TDECQ, OMA, Tcode, stressed receiver sensitivity and associated diagnostics.
  • Expertise in transceiver components including CW lasers, SiPh PICs, EML, DSP, passive subassemblies, their failure modes and characterization.
  • Knowledge of thermal, mechanical, power, signal integrity constraints in dense hardware.
  • Knowledge of SiPh design process, yield improvement and reliability testing.
  • Familiarity with CPO technologies and challenges/risk areas.
  • Familiarity with subcomponent supply chains and global manufacturers, ODMs and CMs.
  • Strong problem-solving skills and ability to thrive in a fast-paced, ambiguous setting.
COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Network Engineer - ML Infrastructure (High-Speed Interconnects) in Palo Alto, CA vacancy
  • $262k - $364k

     ...and coach a distributed engineering team, fostering...  ...reliability.Performance and Network Engineering: Architect...  ...HPC, RDMA, and ML workloads.Strategic Roadmap...  ...products.Experience with C, High Performance Computing,...  ...or Machine Learning Infrastructure.Google's software... 
    Suggested
    Remote work
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $163k - $237k

     ...execution to deliver high-performance network design components...  ...degree in Electrical Engineering, Computer...  ...Experience with high-speed interconnects.Experience developing...  ...shape the future of AI/ML hardware acceleration...  ....The AI and Infrastructure team is redefining... 
    Suggested
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $127.1k - $185k

     ...talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll...  ...intersection of high-performance computing...  ...machine learning infrastructure - building the...  ...networking, or RDMA/high-speed interconnects- Exposure to... 
    Suggested
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $190k - $260k

     ...as good as the speed at which we...  ...– depends on infrastructure that turns thousands...  ...looking for engineers who make model...  ...: Design high-throughput data...  ...across storage, network, CPU...  ...Partner with ML teams to scale...  ...Profiler), and interconnects (NVLink, InfiniBand... 
    Suggested
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    a month ago
  • $150k - $230k

     ...researchers and veteran systems engineers who share a vision for...  ...complex, traditional infrastructure struggles to meet the...  ...for fault-tolerant, high-performance...  ...of GPU systems, high-speed networking, and distributed coordination...  ...(RDMA, InfiniBand) ML framework or runtime internals... 
    Suggested

    Clockwork.io

    Palo Alto, CA
    18 days ago
  • $159k - $230k

     ...failure reproduction.Design network, power and cooling infrastructure for end-to-end system...  ...designers, test engineers, on project planning within...  ...tests of end-to-end highly complex ML system design...  ...specialized ML ASICS and high-speed interconnects. You will also be... 

    Google

    Sunnyvale, CA
    3 days ago
  • $193.93k - $291.15k

     ...approach to autonomous driving, and the ML Infrastructure team builds and operates the...  ...streaming ingestion, storage layout, and high-throughput data generation and storage....  ...or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 3+ years... 
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    27 days ago
  • $193.3k - $261.5k

    We are seeking an experienced engineer to work on distributed AI/ML systems. This role involves...  ...valued, and experience with high-speed networking or HPC interconnects is valued highly.If you like...  ...critical building blocks for EC2 infrastructure. Every instance in EC2 is... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $111.07k - $166.4k

     ...building blocks of the data infrastructure that connects our...  ..., Your ImpactHigh-speed interconnects — from switch and NIC...  ...data centers and AI/ML fabrics move data at...  .... As a member of the Network Validation / Interoperability...  ...-generation, ultra-high-speed networking gear... 
    Permanent employment
    Full time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    4 days ago
  • $148.2k - $222.2k

    Staff Engineer, Infrastructure Platforms Position SummaryThe Staff Engineer, Infrastructure...  ..., enterprise storage, High Performance Computing (HPC),...  ...storage, virtualization, networking, and High Performance...  ...genomics, bioinformatics, AI/ML, or other scientific computing... 
    Full time
    Work from home
    Monday to Friday

    Pacific Biosciences

    Menlo Park, CA
    3 days ago
  • $198k - $326k

     ...largest professional network, built to create...  ...signals throughout the ML lifecycle, process high volumes of model and...  ...must operate at the speed and scale of LinkedIn...  ..., and enforceable engineering capabilities.As a Sr...  ...generation of AI Governance infrastructure. You will solve... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    2 days ago
  •  ...ROLE:We are seeking a highly skilled Director of...  ...and execution of our network infrastructure products. This role is...  ...GPU clusters for AI, ML, and HPC workloads.THE...  ....Collaboration with Engineering Teams: Work closely with...  ...data transfer speeds, network congestion,... 
    Remote work

    AMD

    Santa Clara, CA
    2 days ago
  •  ...training and inference speeds; over 10 times...  ...with ultra high-speed inference.About...  ...the TeamThe Core Infrastructure team builds the software...  ...that power engineering workflows across Cerebras...  ..., filesystems, networking, remote machines,...  ...release, quality, ML systems, and product... 
    Remote work

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...experienced and strategic Senior Manager of Network Engineering to lead the team responsible for the...  .... This position manages our network infrastructure across multiple global sites, data...  ...business deadlines. You will guide a highly capable engineering team while heavily... 
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    14 days ago
  • $180k

     ...of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is...  .... ABOUT THE ROLE: As an ML Infrastructure Engineer, you will play a pivotal...  ...NVIDIA drivers, CUDA toolkits, and networking Comfortable with Linux... 
    Temporary work
    Work experience placement

    SpaceXAI

    Palo Alto, CA
    17 days ago
  •  ...AI Infrastructure Engineer at HOAiHOAi is a fast-growing startup revolutionizing...  ...serving infrastructure, and ML pipelines that power HOAi's productsPerformance...  ...-minded: Balances speed with quality; ships...  ...iteratesBar raiser: Sets and maintains high standards; elevates team... 

    Vantaca

    Redwood City, CA
    5 hours ago
  • $183k - $275k

     ...autonomy code change, from ML model updates to...  ...road. You will own the infrastructure that makes this possible...  ...much much more. Engineers across the company rely...  ...onboard logging rate or network contention as more...  ...visualization layer. This is a high-ownership, high-... 
    Temporary work
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    18 days ago
  • $188.5k - $255k

     ...tripling the speed at which developers...  ...infrastructure, operational intelligence...  ...developer, or engineering leader gets...  ...Consumption, AI/ML Enablement,...  ...Abstraction & Network Infrastructure...  ...across three interconnected areas: enterprise...  ...)Convert high-ambiguity charters... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    3 days ago
  • $165.2k - $223.6k

    We are seeking an experienced engineer to work on distributed AI/ML systems. This role involves...  ...valued, and experience with high-speed networking or HPC interconnects is valued highly.If you like...  ...critical building blocks for EC2 infrastructure. Every instance in EC2 is... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $193.93k - $352.29k

     ...investors. About the Role The Autonomy ML Infrastructure team is responsible for building &...  ...compression.  Work with autonomy engineers to optimize, validate, and deploy large...  ...compiler framework, FTL. Write robust, high quality software to increase our confidence... 
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    23 days ago
  • $207k - $340k

     ...largest professional network, built to create economic...  ...Staff Software Engineer to lead LinkedIn’s GPU...  ...Platform, a foundational AI infrastructure stack that powers...  ...scalable, observable, and highly available multi-tenant...  ...Suggested Skills:AI / ML Infrastructure Technical... 
    For contractors
    Work at office
    Remote work
    Work from home
    Flexible hours

    Linkedin

    Mountain View, CA
    23 hours ago
  • $124k - $250k

     ...of others.A Day in the LifeAs a member of our software engineering infra team, you'll solve technical challenges,...  ...upgrading and implementing state-of-the-art software infrastructure. The team builds a high-performance, high availability, globally distributed ecosystem... 

    AppLovin

    Palo Alto, CA
    4 days ago
  • $1,000 - $2,030 per month

     ...Work on the Machine Learning and AI Infrastructure team, which is focused on providing infrastructure...  ...for data scientists and other engineers to use ML and AI technologies effectively....  ...and have experience building scalable high-performance systems ~ Familiarity with... 
    Full time
    Temporary work
    Work at office
    Flexible hours

    Cloudkitchens

    Mountain View, CA
    23 hours ago
  • $193.3k - $261.5k

     ...seeking an experienced engineer and technical...  ...team that owns the network stack for EC2 distributed AI/ML systems. The team develops...  ...experience with high-speed networking or HPC/RDMA interconnects is highly valued.If...  ...blocks for EC2 infrastructure. Every instance in... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $180k - $250k

     ...blog posts sharing our high-level results for text...  ...research and data engineering necessary to solve this...  ...an experienced Cloud Infrastructure Engineer to join our core...  ...training large-scale ML models Drive cost-efficiency...  ...-level debugging—networking issues, memory leaks,... 
    Work at office
    Relocation package

    Datology

    Redwood City, CA
    3 days ago
  • $150.32k - $225.48k

     ...operations, systems and safety engineering - all dedicated to making a...  ...automated driving technology at the speed of a technology startup....  ...team: The Onboard Platform - High Level OS team is an embedded...  ...of software tools and infrastructure that improve developer experience... 
    Permanent employment
    Full time
    Work at office
    Immediate start
    Visa sponsorship

    Socket

    Palo Alto, CA
    4 days ago
  • $207k - $300k

    Design infrastructure for scalable usage of collected research...  .... Investigate new ML tools and agent platforms...  ...’s degree or PhD in Engineering, Computer Science, or...  ...-scale system design, networking and data storage, security...  ...daily lives.We are a highly collaborative team.... 

    Google

    Mountain View, CA
    23 hours ago
  • $157k - $235k

     ...other digital services.Snap Engineering teams build fun and technically...  ...team, part of the Content ML organization. We develop and...  ...you’ll do:Design and optimize infrastructure systems for machine learning...  ...training data generationDevelop high-performance inference systems... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    23 hours ago
  • $184k - $287.5k

     ...is seeking elite ASIC Infrastructure engineers to deliver the tooling...  ...all HW Design teams Use ML/DL/AI techniques to...  ...productivity Improve the speed, flexibility and...  ...compute farm, filer, and network topology requirements...  ...opportunity employer. As we highly value diversity in our... 
    Full time
    Work experience placement
    Work at office
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-...  ...-driven validation to high-scale, automated cloud validation...  ...function-to-function networking, gRPC bottlenecks, and in-cluster...  ..., preferably involving AI/ML inference, computer vision,... 
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Network Engineer - ML Infrastructure (High-Speed Interconnects). Be the first to apply!