Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Network Engineer - ML Infrastructure (High-Speed Interconnects)

$180k

SpaceXAI

Job Description

Job Description

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

SpaceXAI is building at a furious pace with the latest compute and switching hardware to help people understand the universe. We are looking for exceptional ML Infrastructure Engineers with deep expertise in high-speed interconnect technologies to design, build, and optimize the network fabric that powers large-scale AI training and inference clusters. This strategic role will drive innovation in high-bandwidth, low-latency, power-efficient interconnects critical for AI/ML clusters based on advanced computing platforms.

You will have the opportunity to work on all modalities of interconnects connecting GPUs and switches both inside and between data centers, including our primary front and backend networks that train Grok and that customers use for inference. Engineers will own all aspects from design and development to build and operations. You will be expected to define and improve team processes and to contribute to scaling and maintenance efforts.

You will focus on the physical layer and system-level integration of copper (ACC, AEC, CPC) and optical (FRO, LRO/TRO, LPO, AOC, CPO) interconnects that directly determine the performance, power efficiency, scale, and cost of next-generation AI/ML clusters. This is a highly technical, hands-on role bridging ML cluster requirements with cutting-edge interconnect hardware — ideal for engineers who love both large-scale AI systems and the physics/engineering of 200G+ SerDes, PAM4, photonics, signal integrity and diagnostics.

RESPONSIBILITIES:
  • Design, validate, and productize high-speed copper and optical connectivity solutions for AI clusters (100k+ GPU scale).
  • Own vendor due diligence and onboarding for new 1.6T products including AEC and pluggable optical transceivers (DR4/8, FR4) including rigorous bring-up & characterization.
  • Investigate the opportunity for LPO and LRO in our network.
  • Evaluate early co-packaged and near-packaged engines for switches and GPUs.
  • Pathfinding for new interconnect modalities including VCSEL, microLED, THz radio-based solutions to improve network economics and reliability.
  • Work closely with vendors (transceiver, cable, SerDes, DSP, silicon photonics foundries) to influence roadmaps and ensure timely delivery of next-gen solutions.
  • Collaborate with ML training teams to translate workload communication patterns into concrete interconnect topology and optical reconfigurability requirements.
  • Perform system-level simulation of end-to-end fabric performance.
  • Drive failure analysis, root cause, and corrective actions for interconnect-related issues in production clusters through fleet-level metrics gathering and analysis.
  • Contribute to internal tooling and automation for interconnect health monitoring, telemetry, diagnostics, remediation and automated qualification pipelines.
  • Stay current with industry standards (OIF CMIS, IEEE) and emerging technologies (multi-core/hollow-core fiber, 448G SerDes, TFLN, ring resonators)
BASIC QUALIFICATIONS:
  • At least 8+ years of hands-on experience in designing, deploying and operating high-speed copper and optical interconnects, preferably in a module design role or in a hyperscale datacenter environment.
  • Master's or PhD degree in Electrical Engineering, Photonics or Physics.
  • Deep knowledge of PAM4 SerDes performance, equalization, jitter, crosstalk.
  • Solid operational understanding of FEC, Retimers, TIAs and Drivers.
  • Deep knowledge of optical link budget analysis and performance metrics including TDECQ, OMA, Tcode, stressed receiver sensitivity and associated diagnostics.
  • Expertise in transceiver components including CW lasers, SiPh PICs, EML, DSP, passive subassemblies, their failure modes and characterization.
  • Knowledge of thermal, mechanical, power, signal integrity constraints in dense hardware.
  • Knowledge of SiPh design process, yield improvement and reliability testing.
  • Familiarity with CPO technologies and challenges/risk areas.
  • Familiarity with subcomponent supply chains and global manufacturers, ODMs and CMs.
  • Strong problem-solving skills and ability to thrive in a fast-paced, ambiguous setting.
COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Network Engineer - ML Infrastructure (High-Speed Interconnects) in Palo Alto, CA vacancy
  • $180k

     ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This...  ...looking for exceptional ML Infrastructure Engineers with deep expertise in high-speed interconnect technologies to design,...  ..., and optimize the network fabric that powers large... 
    Suggested
    Temporary work

    SpaceXAI

    Palo Alto, CA
    25 days ago
  • $262k - $364k

     ...and coach a distributed engineering team, fostering...  ...reliability.Performance and Network Engineering: Architect...  ...HPC, RDMA, and ML workloads.Strategic Roadmap...  ...products.Experience with C, High Performance Computing,...  ...or Machine Learning Infrastructure.Google's software... 
    Suggested
    Remote work
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  • $192k - $278k

     ...firmware to control high-speed transceivers, managing...  ...Partner with TPU Systems, Networking Software, and High-...  ...impact Google's AI infrastructure.Minimum...  ...degree in Electrical Engineering, Computer Engineering...  ...shape the future of AI/ML hardware acceleration... 
    Suggested
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  • $163k - $237k

     ...execution to deliver high-performance network design components...  ...degree in Electrical Engineering, Computer...  ...Experience with high-speed interconnects.Experience developing...  ...shape the future of AI/ML hardware acceleration...  ....The AI and Infrastructure team is redefining... 
    Suggested
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • $127.1k - $185k

     ...talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll...  ...intersection of high-performance computing...  ...machine learning infrastructure - building the...  ...networking, or RDMA/high-speed interconnects - Exposure to... 
    Suggested
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  •  ...and inference speeds; over 10 times...  ...workloads with ultra high-speed...  ...looking for a Network Architect to join...  ...our Cluster Engineering Team and help...  ...datacenter and interconnect fabric for the...  ...fabrics for AI/ML and HPC clusters...  ...of network infrastructure using Python,... 

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $190k - $260k

     ...as good as the speed at which we...  ...- depends on infrastructure that turns thousands...  ...looking for engineers who make model...  ...will:Design high-throughput data...  ...across storage, network, CPU...  ...formatsPartner with ML teams to scale...  ...Profiler), and interconnects (NVLink, InfiniBand... 
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    1 day ago
  • $159k - $230k

     ...failure reproduction.Design network, power and cooling infrastructure for end-to-end system...  ...designers, test engineers, on project planning within...  ...tests of end-to-end highly complex ML system design...  ...specialized ML ASICS and high-speed interconnects. You will also be... 

    Google

    Sunnyvale, CA
    4 hours ago
  • $163k - $237k

     ...characterization, High Voltage Screen (HVS...  ...in Electrical Engineering, Computer Engineering...  ...Generation (ATPG)/High-Speed Input/Output (HSIO...  ...Streaming Scan Network (SSN)/Streaming...  ...the future of AI/ML hardware acceleration...  ....The AI and Infrastructure team is redefining... 
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $160.36k - $240.54k

     ...approach to autonomous driving, and the ML Infrastructure team builds and operates the...  ...streaming ingestion, storage layout, and high-throughput data generation and storage....  ...or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 1+ years... 
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago
  • $193.3k - $261.5k

    We are seeking an experienced engineer to work on distributed AI/ML systems. This role involves...  ...valued, and experience with high-speed networking or HPC interconnects is valued highly.If you like...  ...critical building blocks for EC2 infrastructure. Every instance in EC2 is... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $150k - $230k

     ...researchers and veteran systems engineers who share a vision for...  ...complex, traditional infrastructure struggles to meet the...  ...for fault-tolerant, high-performance...  ...of GPU systems, high-speed networking, and distributed coordination...  ...(RDMA, InfiniBand) ML framework or runtime internals... 

    Clockwork.io

    Palo Alto, CA
    more than 2 months ago
  • $148.2k - $222.2k

    Staff Engineer, Infrastructure Platforms Position SummaryThe Staff Engineer, Infrastructure...  ..., enterprise storage, High Performance Computing (HPC),...  ...storage, virtualization, networking, and High Performance...  ...genomics, bioinformatics, AI/ML, or other scientific computing... 
    Full time
    Work from home
    Monday to Friday

    Pacific Biosciences

    Menlo Park, CA
    9 hours ago
  • $152k - $282k

     ...new services at the speed we have been since...  ...this always-on, high-tech, and hyper-connected...  ...our Application Infrastructure org, this is a...  ...Staff Software Engineer to lead the architectural...  ...management, and network I/O to ensure sub-...  ...: Integrate AI/ML models directly into... 
    Temporary work

    Coupang

    Mountain View, CA
    9 hours ago
  •  ...Lead Software Engineer We have an opportunity to impact your career...  ...problems Develops secure high-quality production code, and...  ...improvements in delivery speed, reliability, and code quality...  ...and skills Experience in Infrastructure Architecture designs Experience... 

    Chase

    Palo Alto, CA
    4 days ago
  • $126.1k - $185k

    Do you like to use network and Unix systems engineering to deliver simple, sustainable...  ...IP networks?AWS Infrastructure Services owns the design...  ...customers’ requirement on interconnect solutions. You'll work backwards...  ...teams to maintain a highly available network that delights... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $132k - $190k

     ...degree in Electrical Engineering, Computer Engineering,...  ...Peripheral Component Interconnect Expres (PCIe), Double...  ...Artificial Intelligence (ML/AI) hardware systems projects...  ...center.Our Platforms Infrastructure Engineering team...  ...all the way through to high-volume manufacturing.... 
    Worldwide

    Google

    Sunnyvale, CA
    9 hours ago
  •  ...Job Title: Network Engineer Location: Sunnyvale,...  ...training/inference server infrastructure silicon . As a Network...  ..., RDMA, and RoCE for high-performance...  ...exchanges across high-speed interfaces. Understanding...  ...and their high-speed interconnect fabric interfaces... 
    Full time
    Work experience placement

    Stellent IT LLC

    Sunnyvale, CA
    1 day ago
  • $143.2k - $243k

     ...California in 2004 when a visionary engineer, Fred Luddy, saw the...  ...to a composable, agent-native infrastructure foundation that agents and applications...  ...corpus. ~ Design and run high-throughput ingestion and...  ...with Search Ranking, ML, and Platform engineering teams... 
    Full time
    Work at office
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Mountain View, CA
    22 days ago
  •  ...position as a key member of a high-performing team that delivers infrastructure and performance...  ...As a Lead Infrastructure Engineer at JPMorganChase within...  ...engineering such as hardware, networking terminology, databases,...  ...exploring AI tools to speed up troubleshooting and documentation... 
    Permanent employment

    JP Morgan Chase

    Palo Alto, CA
    4 days ago
  •  ...Overview The AI Infrastructure Engineer at HOAi is responsible for scaling...  ...serving infrastructure, and ML pipelines that power HOAi's products...  ...Velocity-minded: Balances speed with quality; ships incrementally...  ...raiser: Sets and maintains high standards; elevates team... 
    Work at office
    Remote work
    Flexible hours

    Vantaca

    Redwood City, CA
    5 days ago
  • $188.5k - $255k

     ...tripling the speed at which developers...  ...infrastructure, operational intelligence...  ...developer, or engineering leader gets...  ...Consumption, AI/ML Enablement,...  ...Abstraction & Network Infrastructure...  ...across three interconnected areas: enterprise...  ...)Convert high-ambiguity charters... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    9 hours ago
  • $272,000 - $327,000 per day

     ...Technology and Applications - Network Services & Client Support /...  ...Senior Manager of Network Engineering to lead the team responsible...  ...position manages our network infrastructure across multiple global sites...  ...deadlines. You will guide a highly capable engineering team while... 
    Full time
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    9 hours ago
  • $180k - $250k

     ...blog posts sharing our high-level results for...  ...data research and data engineering necessary to solve this...  ...an experienced Cloud Infrastructure Engineer to join our...  ...training large-scale ML models Ensure our infrastructure...  ...-level debugging—networking issues, memory leaks,... 
    Full time
    Work at office
    Relocation package

    Datologyai

    Redwood City, CA
    9 hours ago
  • $207k - $340k

     ...largest professional network, built to create economic...  ...Staff Software Engineer to lead LinkedIn’s GPU...  ...Platform, a foundational AI infrastructure stack that powers...  ...scalable, observable, and highly available multi-tenant...  ...Suggested Skills:AI / ML Infrastructure Technical... 
    For contractors
    Work at office
    Remote work
    Work from home
    Flexible hours

    Linkedin

    Mountain View, CA
    2 days ago
  • $124k - $250k

     ...of others.A Day in the LifeAs a member of our software engineering infra team, you'll solve technical challenges,...  ...upgrading and implementing state-of-the-art software infrastructure. The team builds a high-performance, high availability, globally distributed ecosystem... 

    AppLovin

    Palo Alto, CA
    1 day ago
  • $152k - $228k

     ...autonomy code change, from ML model updates to...  ...road. You will own the infrastructure that makes this possible...  ...much much more. Engineers across the company rely...  ...onboard logging rate or network contention as more...  ...visualization layer. This is a high-ownership, high-... 
    Full time
    Temporary work

    Who We Are Nuro

    Mountain View, CA
    9 hours ago
  • $193.3k - $261.5k

     ...seeking an experienced engineer and technical...  ...team that owns the network stack for EC2 distributed AI/ML systems. The team develops...  ...experience with high-speed networking or HPC/RDMA interconnects is highly valued.If...  ...blocks for EC2 infrastructure. Every instance in... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $192k - $275k

     ...HybridZoox’s Robot Software Infrastructure Deployment team builds...  ...: advanced sensors, high-performance compute...  ..., automotive ECUs, network switches, and user interface...  ...when solving complex engineering challenges. If you...  ...improvements in deployment speed, efficiency, and... 
    Full time
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    3 days ago
  • $157k - $235k

     ...other digital services.Snap Engineering teams build fun and technically...  ...Engineer to join the ML Platform Experience team, part...  ...you’ll do:Design and optimize infrastructure systems for machine learning...  ...training data generationDevelop high-performance inference systems... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Network Engineer - ML Infrastructure (High-Speed Interconnects). Be the first to apply!