Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Network Engineer, Platform, Automation & HPC/AI

Lawrence Berkeley National Laboratory

Join NERSC at Berkeley Lab and help engineer the high-performance network platform powering some of the nation's most advanced scientific computing. As a Network Engineer, Platform, Automation & HPC/AI, you'll work at the intersection of network engineering, automation, and software development while helping advance AI-driven network operations. Our team manages 1 Tb/s of border connectivity to ESnet and an 800G/400G data center network backbone supporting the NERSC-9 and NERSC-10 supercomputers, multi-tier storage, archive, and edge services. Your work will help improve the performance, scalability, automation, and reliability of scientific workflows serving more than 10,000 users. Our mission is to bring science solutions to the world. We welcome candidates from all backgrounds, including those with non-traditional paths. We value a growth mindset and believe skills and experience are transferable. If you're eager to learn and meet the qualifications below, we encourage you to apply. Join our team where your work can have a high impact for an organization associated with 17 Nobel Prizes... and counting. This position may be filled at Level 3 or Level 4. Level 3 is intended for experienced engineers who independently solve complex networking and automation challenges. Level 4 is intended for senior technical leaders who architect solutions, lead modernization efforts, and establish new technical approaches for complex HPC and data center environments. Network Engineer Level 3 will: Implement, operate, maintain, and improve network automation and observability solutions. Contribute to Data Center modernization efforts and NERSC's Smart Facility initiative. Support efforts to design and deliver network services to address emerging needs (e.g., American Science Cloud, new Edge services). Continuously monitor and optimize network performance, focusing on reducing latency, maximizing throughput, and improving fault tolerance. Create and maintain comprehensive network documentation, including physical and logical topology diagrams. Collaborate with the Security Group to implement security measures for data integrity and privacy, ensuring high availability and reliability through redundancy and failover mechanisms. Share on-call rotation with colleagues and serve as an escalation contact for service incidents. Work on and resolve complex issues where analysis of situations or data requires an in-depth evaluation of multiple variables. Exercise judgment in selecting methods, techniques and evaluation criteria for obtaining results. Build effective working relationships with technical partners and stakeholders across disciplines. In addition to the above, the Senior Network Engineer Level 4 will: Architect, develop, and establish technical direction for network automation, observability, and self-healing capabilities. Lead Data Center modernization efforts in support of NERSC's Smart Facility initiative and emerging needs e.g., American Science Cloud Design, develop, and maintain automation frameworks, infrastructure-as-code, and software solutions to manage, optimize, and self-heal the HPC and data center network. Build and integrate AI/ML-driven observability, predictive analytics #J-18808-Ljbffr Lawrence Berkeley National Laboratory

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Network Engineer, Platform, Automation & HPC/AI in Berkeley, CA vacancy
  • Berkeley Lab’s NERSC division seeks a Network Engineer, Platform, Automation & HPC/AI to help engineer the high‑performance network platform powering national scientific computing. You will work at the intersection of network engineering, automation, and software development... 
    Network

    LBL

    Berkeley, CA
    1 day ago
  • $156.86k - $191.72k

    Join NERSC at Berkeley Lab and help engineer the high-performance network platform powering some of the nation’s most advanced scientific computing. As a Network Engineer, Platform, Automation & HPC/AI, you’ll work at the intersection of network engineering, automation,... 
    Network
    Full time
    Remote work
    Flexible hours

    Berkeley Lab

    Berkeley, CA
    2 days ago
  • Lawrence Berkeley National Laboratory's NERSC is seeking a Network Engineer, Platform, Automation & HPC/AI to advance the 1 Tb/s border network and an 800G/400G data center backbone supporting HPC workloads and a wide user base across scientific computing. This role spans... 
    Network

    Lawrence Berkeley National Laboratory

    Berkeley, CA
    1 day ago
  • Berkeley Lab is seeking a Network Engineer to join the team that powers AI-driven network operations and HPC/AI workflows. You’ll help design, implement, and operate automation and observability for a 1 Tb/s border network and an 800G/400G data center backbone supporting... 
    Network

    Berkeley Lab

    Berkeley, CA
    2 days ago
  •  ...Cloud, is a leader in AI cloud infrastructure...  ...a Senior Software Engineer on Lambda’s Core Cloud Platform team, you will build...  ...infrastructure dependencies, networking, and cloud workflows...  ..., infrastructure automation, and service...  ...GPU infrastructure, HPC, Kubernetes, Slurm,... 
    Network
    Full time
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    16 hours ago
  •  ...intelligence to power the AI economy. We partner with leading...  .... Our vast talent network trains frontier AI models in...  ...the Role As a Software Engineer on the Automations team at Mercor, you’ll join...  ...directly with operators and platform engineers to turn repeatable... 
    Network
    Full time
    Work at office
    Immediate start

    Mercor

    San Francisco, CA
    16 hours ago
  •  ...fans out into multiple AI inference requests running...  .... Today, our inference engineers and researchers build...  ...while also managing networking, securing capacity, and...  ...responsibilities we want a dedicated platform team to own. Your job...  ...-LLM. Slurm or other HPC schedulers. GPU kernel... 
    Network
    Shift work

    Neura Market

    San Francisco, CA
    2 days ago
  • $250.6k - $362.6k

     ...U.S. soil.Meet the Team The Platform Engineering organization is responsible...  ...power engineering across Cisco Network Platform. Our team owns the...  ...provide the infrastructure, automation, observability, security,...  ...protect organizations in the AI era - and beyond. We’ve been... 
    Network
    Permanent employment
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    5 days ago
  • $180k - $250k

     ...Nimble Nimble is an AI robotics company...  ...FedEx to build a national network of autonomous...  ...team of the world’s best engineers and operators. If you...  ...backend and robotics platform software that powers...  ...systems. Exposure to automation environments such as warehousing... 
    Network
    Full time
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    16 hours ago
  • $300 per month

     ...vertically integrated AI infrastructure company...  ...seeking Senior Software Engineers to design and develop...  ..., as well as creating automation software to...  ...and enhancing our cloud platform’s overall performance...  ...configuration of servers, network switches, power delivery... 
    Network
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    16 hours ago
  • $156.86k - $191.72k

     ...seeking a System Infrastructure / Platform Engineer to help build and manage HPC systems and Linux-based...  ...clusters, parallel storage, high-speed networking, Slurm, and Kubernetes, balancing...  ...Develop and maintain scripts and automation tools Participate in a 24/7 on-call... 
    Network
    Permanent employment
    Full time
    Remote work
    Flexible hours

    Berkeley Lab

    Berkeley, CA
    3 days ago
  • $207k - $275k

     ...The Essential Cloud for AI™. Built for pioneers...  ...CoreWeave delivers a platform of technology, tools,...  ...You'll Do The Endpoint Engineering team at CoreWeave manages...  ...of engineers, network operators, and support...  ...detection and cleanup automation, and deliver self-service... 
    Network
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    CoreWeave

    San Francisco, CA
    29 days ago
  • $285k - $335k

     ...vertically integrated AI infrastructure...  ...Infrastructure Operating Platform: a unified control...  ...is a Principal Engineer role reporting...  ...running, and drive automated remediation — a...  ...attestation, burn-in, network readiness,...  ...datacenter and GPU/HPC infrastructure — bare... 
    Network
    Temporary work

    Crusoe

    San Francisco, CA
    7 days ago
  •  ...to accelerate the progress of AI applications out into the...  ...Anyscale is looking for a Software Engineer to join the Infrastructure...  ...that powers Anyscale’s cloud platform. You will have the opportunity...  ...~ Deep understanding of networking, security, and authentication... 
    Network
    Full time

    Anyscale

    San Francisco, CA
    16 hours ago
  •  ...Are Koah Labs is building the ad network to power the next generation of AI-native products. Our mission is to...  ...Role We are looking for strong engineers with experience and interest in designing...  ...that make up our adtech platform. You might be a fit if You... 
    Network
    Full time

    Koah

    San Francisco, CA
    16 hours ago
  •  ...inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE...  ..., offering unparalleled learning and networking opportunities. Apply now to... 
    Network
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  • Crusoe Cloud is seeking a seasoned software engineer to design and build internal datacenter tooling and infra management systems. You will create automation for rapid deployment and configuration of servers, network gear, and power delivery units. Candidates should have... 
    Network

    FELICIS

    San Francisco, CA
    16 hours ago
  • $160k - $230k

     ...About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle...  ...the fastest LLM inference engine with state-of-the-art AI cloud...  ...managing our data center compute, networking, and storage. Design and build... 
    Network
    Full time
    Remote work

    Together Ai

    San Francisco, CA
    16 hours ago
  • $200k - $400k

     ...the leading conversational AI platform empowering every brand to deliver...  ...that power Decagon: networking, data, ML serving, developer...  ...looking for a Senior Software Engineer to help build and evolve our...  ...everything from CI/CD and release automation to observability standards... 
    Network
    Full time
    Work at office
    Local area

    Decagon

    San Francisco, CA
    7 hours ago
  • $220k - $260k

     ...interaction — a unified platform that combines streaming...  ...layer enterprise AI agents need in production...  ...experienced Staff Software Engineer to help architect and...  ...tools and services for automated infrastructure provisioning...  ...IaC for system and network infrastructure Familiarity... 
    Network
    Full time
    Work at office

    Redpanda Data

    San Francisco, CA
    16 hours ago
  • $127k - $184k

     ...expertise to prove the value of Google Cloud Platform across the portfolio through complex...  ...or support role.Experience with cloud engineering, on-premise engineering, virtualization...  ...Data Management, Data Analytics, Cloud AI, Networking, Migrations, Security.Experience... 
    Network

    Google

    San Francisco, CA
    1 day ago
  •  ...Description We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and...  ...on is probably the largest anime AI training cluster in the world . You’...  ..., parallel filesystems are serving, network is transmitting, and that the anime models... 
    Network
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    22 days ago
  • $200k

     ...Ready to architect the high-speed networks powering the AI era? Join a trailblazing leader in GPU...  ...Gain exposure to cutting-edge network automation and help shape the future of AI infrastructure...  ...(Hyper-scalers, FinTech, or HPC) where performance is the primary metric... 
    Network
    Full time
    San Francisco, CA
    more than 2 months ago
  • $10k - $20k

     ...services. The company is looking for a Network Automation Engineer to design and build high-speed, low-...  ...800G connectivity, supporting advanced AI and LLM workloads in a high-performance...  ...environments (Hyper-scalers, FinTech, or HPC) where performance is the primary... 
    Network
    Full time
    San Francisco, CA
    7 days ago
  • $172.5k - $260.1k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ....About the TeamThe Platform Orchestration team is...  ...on creating a unified, automated, and scalable platform that enables engineering teams to deliver...  ...domains such as IAM, EC2, or networking.AWS certifications (Professional... 
    Network
    Permanent employment
    Full time

    Salesforce

    San Francisco, CA
    3 days ago
  • BackOps AI is transforming supply chain operations with agentic AI solutions. We are seeking a Senior Software Engineer to architect and implement end-to-end systems powering our agentic solutions. This is a full-time, hybrid role with a core Bay Area presence and remote... 
    Full time
    Remote work

    BackOps AI

    San Francisco, CA
    4 days ago
  • $148.7k - $201.2k

     ...-disciplinary team of scientists, engineers, and technicians, on a mission to...  ...computer.We are looking to hire an HPC Platform Engineer to develop, automate, and maintain high-performance computing...  ...and Lambda- Experience in network fundamentals (DNS, DHCP, TCP/IP, routing... 
    Network
    Local area
    Flexible hours

    Amazon

    San Francisco, CA
    4 days ago
  • $240k - $300k

     ...unprecedented speed and accuracy. Our AI-enabled platform turns siloed and disconnected data into...  ..., while making it easy for other engineers to ship safely and quickly on top of it...  ...are rare and fast to resolve; hardening network and platform security controls to meet... 
    Network
    Local area

    Peregrine Technologies

    San Francisco, CA
    5 days ago
  • $130k - $400k

     ...Full Stack Engineer, Assessment Platform Location: San Francisco, CA Company Stage of Funding: Series C AI Marketplace ($10B Valuation) Office Type: Onsite (5 Days...  ...AI labs and enterprises with a global network of experts who generate the high-quality... 
    Network
    Full time
    Work at office
    Worldwide
    Relocation package

    Recruiting From Scratch

    San Francisco, CA
    16 hours ago
  • $207k - $300k

     ...-performing team of software engineers, actively coaching them to grow...  ....Experience with Generative AI and building Generative AI...  ..., large-scale system design, networking, security, data compression,...  ...highly performant team owning a platform that manages billions of security... 
    Network

    Google

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Network Engineer, Platform, Automation & HPC/AI. Be the first to apply!