Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote Site Reliability Engineer in Network Infrastructure

Full-time

Nebius

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The Role

We’re looking for a Site Reliability Engineer to help build and run the fundamental part of Nebius – the Network – the infrastructure everything else depends on. This is an engineering-first SRE role: you’ll set clear reliability targets, build the tooling and automation to meet them, and make the network safer to operate as we scale quickly.

Your responsibilities will include:

  • Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense)
  • Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards
  • Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting)
  • Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents
  • Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes
  • Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast

We expect you to have:

  • Strong production Linux fundamentals and a structured approach to debugging complex systems
  • Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency/loss, failure domains, etc.)
  • Hands-on experience operating high-availability systems and improving them over time (not just “keeping lights on”)
  • Ability to write and maintain software/automation (Go is common for us; Python is also welcome)
  • Experience with modern infrastructure tooling (e.g., IaC, CI/CD, container platforms) and comfort automating operational workflows

It will be an added bonus if you have:

  • Experience with high-throughput traffic processing: load balancers, tunneling/decap, NAT64, or similar datapath-heavy systems
  • Low-level networking performance/debug background (eBPF/XDP, DPDK, perf/ftrace, kernel networking internals)
  • Experience building network-safe delivery pipelines (testing labs, staged rollouts, automated verification, drift detection)
  • Background with large-scale network observability/telemetry (e.g., routing/flow telemetry, regression detection at scale)

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What’s it like to work at Nebius:

Fast moving – Bold thinking – Constant growth – Meaningful impact – Trust and real ownership – Opportunity to shape the future of AI

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.

If you need accommodations during the application process, please let us know.

Jobicy JobID: 150130
Vacancy posted 11 days ago
Similar jobs that could be interesting for youBased on the Remote Site Reliability Engineer in Network Infrastructure in Remote vacancy
  •  ...team of researchers, engineers, designers, and...  ...performance, scalable and reliable machine learning...  ...are looking for a Site Reliability...  ...and influence the Infrastructure team’s roadmap based...  ...compute/storage/network resource and cost...  ...offices if you are remote, plus an annual company... 
    Remote work
    Full time
    Work experience placement
    Work at office
    Local area
    Home office

    Cohere

    New York, NY
    5 days ago
  • $139k - $155k

     ...Office expectations**For Remote Roles: If this role is...  ...focused group of SRE engineers dedicated to making...  ...s AI and ML platforms reliable, secure, and operationally...  ...patients. We own the infrastructure operability of AI data...  ...controls — including network segmentation, secrets... 
    Remote work
    Full time
    Work at office

    PointClickCare

    Salt Lake City, UT
    3 days ago
  • $159k - $272k

     ...Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help...  ...operateMaintains a broad internal professional network and knows when to engage/activate...  ...Maryland, Colorado, Washington and remote workers$175,000.00 - $299,000.00... 
    Remote work
    Full time
    Private practice
    Local area
    Work from home
    3 days per week

    T. Rowe Price

    Owings Mills, MD
    2 days ago
  • $127k - $249k

     ...experienced Senior or Staff Engineer for our SRE, InfraSec team,...  ...security of our cloud-based infrastructure. As a Staff SRE, you will be...  ...hybrid basis, or it can be fully remote while working from a...  ...AWS, Azure, GCP), including network and compute security, identity... 
    Remote work
    Local area
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    5 days ago
  • $195k - $285k

     ...build and lead Our Client's Site Reliability Engineering function from the ground up — owning the infrastructure that development,...  ...Deep Linux systems expertise: networking (TCP/IP, RDMA, bonding), kernel...  ...production scale — module design, remote state, environment... 
    Remote work

    Phizenix

    Santa Clara, CA
    a month ago
  • $215k - $275k

     ...date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation...  ...Kubernetes-based deploymentsDeep understanding of networking, security, and authentication mechanisms in cloud environmentFamiliarity... 
    Work at office

    Anyscale

    San Francisco, CA
    2 days ago
  • $135k - $200k

     ...The Role As a Senior Software Engineer on Network Infrastructure you will be joining a team whose mission...  ...(architecture, design patterns, reliability and scaling) of new and existing systems...  ...are a few roles that allow for “Remote” work on an exceptional basis. If you... 
    Remote work
    Full time
    Work experience placement
    Work at office
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    1 day ago
  •  ...Job Title: Software Engineer - Senior Level (IE Platform Infrastructure) Location: Arlington, VA Clearance: Top...  .... This role focuses on designing reliable integration environments, managing...  ...Development Training Reimbursement ~ Flexible/remote work schedule.... 
    Remote work
    Full time
    Flexible hours

    Vivsoft Technologies

    Arlington, VA
    1 day ago
  • $164k - $270k

     ...understanding when and how to act.Partner with Security, IT, and Infrastructure to translate compliance requirements (CMMC) into enforceable,...  ...building self-healing or auto-remediation platforms (remote actions, osquery + response, custom agents). Exposure to OT (Operational... 
    Remote work
    Permanent employment
    Full time
    Local area
    Flexible hours

    Hadrian

    Los Angeles, CA
    5 days ago
  • $135k - $200k

     ...more. The Role Software Engineers at Palantir build software...  ...aspect of a product. Our infrastructure teams are responsible for the...  ...you’re motivated to develop reliable, performant, and scalable systems...  ...a few roles that allow for “Remote” work on an exceptional... 
    Remote work
    Full time
    Temporary work
    Work experience placement
    Work at office
    Work from home
    Relocation package

    Palantir Technologies

    New York, NY
    1 day ago
  •  ...Your Role: Software Engineer - Platform / Core Infrastructure We’re looking for a Software...  ...systems that scale reliably and securely, and can be...  ...across compute, storage, networking, and observability to drive...  ...should know # Location: Remote in the US or EMEA (all team... 
    Remote job
    Full time
    Contract work
    Immediate start

    Xbow

    Remote
    1 day ago
  •  ...and geographies, a scalable, reliable infrastructure foundation is critical to...  ...seeking a Cloud Infrastructure Engineer for our Cellular...  ...infrastructure setup (compute, networking, IAM, etc.) and complete application...  ...to thrive in a fast-paced, remote-first environment—balancing... 
    Remote work
    Full time

    Abnormal Company

    United States
    1 day ago
  • As a Sr Cloud Infrastructure Engineer on the Infrastructure Platform Team, you will be working on the...  .... This role can be filled by a fully remote employee.Responsibilities include: Designing...  ...servers, cloud-native applications, networks, firewalls, load balancers, etc.... 
    Remote work
    Work from home

    Genesys

    Austin, TX
    5 days ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and...  ...looking for a hands-on senior engineer to own the lifecycle and automation...  ...: US, CA, Santa Clara; US, IL, Remote; US, WA, Remote; US, Remote; US... 
    Remote work
    Full time
    Weekend work

    Nvidia

    Illinois
    3 days ago
  •  ...DescriptionIn the assigned Job Role of Infrastructure Consultant 2, your Area Of Responsibility...  ...- authoring reusable modules, managing remote state and workspaces, and structuring...  ...integrations.• Experience in defining zero-trust networking principles, secrets management, and... 
    Remote work
    Full time
    Temporary work
    Relocation

    Infosys Technologies

    Charlotte, NC
    3 days ago
  • $280k - $380k

     ...scale. We focus on reliability and automation, engineering systems that perform...  ...turning complex infrastructure into reliable, well...  ...experienced DevOps/SRE (Site Reliability...  ...toolsSolid understanding of networking, security, and...  ...generally flexible for remote work, except for... 
    Remote work
    Work at office
    Local area
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    5 days ago
  • $142.5k - $198.75k

     ...industry and social infrastructure around the globe...  ...assist in the design, engineering, development,...  ...defined SLO's. Manage site stability, performance, reliability, and maintain...  ...development, systems, networking, and cloud...  ...WHILE THIS ROLE IS REMOTE, YOU MUST BE A US... 
    Remote work
    Full time
    Flexible hours

    Hitachi

    Greenville, SC
    5 days ago
  •  ...thinking organization, apply now.We are currently seeking a Network Engineer (Epic infrastructure) - Remote to join our team in Plano, Texas (US-TX), United...  ....Support disaster recovery (DR) testing and redundant site configuration (active/passive or active/active designs... 
    Remote work
    Work at office
    Flexible hours

    NTT DATA

    Plano, TX
    3 days ago
  • $112.9k - $257k

     ...Systems Platform and Infrastructure EngineerThe Opportunity...  ...cloud architects and engineers specializing in GitLab...  ...translate mission and reliability requirements into resilient...  ...page on our Careers site and reviewing Our...  ...cameras on during meetings.Remote: If this position is... 
    Remote work
    Full time
    Contract work
    Part time
    Work at office
    Local area

    Booz Allen Hamilton

    Hampton, VA
    1 day ago
  • $80 - $90 per hour

     ...LaSalle Network is hiring for a Senior Site Reliability Engineer (Compute Platform) with a leading infrastructure and platform engineering firm known for innovation and cutting-edge technological...  ...solutions. This opportunity is a remote role focused on deep infrastructure and... 
    Remote work
    Hourly pay
    Contract work
    Temporary work

    LaSalle Network

    Chicago, IL
    8 days ago
  • $180k - $220k

     ...possible.The Staff Platform Infrastructure Engineer will be instrumental in...  ...Washington, DC and will be fully remote.In this role, you will have...  ...tooling.Solid security and networking fundamentals for...  ...observability, incident response, and reliability practices.Proven ability to... 
    Remote work
    Full time
    Work at office
    Work from home
    Flexible hours

    Danaher Corporation

    New York, NY
    4 days ago
  • $155k - $170k

     ...getting started!OverviewThe Lead Platform Engineer, Cloud Infrastructure plays a critical role in advancing...  ..., and a proactive mindset to deliver reliable and efficient platform solutions.This...  ...of one of these locations. Fully remote work is not available for this role.ResponsibilitiesCollaborate... 
    Remote work
    Work at office
    Local area
    Work from home

    Planet Fitness

    Hampton, NH
    3 days ago
  • $120 - $130 per hour

     ...Narrative Art is seeking a highly experienced Network Engineer to perform a comprehensive assessment,...  ...of its enterprise network infrastructure.This role is intended for a level Cisco...  ...Sick Leave)Workplace TypeThis is a fully remote position.Application DeadlineThis position... 
    Remote work
    Contract work
    Temporary work
    Interim role
    Immediate start

    TEKsystems

    Los Angeles, CA
    1 day ago
  • $80 - $90 per hour

     ...LaSalle Network is hiring for a Senior Site Reliability Engineer (Storage Platforms) with a storage-focused, enterprise...  ...innovation and impactful cloud infrastructure. Join a dedicated team managing...  ...is contract-based, offering remote work with potential onsite collaboration... 
    Remote work
    Hourly pay
    Contract work
    Temporary work

    LaSalle Network

    Chicago, IL
    8 days ago
  • $75.8k - $144.2k

     ....This role is primarily On-Site, with flexibility at hiring...  ...SIGINT domain, is seeking an Infrastructure Engineer II who strives for...  ...AdministrationExperience with enterprise grade networking equipment, such as Cisco,...  ...as on-site, hybrid or remote.The salary range for this... 
    Remote work
    Temporary work
    Work experience placement
    Work at office
    Relocation
    Flexible hours

    Raytheon

    Annapolis Junction, MD
    5 days ago
  • $262k - $364k

     ...and coach a distributed engineering team, fostering innovation...  ...stack for efficiency and reliability.Performance and Network Engineering: Architect performant...  ...Performance Computing, Remote Direct Memory Access,...  ...or Machine Learning Infrastructure.Google's software engineers... 
    Remote work
    Worldwide

    Google

    Sunnyvale, TX
    1 day ago
  •  ...Role Summary: As an Infrastructure Engineer, you will have the opportunity...  ...improvement of our infrastructure reliability and security while...  ...Demonstrable experience with networks, security, load balancers,...  ...least three days a week. For remote employees, occasional travel... 
    Remote work
    Full time
    Work experience placement
    Work at office
    3 days per week

    Notable

    Remote
    1 day ago
  • $210k - $220k

     ...security, compliance, and engineering teams to automate...  ...compliance-fluent, and infrastructure-minded Senior Site Reliability Engineer - Government Cloud...  ...a permanent, full-time remote configuration based in the...  ...who handles distributed network topologies fluidly natively... 
    Remote work
    Permanent employment
    Full time
    Shift work

    Tines

    Remote
    a month ago
  •  ...wide range of customer infrastructure, including health...  ...operations. As an experienced engineer you will have...  ...complex technical, multi-site and multi-discipline incidents...  ...including virtual / remote working.Experience of...  ...learn, and build your network. Make this the place... 
    Remote work
    Flexible hours

    NTT DATA

    Birmingham, AL
    3 days ago
  •  ...transformation, bringing in engineering practices to an IT organization...  ...a background in Cloud Based Infrastructure on AWSto join our team. You...  ...APIs Good understanding of networking and security concepts The...  ...internally (not on this external site). If you have any questions... 
    Remote work
    Full time
    For contractors

    Autodesk

    Portland, OR
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote Site Reliability Engineer in Network Infrastructure. Be the first to apply!