Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Network & Site Reliability Engineer

Alembic

About UsAlembic is the pioneering Causal AI platform. We help the world's largest enterprises move past correlation to prove what actually drives business outcomes — the question marketing and growth teams have never been able to answer with confidence. Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence.We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture. Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure — one of the fastest private supercomputers in the world. (We've melted GPUs getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'll design and operate the global network and reliability layer behind one of the world's fastest private supercomputers — the fabric powering distributed compute, ML workloads, real-time analytics, and mission-critical enterprise systems. You'll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions.It's a strong fit if you like solving deep infrastructure problems, building resilient systems, automating everything repetitive, and owning architecture rather than just maintaining it.What You'll DoArchitect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads.Own network device configuration management end to end, ensuring consistency and reliability across the fleet.Improve system and network reliability and performance through automation, observability, and proactive capacity planning.Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.Build and maintain comprehensive monitoring, alerting, and incident response — SLOs, runbooks, and on-call rotations — and drive post-incident analysis and continuous improvement.Ensure security, compliance, and operational readiness across our network and cloud infrastructure.Partner across engineering and data science to drive a culture of performance and reliability.What Will Help You Succeed8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration.A strong background in network security, architecture, design, and operations.Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols — BGP, QoS, MPLS, and IPsec VPNs.Experience designing and operating modern datacenter network fabrics (spine-leaf, EVPN/VXLAN, ECMP).Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).WAN engineering — carrier circuit provisioning and external network peering.Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure.Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross-functional communication.Also HelpfulNVIDIA networking technologies — Cumulus Linux, InfiniBand, Spectrum-X, and BlueField DPUs (this is the fabric behind our SuperPOD).Familiarity with data-intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI).Security practices for applications and infrastructure, and experience in high-compliance or SOC 2 environments.The Role Is Right for You IfYou want to own mission-critical network and infrastructure end to end — from architecture to incident management — not just keep it running.You'd rather build and automate than direct from a distance, and you want meaningful influence over how a high-performance platform scales.Why You Might Be Excited About AlembicHard problems with real impact: You'll own the network and reliability layer behind systems that influence multimillion-dollar decisions at Fortune 100 companies.Cutting-edge technology: Operate our own NVIDIA DGX SuperPOD on Grace Blackwell — one of the fastest private supercomputers in the world — and run a fabric (InfiniBand, Spectrum-X, BlueField) almost no company has in-house.Technical autonomy: Ownership over architecture decisions and the freedom to solve hard infrastructure problems your way.Elite team: Join top engineers who thrive on hard problems and high-impact work.Series B momentum, real ownership: Meaningful equity at a Series B company that's raised $145M, with proven product-market fit and Fortune 100 traction.Why You Might Not Be ExcitedIf you only want to tell people what to build instead of building and automating alongside them, this isn't the environment for you.You prefer companies with 100% built-out process for every detail.You prefer static over dynamic — projects and priorities adapt as we grow. We have real paying customers and a playbook, and we still move at startup speed at Series B scale.Compensation Range: $210K - $240KLocationSan Francisco HQAddressSan Francisco, CaliforniaEmployment TypeFull timeLocation TypeOn-siteDepartmentEngineeringTechnical OperationsCompensationOur compensation philosophy ensures new hires earn 50%+ current benchmarks. We'll discuss further with you. Most recent San Francisco benchmark data: $210K – $240KOur Compensation Philosophy:Market-based: Our formula ensures new hires earn at or above real-time benchmarks. Ownership: Our generous equity program ensures new hires are owners, not just employees.Transparent: We openly discuss salary expectations to avoid surprises later in the process. Data-driven: We use objective data to remove bias and ensure consistency in compensation decisions.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Network & Site Reliability Engineer in San Francisco, CA vacancy
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and...  ...of experience with datacenter operations and/or system and network administrationExperience with containerization (Docker),... 
    Senior
    Network

    Alembic

    San Francisco, CA
    4 days ago
  • $165k - $241.4k

     ...digital experiences across every network - even the ones they don’t...  ....We’re looking for talented engineers with a software or operations...  ...development teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Senior
    Network
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $152.5k - $205k

     ...platform includes the world’s largest regulated stablecoin network anchored by USDC, Circle Payments Network for global...  ...everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and... 
    Senior
    Network
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range...  ...cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-...  ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... 
    Senior
    Network
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...Kubernetes, cloud-native architectures, APIs, load balancing, networking, DNS, and distributed systemsExperience with... 
    Senior
    Network
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    6 hours ago
  • $165k - $225.6k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability...  ...technical documentation, including network diagrams, runbooks, and disaster recovery procedures... 
    Senior
    Network
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $165k - $241.4k

     ...organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an...  ...portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background... 
    Senior
    Network
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    6 hours ago
  • $227.2k - $324.5k

    About the Role:Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are...  ...seeking an experienced and visionary Senior SRE Manager to lead and grow our newly...  ...knowledge of AWS services (especially networking, IAM, EKS, ALBs/NLBs, Route 53,... 
    Senior
    Network
    Full time
    Contract work
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    4 days ago
  • $215k - $275k

     ...raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide...  ...) and Kubernetes-based deploymentsDeep understanding of networking, security, and authentication mechanisms in cloud... 
    Senior
    Network
    Work at office

    Anyscale

    San Francisco, CA
    1 day ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure....  ...solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security... 
    Senior
    Network
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  •  ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools...  ...experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient than most... 
    Senior
    Network

    TechChain Talent

    San Francisco, CA
    4 days ago
  •  ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses...  ...This means you can debug most normal issues with performance, networking, kernel drivers, package management, etc. or have a good idea... 
    Senior
    Network
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    3 days ago
  • $175k - $250k

     ...50,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,...  ...unavailable. Modality: On-Site only. Must live within commuting...  ..., performance, and reliability across environments....  ...distributed storage, and networking Ensure reliability, scalability... 
    Senior
    Network
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    10 hours ago
  • $232k - $319k

     ...scale the service with great people and reliable, cost-effective, and efficient...  ...oversee multiple teams focused on Edge networking, K8s platform, Observability, automation...  ...Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful... 
    Senior
    Network
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    4 days ago
  • $15k

     ...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute...  ...for HPC/batch compute environmentsExperience with HPC networking (InfiniBand, RDMA)Solid security/IAM foundations (... 
    Senior
    Network
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    2 days ago
  • $139.76k - $287.75k

     ...their business.We are seeking a Senior Site ReliabilityEngineer to help...  ...in advancing the reliability, scalability, automation, observability...  ...is a highly hands-on engineer with strong production experience...  ...Linux, containers, IAM, networking, and distributed systemsExperience... 
    Senior
    Network
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  •  ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer...  ...(EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms. ~ Deep understanding... 
    Senior
    Network

    Kody

    San Francisco, CA
    16 days ago
  • $300k

     ...full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation...  ...-availability GPU workloads. Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU... 
    Senior
    Network
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $250k

     ...in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments...  ...modern GPU cloud providers Strong understanding of networking fundamentals (DNS, TCP/IP, routing, performance... 
    Senior
    Network
    Permanent employment
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ...ownership, working together to build scalable, reliable, and secure products that empower...  ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Senior
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    4 days ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations... 
    Senior
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • $220k - $235k

     ...are seeking a strategic, high-output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role,...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... 
    Senior
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    3 days ago
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services...  ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role... 
    Senior

    Socket

    San Francisco, CA
    10 hours ago
  • $195k - $257.5k

     ...world’s largest regulated stablecoin network anchored by USDC, Circle Payments Network...  ...you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll...  ...:Staff Site Reliability Engineer (IV) Senior Site Reliability Engineer (III)What you... 
    Network
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  •  ...and services they want to use. Plaid’s network covers 12,000 financial institutions...  ...builds the platforms and tooling that help engineering teams develop, deploy, and operate...  ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'... 
    Network
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    3 days ago
  • $155k - $222.6k

     ...cloud platform. As a team of six engineers distributed across the US,...  ...strong focus on automation, reliability, and operational excellence....  ...~2+ years of experience in Site Reliability Engineering, DevOps...  ...solutions. Add to that our worldwide network of doers and experts, and you... 
    Network
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    Daly City, CA
    5 days ago
  • $194k - $267k

     ...If you are too, let's talk.The TeamThe Site Reliability team is dedicated to architecting and...  ...that maximize platform reliability and engineering velocity.The ideal candidate is...  ...are part systems administrator, part network administrator, and part developer. What... 
    Network
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $204k - $306k

     .... If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure...  ...Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes...  ...oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation... 
    Network
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...s talk.We are seeking a highly technical ObservabilitySite Reliability Engineer with a specialty in Google Cloud, to own and expand our Observability...  ...Systems: Deep understanding of Linux internals, networking (TCP/IP, DNS, Load Balancing), and container orchestration... 
    Network
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Network & Site Reliability Engineer. Be the first to apply!