Senior Network & Site Reliability Engineer
Alembic
About UsAlembic is the pioneering Causal AI platform. We help the world's largest enterprises move past correlation to prove what actually drives business outcomes — the question marketing and growth teams have never been able to answer with confidence. Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence.We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture. Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure — one of the fastest private supercomputers in the world. (We've melted GPUs getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'll design and operate the global network and reliability layer behind one of the world's fastest private supercomputers — the fabric powering distributed compute, ML workloads, real-time analytics, and mission-critical enterprise systems. You'll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions.It's a strong fit if you like solving deep infrastructure problems, building resilient systems, automating everything repetitive, and owning architecture rather than just maintaining it.What You'll DoArchitect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads.Own network device configuration management end to end, ensuring consistency and reliability across the fleet.Improve system and network reliability and performance through automation, observability, and proactive capacity planning.Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.Build and maintain comprehensive monitoring, alerting, and incident response — SLOs, runbooks, and on-call rotations — and drive post-incident analysis and continuous improvement.Ensure security, compliance, and operational readiness across our network and cloud infrastructure.Partner across engineering and data science to drive a culture of performance and reliability.What Will Help You Succeed8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration.A strong background in network security, architecture, design, and operations.Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols — BGP, QoS, MPLS, and IPsec VPNs.Experience designing and operating modern datacenter network fabrics (spine-leaf, EVPN/VXLAN, ECMP).Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).WAN engineering — carrier circuit provisioning and external network peering.Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure.Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross-functional communication.Also HelpfulNVIDIA networking technologies — Cumulus Linux, InfiniBand, Spectrum-X, and BlueField DPUs (this is the fabric behind our SuperPOD).Familiarity with data-intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI).Security practices for applications and infrastructure, and experience in high-compliance or SOC 2 environments.The Role Is Right for You IfYou want to own mission-critical network and infrastructure end to end — from architecture to incident management — not just keep it running.You'd rather build and automate than direct from a distance, and you want meaningful influence over how a high-performance platform scales.Why You Might Be Excited About AlembicHard problems with real impact: You'll own the network and reliability layer behind systems that influence multimillion-dollar decisions at Fortune 100 companies.Cutting-edge technology: Operate our own NVIDIA DGX SuperPOD on Grace Blackwell — one of the fastest private supercomputers in the world — and run a fabric (InfiniBand, Spectrum-X, BlueField) almost no company has in-house.Technical autonomy: Ownership over architecture decisions and the freedom to solve hard infrastructure problems your way.Elite team: Join top engineers who thrive on hard problems and high-impact work.Series B momentum, real ownership: Meaningful equity at a Series B company that's raised $145M, with proven product-market fit and Fortune 100 traction.Why You Might Not Be ExcitedIf you only want to tell people what to build instead of building and automating alongside them, this isn't the environment for you.You prefer companies with 100% built-out process for every detail.You prefer static over dynamic — projects and priorities adapt as we grow. We have real paying customers and a playbook, and we still move at startup speed at Series B scale.Compensation Range: $210K - $240KLocationSan Francisco HQAddressSan Francisco, CaliforniaEmployment TypeFull timeLocation TypeOn-siteDepartmentEngineeringTechnical OperationsCompensationOur compensation philosophy ensures new hires earn 50%+ current benchmarks. We'll discuss further with you. Most recent San Francisco benchmark data: $210K – $240KOur Compensation Philosophy:Market-based: Our formula ensures new hires earn at or above real-time benchmarks. Ownership: Our generous equity program ensures new hires are owners, not just employees.Transparent: We openly discuss salary expectations to avoid surprises later in the process. Data-driven: We use objective data to remove bias and ensure consistency in compensation decisions.
- ...well. You know Linux. Even Kubernetes still runs on computers! This means you can debug most normal issues with performance, networking, kernel drivers, package management, etc. or have a good idea where to start You have used at least one of the major cloud providers...SeniorNetworkTemporary workWork experience placement
$210k - $240k
Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual... ...Terraform, Ansible) Strong knowledge of Linux systems, networking, and systems performance tuning Experience with monitoring...SeniorNetworkFull time$175k - $250k
...50,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,... ...unavailable. Modality: On-Site only. Must live within commuting... ..., performance, and reliability across environments.... ...distributed storage, and networking Ensure reliability, scalability...SeniorNetworkFull timeRemote workRelocationRelocation package$250k
...in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... ...modern GPU cloud providers Strong understanding of networking fundamentals (DNS, TCP/IP, routing, performance...SeniorNetworkFull timeRemote work$181k - $263k
...organizations, between brands, and across its premier global network of top-quality partners. Hundreds of global... ...providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability...SeniorNetworkFull timeWork from homeWorldwideFlexible hoursNight shift$165k - $227k
...this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products... ...environments. Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload...SeniorNetworkLocal areaWorldwideFlexible hours$232k - $319k
...scale the service with great people and reliable, cost-effective, and efficient... ...oversee multiple teams focused on Edge networking, K8s platform, Observability, automation... ...Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful...SeniorNetworkPermanent employmentLocal areaWorldwideFlexible hours$174.92k - $209.91k
...access to data as simple and reliable as electricity. With Fivetran... ...and ready to query, with no engineering or maintenance required. We’re... ...our teams, systems, and career sites. About the Role... ...~ Experience with cloud networking like VPNs, Privatelinks, and...SeniorNetworkFull timeWork at officeRemote work$300k
...full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation... ...-availability GPU workloads. Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU...SeniorNetworkPermanent employment$167.7k - $245.2k
...digital experiences across every network - even the ones they don’t... ....We’re looking for talented engineers with a software or operations... ...development teams to ensure the reliability, performance and security of... ...Please see the Cisco careers site to discover more benefits and...SeniorNetworkFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$190.8k - $267.1k
...Reddit grow its business. The reliability of our Ads systems directly... ...partners closely with Ads Engineering to improve reliability,... ...trust. We’re looking for a Senior Site Reliability Engineer to build... ...applications, infrastructure, networking, and services.Experience...SeniorNetworkFor contractorsWork experience placement- ...The Team Platform Engineering is the department within SRE that is responsible... ...responsibilities encompass network architecture, service mesh,... ...and maintaining the reliable and globally connected multi-... ...Overview We are seeking a talented Site Reliability Engineer (SRE)...SeniorNetworkFull timeWork at officeRemote workWorldwide
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range... ...cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-... ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...SeniorNetworkWork at officeLocal areaRemote workWorldwideFlexible hours$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure... ...team manages the global network substrate that enables... ...role in engineering the reliable, globally connected, multi-cloud... ...are seeking a talented Senior Site Reliability Engineer (SRE) with...SeniorNetworkLocal areaRemote workWorldwideFlexible hours$117k - $209.33k
...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...Kubernetes, cloud-native architectures, APIs, load balancing, networking, DNS, and distributed systemsExperience with...SeniorNetworkFull timeFor contractors$152.5k - $205k
...platform includes the world’s largest regulated stablecoin network anchored by USDC, Circle Payments Network for global money... ...everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries...SeniorNetworkFlexible hours$167.7k - $245.2k
...organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an... ...portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background...SeniorNetworkFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure.... ...solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security...SeniorNetworkLocal areaRemote workWorldwideFlexible hours- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services... ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role...Senior
- ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage...Senior
$200k - $240k
...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to... ...and Helm ~ Understand the importance of performant and reliable systems ~ Education - Ideally looking for a B.A. / B.S. degree...SeniorWork at officeImmediate start3 days per week$200.7k - $250.9k
...washed away in a flood in 1942, the Royal Engineers rebuilt it. Then it washed away again in... ...opportunities for improvements in reliability/observability/performance/preparedness and... ...candidate for the role: Has past Site Reliability Engineering or DevOps experience...Senior$200k - $240k
...Senior Site Reliability Engineer (SRE)Location: San Francisco, CAWork Model: OnsiteIndustry: Renewable EnergyComp: $200,000 - $240,000 We’re... ...ownership of critical production infrastructure spanning cloud, networking, and physical environments. This is not a traditional...Network$350k
...infrastructure providers to build reliable, high-performance platforms... ...opportunity is for a Staff Site Reliability Engineer to lead the reliability of large... ...health, distributed training, networking, and incident response. Working as a senior technical leader, the role focuses...Network$170k - $220k
...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes. Our innovative approach blends technology with deep legal expertise, making us a leader in our field. We go beyond surface-level...SeniorWork at officeRemote workFlexible hours- ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here,... ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...SeniorWorldwideWeekend work
- ...looking for a frontend focused engineer. The Role We’re... ...obsessed with building fast, reliable, and secure systems—and giving... ...from monitoring and alerting to networking, security, and CI/CD.... ...started/sold ed-tech company, Senior Class President at Stanford,...SeniorNetworkFull time
$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...SeniorTemporary workImmediate startFlexible hoursShift work- ...the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology,... ...fundamentals Linux/Unix or Windows Server administration Browser/network troubleshooting tools (e.g., Firebug, Chrome DevTools,...Network
$194k - $237k
...sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and... ...technical direction, develops senior technical leaders, and demonstrates impact... ...AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures...NetworkHourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Network & Site Reliability Engineer. Be the first to apply!
- lead network engineer San Francisco, CA
- junior network engineer no experience San Francisco, CA
- cisco network engineer San Francisco, CA
- production network engineer San Francisco, CA
- network engineer San Francisco, CA
- network engineer full time San Francisco, CA
- remote cisco network engineer San Francisco, CA
- network consulting engineer San Francisco, CA
- network engineer internship San Francisco, CA
- network applications engineer San Francisco, CA



