Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Observability Lead - Cloud SRE & Network Reliability

$114k - $253k

Lam Research Corporation

The group you’ll be a part ofThe Global Information Systems Group is dedicated to the success of Lam through providing best-in-class and innovative information system solutions and services. Together, we support users globally with data, information, and systems to achieve their business objectives. The impact you’ll makeOur team at Lam is seeking a hands-on Observability Lead with a strong Site Reliability Engineering (SRE) and multi-cloud networking foundation to join our GIS Infrastructure Platform Engineering team. You will lead engineers in delivering robust observability frameworks, SLA/SLO/SLI disciplines, DR/BCP programs, backup and restore operations, and end-to-end network reliability across Azure, AWS, and GCP. You will own the full-stack delivery of observability, reliability, and resilience capabilities across a global multi-cloud enterprise.What you’ll doLead and grow a team delivering a world-class observability platform across global, multi-cloud production environments, including Azure, AWS, and GCP.Define and enforce SLA, SLO, and SLI frameworks across all infrastructure and network domains, driving continuous improvement through effective error budget management.Own end-to-end multi-cloud network observability, including VNet and VPC traffic flows, Transit Gateway routing, BGP peering health, and inter-region connectivity.Design and govern multi-cloud networking architectures, including Azure VNet, AWS VPC and Transit Gateway, GCP VPC, and hybrid connectivity solutions such as ExpressRoute, Direct Connect, and Cloud Interconnect.Design and implement agentic AI workflows using LLM-based agents, RAG patterns, and orchestration frameworks to enable AIOps-driven fault detection and remediation.Own disaster recovery (DR) and business continuity planning (BCP) strategy, including runbook authorship, multi-cloud failover validation, and periodic DR drills to ensure RTO and RPO commitments are met.Lead backup and restore operations across multi-cloud and hybrid environments, incorporating automated validation and cross-cloud recovery workflows.Build robust monitoring and alerting pipelines by integrating Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, and Google Cloud Operations into a unified observability stack.Drive automation-first practices through self-healing pipelines, remediation playbooks, and infrastructure-as-code (IaC) patterns to reduce toil and improve MTTR.Lead P1, P2, and P3 incident response efforts, including structured post-mortems and action tracking.Define and drive the multi-quarter roadmap for observability, reliability, networking, DR/BCP, and AI-assisted operations.Support hiring, performance management, and career development for the team.Who we’re looking forA BS, MS, or PhD in Computer Science, Engineering, or a related field (or equivalent experience), with 12+ years of overall experience in Infrastructure, SRE, DevOps, or Network Engineering and 6+ years of experience leading high-performing SRE, Observability, or Platform Engineering teams.Proven expertise in defining, enforcing, and operating SLA, SLO, and SLI frameworks, including effective error budget management.Hands-on experience with disaster recovery (DR) and business continuity planning (BCP), including RTO/RPO planning, failover testing, and continuity documentation.Deep expertise in backup and restore operations across multi-cloud and hybrid environments.Strong multi-cloud networking skills across Azure (VNet, ExpressRoute, Virtual WAN), AWS (VPC, Transit Gateway, Direct Connect), and GCP (VPC, Cloud Interconnect, VPC-SC).Experience building and operating observability platforms, including tools such as Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Splunk, or equivalent solutions, with a focus on network telemetry and flow analysis.Deep expertise in automation, including Ansible, Terraform, Python, and self-healing infrastructure pipelines.Hands-on experience with infrastructure as code (IaC), CI/CD pipelines, Kubernetes (AKS, EKS, GKE), and all three major cloud platforms.Strong programming skills in Python or Go for tooling, automation, and system integrations.Experience leading P1, P2, and P3 incident management, including ITSM integration (ServiceNow preferred).Exceptional communication skills, with the ability to translate complex technical concepts into clear business value for engineering, product, and executive stakeholders.Preferred qualificationsExperience with AIOps, including AI-assisted network fault detection, anomaly correlation, and auto-remediation.Familiarity with agentic AI workflows, including LLM-based agents and RAG patterns, applied to observability and operational use cases.Background in global WAN architectures, including MPLS and resilience strategies for multi-region enterprise environments.Experience with compliance-driven disaster recovery and business continuity (DR/BCP) programs, including InfoSec audits, SOX, and ISO 22301 requirements.Experience with FinOps and multi-cloud cost observability, including network egress visibility and cost optimization across Azure, AWS, and GCP.Relevant cloud certifications, such as Azure AZ-700 or AZ-305, AWS ANS-C01 or SAP-C02, and GCP Professional Cloud Network Engineer or Architect.Background in HPC, on-premises, or hybrid cloud environments.Our commitmentWe believe it is important for every person to feel valued, included, and empowered to achieve their full potential. By bringing unique individuals and viewpoints together, we achieve extraordinary results.Lam Research ("Lam" or the "Company") is an equal opportunity employer. Lam is committed to and reaffirms support of equal opportunity in employment and non-discrimination in employment policies, practices and procedures on the basis of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex (including pregnancy, childbirth and related medical conditions), gender, gender identity, gender expression, age, sexual orientation, or military and veteran status or any other category protected by applicable federal, state, or local laws. It is the Company's intention to comply with all applicable laws and regulations. Company policy prohibits unlawful discrimination against applicants or employees.Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on-site collaboration with colleagues and the flexibility to work remotely and fall into two categories – On-site Flex and Virtual Flex. ‘On-site Flex’ you’ll work 3+ days per week on-site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. ‘Virtual Flex’ you’ll work 1-2 days per week on-site at a Lam or customer/supplier location, and remotely the rest of the time.#LI-DM1SalaryCA San Francisco Bay Area Salary Range for this position: $114,000.00 - $253,000.00.The above salary range for this position is relevant to applicants that reside or work onsite in the California, San Francisco Bay Area only. Salary offers will depend on factors that include the location you work from, your level, education, training, specific skills, years of experience and comparison to other employees already in this role. Actual salary may vary from salary offered due to numerous factors including but not limited to unpaid time off, unpaid leave, company mandated shutdown, and other relevant factors.Our Perks and BenefitsAt Lam, our people make amazing things possible. That’s why we invest in you throughout the phases of your life with a comprehensive set of outstanding benefits.Department:Information Systems

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Observability Lead - Cloud SRE & Network Reliability in Fremont, CA vacancy
  • $186.9k - $267.7k

     ...intended, improving reliability and reducing risks...  ...with enhanced observability and control.As a Staff...  ...Engineer (SRE), you will provide...  ...reliability strategy, lead major infrastructure...  ...operating large-scale cloud and on-prem...  ...infrastructure, databases, and networking.Drive automation... 
    Cloud
    Network
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Fremont, CA
    29 minutes ago
  • $114k - $253k

     ...push for the next big semiconductor breakthrough. We lead the way in one of the most critical and fast-moving industries...  ...for a better world together, anything is possible. Observability Lead - Cloud SRE & Network Reliability Date: Jul 21, 2026 Location: Fremont, CA, US, 94538... 
    Cloud
    Network
    Local area
    Remote work
    Flexible hours
    2 days per week
    3 days per week
    1 day per week

    Lam Research Salzburg GmbH

    Fremont, CA
    1 day ago
  • $185k

     ...innovation. Backed by leading technology investors...  ...: A Senior Site Reliability Engineer (SRE) is expected to own...  ...performance ofJuul’s hybrid cloud infrastructure (...  ...Implement network micro-segmentation using...  ...Run with service mesh,observability, and security best practices... 
    Cloud
    Network
    Remote work

    GrabJobs

    Fremont, CA
    1 day ago
  •  ...itD is seeking a Site Reliability Engineer to develop and enhance automation...  ...efficiency of large-scale cloud infrastructure. The ideal...  .... Attend internal itD networking events (in person and virtual...  ...environments. Knowledge of monitoring, observability, and site reliability... 
    Cloud
    Network
    Work experience placement
    Remote work

    GrabJobs

    Fremont, CA
    1 day ago
  • $82.3k - $228.8k

     ...customer experience. Five9 is a leading provider of cloud contact center software,...  ...experienced Senior Site Reliability Engineer – Compute Platforms...  ...collaborate with platform and SRE teams to maintain secure,...  ..., Kubernetes, hypervisors, networking, and Linux systems Partner... 
    Cloud
    Network
    Temporary work
    Work at office
    Remote work
    Worldwide
    3 days per week

    GrabJobs

    Fremont, CA
    1 day ago
  • $141k - $307k

     ...platform capabilities are applied consistently.Lead solution and design reviews, and define...  ...tree learning, artificial neural networks, etc.) and their real-world advantages/drawbacksDemonstrated...  ...developing AI solutions on cloud platforms such as Azure and/or Databricks... 
    Cloud
    Network
    Work at office
    Local area
    Remote work
    Flexible hours
    2 days per week
    3 days per week
    1 day per week

    Lam Research Corporation

    Fremont, CA
    3 days ago
  • $115.2k - $172.8k

     ...with an expanded portfolio of leading-edge technologies that...  ...device performance, quality, and reliability issues. Onto Innovation strives...  ...cross functional AI champion network.Develop reference examples, patterns...  ...working across both cloud and on-prem systems. Why Join... 
    Cloud
    Network
    Permanent employment
    Full time

    Onto Innovation

    Milpitas, CA
    4 days ago
  • $265k

     ...Software Engineer to lead the architecture...  ...services, cloud infrastructure, and...  ...engineering team to ship reliable, secure, and...  ...architecture across compute, networking, service...  ...multi-tenancy, and observability, and build the systems...  ...practices and SRE principles (SLOs,... 
    Cloud
    Network
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours

    GrabJobs

    Fremont, CA
    1 day ago
  • $137.8k - $234.3k

     ...work together with the world’s leading technology providers to accelerate...  ...Lead enterprise network operational excellence by improving service reliability, availability, and performance....  ...manufacturing sites, data centers, cloud environments, and remote users.... 
    Cloud
    Network
    Remote work
    Flexible hours

    KLA-Belgium

    Milpitas, CA
    4 days ago
  • $195k - $250k

     ...Generalist who is able to lead development of a...  ...server‑based compute/networking/deployment solutions...  ...best cycles focusing on reliability enhancements,...  ...site infrastructure to cloud environments, ensuring...  ...Champion Infrastructure Observability (SRE Mindset): Propose, design... 
    Cloud
    Network
    Permanent employment
    Relocation
    Visa sponsorship
    Flexible hours
    Afternoon shift
    3 days per week

    Third Wave Automation

    Union City, CA
    2 days ago
  • $186.9k - $267.7k

     ...operations partners to ensure reliability and performance.Webex is...  ...will collaborate with technical leads and architects across the entire...  ...environments (AWS & Private Cloud). Automate Everything: Lead...  ...solutions. Add to that our worldwide network of doers and experts, and you... 
    Cloud
    Network
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours
    Shift work

    CISCO Systems

    Milpitas, CA
    3 hours ago
  • $209.7k - $272.6k

     ...transformational Director of Network Engineering to lead the strategy,...  ..., data centers, and cloud environments. This...  ...resiliency, scalability, observability, and operational...  ...to deliver secure, reliable, and high-performing...  ...platform engineering, and SRE-aligned operating... 
    Cloud
    Network
    Minimum wage

    Gap

    Pleasanton, CA
    4 days ago
  • $141k - $307k

     ...foundations that power secure, reliable, and production-ready...  ...a hands-on technical lead, you will define and...  ...services, including networking, identity, access,...  ...appropriate.Build and enhance observability across Azure AI...  ...building, and supporting cloud platforms in Microsoft... 
    Cloud
    Network
    Local area
    Remote work
    Flexible hours
    2 days per week
    3 days per week
    1 day per week

    Lam Research Corporation

    Fremont, CA
    2 days ago
  •  ...Join Our Team as a Site Reliability Engineer (SRE)! About Us At Energy Worldnet...  ...resilient systems, improving observability, and ensuring customers can...  ..., and resilience testing • Lead incident response efforts,...  ...alerting) • Experience with cloud platforms (Azure, AWS,... 
    Cloud
    Temporary work
    Casual work
    H1b
    Work at office
    Remote work
    Home office
    Monday to Friday
    Flexible hours
    Shift work

    GrabJobs

    Fremont, CA
    3 days ago
  •  ...critical data center infrastructure for AI, cloud, and hybrid cloud and advances...  ...a talented team and a strategic global network, Celestica helps its customers achieve competitive...  ...in complex, regulated and high-reliability markets such as Industrial & Smart Energy... 
    Cloud
    Network
    Worldwide
    Shift work

    Celestica International LP

    Fremont, CA
    2 days ago
  •  ..., and analytics across both cloud platforms and deployed edge...  ...driving system performance, reliability, and usability. This role is...  ...monitoring, deployment, and observability practices to ensure stable and...  ...distributed systems and network-level interactions (e.g., TCP... 
    Cloud
    Network
    Remote work

    GrabJobs

    Fremont, CA
    2 days ago
  •  ...improvements in data integrity and reliability* Develop automated reporting...  ...-built server, storage, and networking solutions to meet datacenter...  ...**TD SYNNEX (NYSE: SNX) is a leading global distributor and...  ...technology vendors. Our edge-to-cloud portfolio is anchored in some... 
    Cloud
    Network
    Worldwide

    SYNNEX Corporation

    Fremont, CA
    1 day ago
  • $210.6k - $305.1k

     ...neocloud providers, sovereign cloud environments, and...  ...complex GPU, server, network, security, and DPU-enabled...  ...to deploy, operate, observe, and scale.We are seeking...  ...Engineering Manager to lead a team working at the...  ...tradeoffs across performance, reliability, security, portability,... 
    Cloud
    Network
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    CISCO Systems

    Milpitas, CA
    3 days ago
  • $185k - $215k

     ...hands‑on Director to lead the strategy, architecture...  ...platforms across cloud, on‑premises, and high...  ...development infrastructure, observability, and the practical...  ...observability, and overall system reliability. Drive performance...  ...expertise in Linux, networking, storage, and... 
    Cloud
    Network
    Full time

    black.ai

    Milpitas, CA
    4 days ago
  •  ...manufacturing, and delivering custom Server, Storage, and Networking Solutions to the world’s largest Cloud, social media, and Enterprise companies. We pride...  ..., and IR sensor platforms to ensure operational reliability and security compliance. Experience with Axis, Hanwha... 
    Cloud
    Network
    Full time
    Work at office

    Hyve Design Solutions

    Fremont, CA
    5 days ago
  •  ...maintain modern, scalable cloud-native applications using ....  ...transaction volumes, and exceptional reliability. You will work cross-...  ...and AI-assisted testing and observability platforms as foundational...  ...Enhances relationships and networks with senior internal/... 
    Cloud
    Network
    Remote work
    Flexible hours
    Shift work

    GrabJobs

    Fremont, CA
    1 day ago
  •  ...performance and ensure system availability and reliability. Implement and maintain server...  ...scripting (Bash, Python). Familiarity with networking protocols and services. Ability to...  ...(e.g., RHCE). Experience with cloud platforms (e.g., AWS, Azure, GCP). Knowledge... 
    Cloud
    Network

    Info Way Solutions

    Fremont, CA
    1 day ago
  •  ...will have hands-on experience building cloud-native Kubernetes applications and administration...  ..., and optimize cluster performance and reliability. Automate operational tasks using...  ...Platform (preferred). Solid understanding of networking, storage, and security within Kubernetes... 
    Cloud
    Network
    Contract work
    Work experience placement
    Remote work

    GrabJobs

    Fremont, CA
    3 days ago
  •  ...support of their needs, we are looking for a Network Engineer Job Description Job Title:...  ...our campus in Fremont, CA and our public cloud providers. As a Senior Network Engineer,...  ..., focusing on scalability, reliability, and security. Perform network maintenance... 
    Cloud
    Network
    Local area
    Immediate start
    Remote work

    Maxonic

    Fremont, CA
    4 days ago
  •  ...for technology innovators building AI, cloud, and connected infrastructure. As a...  ...role in validating the performance and reliability of our cutting-edge server and...  ...procedures for server, storage, and networking hardware.* Lead the debug and root cause analysis of... 
    Cloud
    Network
    Full time
    Flexible hours

    Hyve Solutions

    Fremont, CA
    4 days ago
  • $200k - $275k

     ...bandwidth, and flexible optical networks of telecom, datacom,...  ...technically accomplished Director to lead the design and development of...  ...functional, performance, reliability, compliance, and manufacturability...  ..., from data centers and cloud networks to AI-scale infrastructure... 
    Cloud
    Network
    Flexible hours

    Molex

    Fremont, CA
    3 days ago
  • $142k - $169k

     ...technology innovators building AI, cloud, and connected infrastructure...  ..., storage platforms, networking hardware, and associated components...  ...Engineering, Quality, Reliability, Supplier Quality, Test Engineering...  ...identify corrective actions.Lead Root Cause Analysis (RCA)... 
    Cloud
    Network
    Full time
    Flexible hours

    Hyve Solutions

    Fremont, CA
    23 hours ago
  •  ...protect modern applications. Backed by leading cybersecurity investors and...  ...partnering with industry leaders across cloud, security, and observability ecosystems to transform application...  ...partners, cloud providers, and personal networks to generate pipeline and close... 
    Cloud
    Network
    Remote work

    GrabJobs

    Fremont, CA
    4 days ago
  • $186.9k - $267.7k

     ...dedicated to driving innovation in networking technologies. Our focus is on...  ..., including those in AI, cloud computing, and enterprise...  ...working relationships.• May lead projects with limited complexity...  .... Write code enabling scale, reliability, and velocity in product releases... 
    Cloud
    Network
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Milpitas, CA
    2 days ago
  • $183.8k - $263.6k

     ...are dedicated to driving innovation in networking technologies. Our focus is on developing...  ...infrastructures, including those in AI, cloud computing, and enterprise environments....  ...boosts the performance, scalability, and reliability of Cisco's Nexus switches. By engaging with... 
    Cloud
    Network
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    CISCO Systems

    Milpitas, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Observability Lead - Cloud SRE & Network Reliability. Be the first to apply!