Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$114k - $253k

Lam Research Salzburg GmbH

In your career, let’s prove what’s possible. At Lam Research, we create equipmentthat drives technological advancements in the semiconductor industry. Our innovative solutions enable chipmakers to power progress in nearly all aspects of modern life, and it takes each member of our team to make it possible. Across our organization, our employees come to work and change the world. We take on the toughest challenges with precision and accuracy. We push for the next big semiconductor breakthrough. We lead the way in one of the most critical and fast-moving industries on the planet. And we do it together, with deep connections and limitless collaboration. The impact we have on the world is made possible by focusing on our people. So we recognize and celebrate our teams’ achievements. We strive to create an inclusive and diverse culture where everyone’s contribution and voice has value. We evaluate and evolve our offerings, so our people receive the support and empowerment to do meaningful things for their lives, careers, and communities. Because at Lam, we believe that when people are the priority and they’re inspired to unleash the power of innovation for a better world together, anything is possible. Observability Lead - Cloud SRE & Network Reliability Date: Jul 21, 2026 Location: Fremont, CA, US, 94538 Worker Category: On-site Flex The group you’ll be a part of The Global Information Systems Group is dedicated to the success of Lam through providing best-in-class and innovative information system solutions and services. Together, we support users globally with data, information, and systems to achieve their business objectives. The impact you’ll make Our team at Lam is seeking a hands-on Observability Lead with a strong Site Reliability Engineering (SRE) and multi-cloud networking foundation to join our GIS Infrastructure Platform Engineering team. You will lead engineers in delivering robust observability frameworks, SLA/SLO/SLI disciplines, DR/BCP programs, backup and restore operations, and end-to-end network reliability across Azure, AWS, and GCP. You will own the full-stack delivery of observability, reliability, and resilience capabilities across a global multi-cloud enterprise. What you’ll do Lead and grow a team delivering a world-class observability platform across global, multi-cloud production environments, including Azure, AWS, and GCP. Define and enforce SLA, SLO, and SLI frameworks across all infrastructure and network domains, driving continuous improvement through effective error budget management. Own end-to-end multi-cloud network observability, including VNet and VPC traffic flows, Transit Gateway routing, BGP peering health, and inter-region connectivity. Design and govern multi-cloud networking architectures, including Azure VNet, AWS VPC and Transit Gateway, GCP VPC, and hybrid connectivity solutions such as ExpressRoute, Direct Connect, and Cloud Interconnect. Design and implement agentic AI workflows using LLM-based agents, RAG patterns, and orchestration frameworks to enable AIOps-driven fault detection and remediation. Own disaster recovery (DR) and business continuity planning (BCP) strategy, including runbook authorship, multi-cloud failover validation, and periodic DR drills to ensure RTO and RPO commitments are met. Lead backup and restore operations across multi-cloud and hybrid environments, incorporating automated validation and cross-cloud recovery workflows. Build robust monitoring and alerting pipelines by integrating Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, and Google Cloud Operations into a unified observability stack. Drive automation-first practices through self-healing pipelines, remediation playbooks, and infrastructure-as-code (IaC) patterns to reduce toil and improve MTTR. Lead P1, P2, and P3 incident response efforts, including structured post-mortems and action tracking. Define and drive the multi-quarter roadmap for observability, reliability, networking, DR/BCP, and AI-assisted operations. Support hiring, performance management, and career development for the team. Who we’re looking for A BS, MS, or PhD in Computer Science, Engineering, or a related field (or equivalent experience), with 12+ years of overall experience in Infrastructure, SRE, DevOps, or Network Engineering and 6+ years of experience leading high-performing SRE, Observability, or Platform Engineering teams. Proven expertise in defining, enforcing, and operating SLA, SLO, and SLI frameworks, including effective error budget management. Hands-on experience with disaster recovery (DR) and business continuity planning (BCP), including RTO/RPO planning, failover testing, and continuity documentation. Deep expertise in backup and restore operations across multi-cloud and hybrid environments. Strong multi-cloud networking skills across Azure (VNet, ExpressRoute, Virtual WAN), AWS (VPC, Transit Gateway, Direct Connect), and GCP (VPC, Cloud Interconnect, VPC-SC). Experience building and operating observability platforms, including tools such as Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Splunk, or equivalent solutions, with a focus on network telemetry and flow analysis. Deep expertise in automation, including Ansible, Terraform, Python, and self-healing infrastructure pipelines. Hands-on experience with infrastructure as code (IaC), CI/CD pipelines, Kubernetes (AKS, EKS, GKE), and all three major cloud platforms. Strong programming skills in Python or Go for tooling, automation, and system integrations. Experience leading P1, P2, and P3 incident management, including ITSM integration (ServiceNow preferred). Exceptional communication skills, with the ability to translate complex technical concepts into clear business value for engineering, product, and executive stakeholders. Preferred qualifications Experience with AIOps, including AI-assisted network fault detection, anomaly correlation, and auto-remediation. Familiarity with agentic AI workflows, including LLM-based agents and RAG patterns, applied to observability and operational use cases. Background in global WAN architectures, including MPLS and resilience strategies for multi-region enterprise environments. Experience with compliance-driven disaster recovery and business continuity (DR/BCP) programs, including InfoSec audits, SOX, and ISO 22301 requirements. Experience with FinOps and multi-cloud cost observability, including network egress visibility and cost optimization across Azure, AWS, and GCP. Relevant cloud certifications, such as Azure AZ-700 or AZ-305, AWS ANS-C01 or SAP-C02, and GCP Professional Cloud Network Engineer or Architect. Background in HPC, on-premises, or hybrid cloud environments. Our commitment We believe it is important for every person to feel valued, included, and empowered to achieve their full potential. By bringing unique individuals and viewpoints together, we achieve extraordinary results. Lam Research ("Lam" or the "Company") is an equal opportunity employer. Lam is committed to and reaffirms support of equal opportunity in employment and non-discrimination in employment policies, practices and procedures on the basis of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex (including pregnancy, childbirth and related medical conditions), gender, gender identity, gender expression, age, sexual orientation, or military and veteran status or any other category protected by applicable federal, state, or local laws. It is the Company's intention to comply with all applicable laws and regulations. Company policy prohibits unlawful discrimination against applicants or employees. Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on‑site collaboration with colleagues and the flexibility to work remotely and fall into two categories – On‑site Flex and Virtual Flex. ‘On‑site Flex’ you’ll work 3+ days per week on‑site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. ‘Virtual Flex’ you’ll work 1-2 days per week on‑site at a Lam or customer/supplier location, and remotely the rest of the time. #LI-DM1 CA San Francisco Bay Area Salary Range for this position: $114,000.00 - $253,000.00. The above salary range for this position is relevant to applicants that reside or work onsite in the California, San Francisco Bay Area only. Salary offers will depend on factors that include the location you work from, your level, education, training, specific skills, years of experience and comparison to other employees already in this role. Actual salary may vary from salary offered due to numerous factors including but not limited to unpaid time off, unpaid leave, company mandated shutdown, and other relevant factors. Our Perks and Benefits At Lam, our people make amazing things possible. That’s why we invest in you throughout the phases of your life with a comprehensive set of outstanding benefits. Nearest Major Market: San Francisco Nearest Secondary Market: Oakland Job Segment: Network, BPO, Computer Science, Recruiting, Network Engineer, Technology, Operations, Human Resources, Engineering #J-18808-Ljbffr

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Fremont, CA vacancy
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Suggested
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Pleasanton, CA
    1 day ago
  • $84.9k - $209.5k

     ...Health Data Intelligence Platform Engineer Building off our Cloud momentum, Oracle has formed a new organization - Health Data Intelligence Platform. This team will focus on product development and product strategy for Oracle Health, while building out a complete platform... 
    Suggested
    Temporary work
    Immediate start
    Flexible hours

    Oracle

    Pleasanton, CA
    3 days ago
  •  ...Reliability Engineer – (Pivotal Systems) The Reliability Engineer is responsible for developing and executing reliability programs for semiconductor process control products and components. This role works closely with engineering, manufacturing, quality, and suppliers... 
    Suggested

    Pivotal Systems

    Fremont, CA
    4 days ago
  • $154k - $211.75k

     ...place for you. Responsibilities As a key member of our engineering team, you would be designing, implementing, and testing software...  ...distributed among software components and outgoing traffic is reliably transported to reach their intended destinations. Being able... 
    Suggested
    Full time
    Immediate start

    Lucid Motors

    Newark, CA
    8 hours ago
  • Partner with product management and engineering leadership to define long-term platform strategy. Design and build scalable, fault tolerant microservices using Python and modern cloud patterns. Architect systems using Kubernetes, container orchestration, and infrastructure... 
    Suggested
    Full time
    Shift work

    Lam Research

    Fremont, CA
    8 hours ago
  • $138k - $269k

     ...feedback about new features from users. The work in the team is multidisciplinary and team members have diverse backgrounds in software engineers, design engineering, ML engineering and neuro-engineering. Job Description and Responsibilities: As a Software Engineer in... 
    Full time
    Temporary work
    Flexible hours

    Neuralink

    Fremont, CA
    8 hours ago
  • $35 per hour

     ...various divisions within the company. We are looking for versatile engineers who are interested in architecting and implementing elegant...  ..., but in the end, you know that what matters is delivering reliable solutions. (Our ultimate aim is to help people; the “right” solution... 
    Hourly pay
    Full time
    Temporary work
    Internship
    Flexible hours

    Neuralink

    Fremont, CA
    8 hours ago
  • $99k - $220k

     ...n## The impact you\u2019ll make\n\nJoin our team as a software engineer to build the Dextro Software platform. Dextro\u2122 - Lam Research...  ...from concept through completion.\n * May visit customer sites to provide support; travel is expected to be less than 10%.\n\n... 
    Full time
    Local area
    Remote work
    Flexible hours
    2 days per week
    3 days per week
    1 day per week

    Lam Research

    Fremont, CA
    8 hours ago
  •  ...cross functional teams to design and develop software programs. Provide technical guidance and mentoring for more junior engineers. May visit customer site to provide support and have ability to travel (total is less than 10%). Bachelor's degree in Computer Engineering,... 
    Full time

    Lam Research

    Fremont, CA
    8 hours ago
  • $135k - $216k

     ...and there are stretches where the pace is high. This isn't for everyone. Job Description and Responsibilities: As a Software Engineer on the Robot Manufacturing Team, you'll work directly with robot engineers and surgical engineers to understand what they need,... 
    Full time
    Temporary work
    Work experience placement
    Flexible hours
    Day shift

    Neuralink

    Fremont, CA
    8 hours ago
  •  ...relocation to the area. We recruit nationally and provide financial relocation assistance. Responsibilities As a Technical Solutions Engineer at Epic, you’ll work on software that impacts 305 million patients around the world. Together with customer counterparts, you’ll... 
    Relocation
    Visa sponsorship
    Relocation package

    Epic

    Fremont, CA
    1 day ago
  • $170k - $220k

     ...Join to apply for the Platform Engineer role at DTEX. DTEX Systems helps hundreds of organizations worldwide better understand their workforce, protect their data, and make human‑centric operational investments. At DTEX, our philosophy towards our business is the... 
    Full time
    Remote work
    Work from home
    Worldwide
    Flexible hours

    DTEX

    Fremont, CA
    3 days ago
  •  ...Senior AI Engineer Dexmate is building the foundation for physical AI — a unified platform that combines high-quality robotic hardware...  ...loops, observability, and lifecycle control) that makes agents reliable in production; the model is a component, the harness is the... 

    Dexmate

    Fremont, CA
    3 days ago
  • $196k

     ...infrastructure is evolving to meet the needs of our fast growing engineering org. We are looking for backend engineers to join our team to...  ...in a new region ~ Building custom Kubernetes operators for reliably managing some of our most critical workloads Data... 
    For contractors
    Work at office
    Remote work
    Relocation
    Flexible hours

    GrabJobs

    Fremont, CA
    5 days ago
  •  ...but our team is uniquely positioned to do it. Some of the best engineers, design engineers, and growth engineers in the world. World‑...  ...data models, queues, workflows, and infrastructure that support reliable product experiences at scale Publishing & performance — Help... 

    Energy Jobline ZR

    Fremont, CA
    9 hours ago
  •  ...We are seeking a Senior DSP Software Engineer to join a small, high-impact embedded software team at a product development company specializing in advanced spectrum monitoring solutions. These platforms serve both defense applications and commercial spectrum monitoring... 

    BrightHire Search Partners

    Fremont, CA
    5 days ago
  • $140k - $160k

     ...We are seeking a passionate and talented Software Engineer with a strong interest in the semiconductor manufacturing industry to support...  .... Analyze and troubleshoot software issues, ensuring high reliability and performance. Participate in design reviews, code reviews,... 

    NOVA

    Fremont, CA
    3 days ago
  •  ...largest companies make smarter, fairer, and more transparent pay decisions—powered by live data and thoughtful design. As a Software Engineer, you’ll have significant ownership and autonomy to build the systems and infrastructure that power Compa’s core products. You’ll... 
    Remote work

    GrabJobs

    Fremont, CA
    5 days ago
  • $125k - $160k

     ...network devices. Work with REST APIs, Docker and Microservices to support the automation testing Collaborate with Network Engineering, DevOps, Test Engineering to support manufacturing process Troubleshoot automation-related issues and optimize performance.... 
    Work experience placement

    Hyve Solutions

    Fremont, CA
    2 days ago
  •  ...platforms. Collaborate closely with data, platform, and other engineers to ensure seamless, end-to-end integration of new tools into the...  ...Ability to travel domestically and internationally. Work can be on site or in hybrid mode Why us? Working at Siemens Software means... 
    Full time
    Work at office
    Work from home

    Siemens AG

    Fremont, CA
    5 days ago
  •  ...world of chip, board, and system design. Position Overview Join the QuestaSim (Simulation) R&D team at Siemens EDA as a Principal Engineer and drive innovation in simulation technology. In this role, you'll architect cutting-edge algorithms and software solutions that... 

    Siemens EDA (Siemens Digital Industries Software)

    Fremont, CA
    2 days ago
  • $140k - $190k

     ...Your Job As a Senior Software Engineer, this person will lead application design in Python for test frameworks, optical transceiver system and parametric tests. This Engineer will also create architecture of test software and stations through design, layout, hardware... 
    Flexible hours

    Molex

    Fremont, CA
    9 hours ago
  • $129.6k - $233.3k

     ...Senior Software Engineer - VSD - Freemont CA Hybrid Job ID 515221 Posted since 23-Jul-2026 Organization Field of work Research & Development...  ...to travel domestically and internationally. Work can be on site or in hybrid mode Why us? Working at Siemens Software means flexibility... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    Siemens Mobility

    Fremont, CA
    5 days ago
  • $80 - $100 per hour

     ...evaluation pipelines used to test frontier AI models on real software engineering work: Design coding benchmarks that evaluate frontier...  ...workflows Analyze model-generated code for correctness, reliability, and edge-case failures Construct structured evaluation... 
    Full time
    Contract work
    For contractors
    Remote work

    GrabJobs

    Fremont, CA
    9 hours ago
  •  ...Java Full stack developers, Python/Java developers, Data analysts/ Data Scientists. Who Should Apply: Recent Computer Science/Engineering /Mathematics/Statistics or Science Graduates looking to make their careers in IT Industry We welcome candidates with all visas... 
    H1b
    Remote work

    SynergisticIT

    Fremont, CA
    4 days ago
  • $100k

     ...level software programmers, Java full-stack developers, Python/Java developers, data analysts/data scientists, and machine learning engineers for full-time positions with clients. Who should apply? Recent computer science/engineering/mathematics/statistics or science... 
    Full time
    H1b
    Remote work

    SynergisticIT

    Fremont, CA
    5 days ago
  •  ...more) Please apply and see more job requisitions at: Essential Duties and Responsibilities: Transcard is seeking a Senior Software Engineer to join our mixed local and remote team. The ideal candidate will write, test, secure, and maintain code for our suite of payments... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to Friday

    GrabJobs

    Fremont, CA
    3 days ago
  • $226.14k

     ...Software Engineer - SWE2501 Intelliswift Software, Inc. is seeking a qualified professional to fill the position of Software Engineer...  ...and modify software programs to ensure technical accuracy, reliability and scalability of programs. Analyze specifications,... 

    Intelliswift

    Newark, CA
    2 days ago
  •  ...Jr Software Engineer To assist in the preparation of plans, scripts, measures, and other items needed for testing while applying company and industry standards. Review requirements and provide inputs to Leads. Build test cases, test data & ensure test environment is... 

    Keylent Inc

    Fremont, CA
    5 days ago
  •  ...Software Engineer In the Global Products Group, we are dedicated to excellence in the design and engineering of Lam's etch and deposition...  ...and mentoring for more junior engineers May visit customer site to provide support and have ability to travel (total is less... 
    Local area
    Remote work
    Flexible hours
    2 days per week
    3 days per week
    1 day per week

    Lam Research

    Fremont, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!