Lead Site Reliability Engineer
$114k - $253kLam Research Salzburg GmbH
In your career, let’s prove what’s possible. At Lam Research, we create equipmentthat drives technological advancements in the semiconductor industry. Our innovative solutions enable chipmakers to power progress in nearly all aspects of modern life, and it takes each member of our team to make it possible. Across our organization, our employees come to work and change the world. We take on the toughest challenges with precision and accuracy. We push for the next big semiconductor breakthrough. We lead the way in one of the most critical and fast-moving industries on the planet. And we do it together, with deep connections and limitless collaboration. The impact we have on the world is made possible by focusing on our people. So we recognize and celebrate our teams’ achievements. We strive to create an inclusive and diverse culture where everyone’s contribution and voice has value. We evaluate and evolve our offerings, so our people receive the support and empowerment to do meaningful things for their lives, careers, and communities. Because at Lam, we believe that when people are the priority and they’re inspired to unleash the power of innovation for a better world together, anything is possible. Observability Lead - Cloud SRE & Network Reliability Date: Jul 21, 2026 Location: Fremont, CA, US, 94538 Worker Category: On-site Flex The group you’ll be a part of The Global Information Systems Group is dedicated to the success of Lam through providing best-in-class and innovative information system solutions and services. Together, we support users globally with data, information, and systems to achieve their business objectives. The impact you’ll make Our team at Lam is seeking a hands-on Observability Lead with a strong Site Reliability Engineering (SRE) and multi-cloud networking foundation to join our GIS Infrastructure Platform Engineering team. You will lead engineers in delivering robust observability frameworks, SLA/SLO/SLI disciplines, DR/BCP programs, backup and restore operations, and end-to-end network reliability across Azure, AWS, and GCP. You will own the full-stack delivery of observability, reliability, and resilience capabilities across a global multi-cloud enterprise. What you’ll do Lead and grow a team delivering a world-class observability platform across global, multi-cloud production environments, including Azure, AWS, and GCP. Define and enforce SLA, SLO, and SLI frameworks across all infrastructure and network domains, driving continuous improvement through effective error budget management. Own end-to-end multi-cloud network observability, including VNet and VPC traffic flows, Transit Gateway routing, BGP peering health, and inter-region connectivity. Design and govern multi-cloud networking architectures, including Azure VNet, AWS VPC and Transit Gateway, GCP VPC, and hybrid connectivity solutions such as ExpressRoute, Direct Connect, and Cloud Interconnect. Design and implement agentic AI workflows using LLM-based agents, RAG patterns, and orchestration frameworks to enable AIOps-driven fault detection and remediation. Own disaster recovery (DR) and business continuity planning (BCP) strategy, including runbook authorship, multi-cloud failover validation, and periodic DR drills to ensure RTO and RPO commitments are met. Lead backup and restore operations across multi-cloud and hybrid environments, incorporating automated validation and cross-cloud recovery workflows. Build robust monitoring and alerting pipelines by integrating Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, and Google Cloud Operations into a unified observability stack. Drive automation-first practices through self-healing pipelines, remediation playbooks, and infrastructure-as-code (IaC) patterns to reduce toil and improve MTTR. Lead P1, P2, and P3 incident response efforts, including structured post-mortems and action tracking. Define and drive the multi-quarter roadmap for observability, reliability, networking, DR/BCP, and AI-assisted operations. Support hiring, performance management, and career development for the team. Who we’re looking for A BS, MS, or PhD in Computer Science, Engineering, or a related field (or equivalent experience), with 12+ years of overall experience in Infrastructure, SRE, DevOps, or Network Engineering and 6+ years of experience leading high-performing SRE, Observability, or Platform Engineering teams. Proven expertise in defining, enforcing, and operating SLA, SLO, and SLI frameworks, including effective error budget management. Hands-on experience with disaster recovery (DR) and business continuity planning (BCP), including RTO/RPO planning, failover testing, and continuity documentation. Deep expertise in backup and restore operations across multi-cloud and hybrid environments. Strong multi-cloud networking skills across Azure (VNet, ExpressRoute, Virtual WAN), AWS (VPC, Transit Gateway, Direct Connect), and GCP (VPC, Cloud Interconnect, VPC-SC). Experience building and operating observability platforms, including tools such as Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Splunk, or equivalent solutions, with a focus on network telemetry and flow analysis. Deep expertise in automation, including Ansible, Terraform, Python, and self-healing infrastructure pipelines. Hands-on experience with infrastructure as code (IaC), CI/CD pipelines, Kubernetes (AKS, EKS, GKE), and all three major cloud platforms. Strong programming skills in Python or Go for tooling, automation, and system integrations. Experience leading P1, P2, and P3 incident management, including ITSM integration (ServiceNow preferred). Exceptional communication skills, with the ability to translate complex technical concepts into clear business value for engineering, product, and executive stakeholders. Preferred qualifications Experience with AIOps, including AI-assisted network fault detection, anomaly correlation, and auto-remediation. Familiarity with agentic AI workflows, including LLM-based agents and RAG patterns, applied to observability and operational use cases. Background in global WAN architectures, including MPLS and resilience strategies for multi-region enterprise environments. Experience with compliance-driven disaster recovery and business continuity (DR/BCP) programs, including InfoSec audits, SOX, and ISO 22301 requirements. Experience with FinOps and multi-cloud cost observability, including network egress visibility and cost optimization across Azure, AWS, and GCP. Relevant cloud certifications, such as Azure AZ-700 or AZ-305, AWS ANS-C01 or SAP-C02, and GCP Professional Cloud Network Engineer or Architect. Background in HPC, on-premises, or hybrid cloud environments. Our commitment We believe it is important for every person to feel valued, included, and empowered to achieve their full potential. By bringing unique individuals and viewpoints together, we achieve extraordinary results. Lam Research ("Lam" or the "Company") is an equal opportunity employer. Lam is committed to and reaffirms support of equal opportunity in employment and non-discrimination in employment policies, practices and procedures on the basis of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex (including pregnancy, childbirth and related medical conditions), gender, gender identity, gender expression, age, sexual orientation, or military and veteran status or any other category protected by applicable federal, state, or local laws. It is the Company's intention to comply with all applicable laws and regulations. Company policy prohibits unlawful discrimination against applicants or employees. Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on‑site collaboration with colleagues and the flexibility to work remotely and fall into two categories – On‑site Flex and Virtual Flex. ‘On‑site Flex’ you’ll work 3+ days per week on‑site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. ‘Virtual Flex’ you’ll work 1-2 days per week on‑site at a Lam or customer/supplier location, and remotely the rest of the time. #LI-DM1 CA San Francisco Bay Area Salary Range for this position: $114,000.00 - $253,000.00. The above salary range for this position is relevant to applicants that reside or work onsite in the California, San Francisco Bay Area only. Salary offers will depend on factors that include the location you work from, your level, education, training, specific skills, years of experience and comparison to other employees already in this role. Actual salary may vary from salary offered due to numerous factors including but not limited to unpaid time off, unpaid leave, company mandated shutdown, and other relevant factors. Our Perks and Benefits At Lam, our people make amazing things possible. That’s why we invest in you throughout the phases of your life with a comprehensive set of outstanding benefits. Nearest Major Market: San Francisco Nearest Secondary Market: Oakland Job Segment: Network, BPO, Computer Science, Recruiting, Network Engineer, Technology, Operations, Human Resources, Engineering #J-18808-Ljbffr
$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The... ...better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives...SuggestedTemporary workImmediate startFlexible hoursShift work$25.5 per hour
...FOH Lead Supervisor Morrison Living is hiring immediately for a full time FOH Lead Supervisor position. Location: Masonic Homes of CA - 34400 Mission Boulevard, Union City, CA 94587. Schedule: Full time schedule. Hours and days may vary; weekends required. Further...SuggestedHourly payFull timePart timeLocal areaImmediate startRemote workFlexible hoursShift workWeekend work$28 - $36 per hour
...microns, our equipment is top-of-the-line, and our team is passionate about quality craftsmanship. Summary The Second Shift Lead is responsible for overseeing daily operations on the second shift in a high-precision, high-mix CNC machining environment serving demanding...SuggestedHourly payAll shiftsShift workAfternoon shift$21.9 per hour
...Job Description Job Description Description: We are seeking an experienced Bilingual Janitorial Lead to assist in overseeing the daily operations of our warehouse janitorial team. This position supports management by helping to maintain efficiency across all areas...SuggestedWork experience placementFlexible hours- ...College coursework or a Bachelor's degree is preferred. Tech-savvy and comfortable using online platforms and technology. Reliable and available throughout the school year. We expect consistent attendance during our busiest program periods. Physically active...SuggestedImmediate start
$25 - $50 per hour
...Role Overview TSA is accepting applications for Lead and Supervisory Transportation Security Officers at airports in Fremont. These roles are ideal for individuals looking to step into leadership positions within airport security operations. TSA provides training to...Shift workNight shiftWeekend work$25 - $50 per hour
...Role Overview TSA is accepting applications for Lead and Supervisory Transportation Security Officers at airports in Newark. These roles are ideal for individuals looking to step into leadership positions within airport security operations. TSA provides training to...Shift workNight shiftWeekend work$140k - $160k
...seeking a passionate and talented Software Engineer with a strong interest in the... ...troubleshoot software issues, ensuring high reliability and performance. Participate in design... ...NASDAQ: NVMI) is a global company and a leading provider of innovative metrology solutions...$21 - $24.25 per hour
...Catering Lead At Panera, our people come first. If you’re looking for a place where you can grow, feel supported, be yourself, enjoy... ...tips Free on-shift meals & unlimited fountain beverages Flexible & reliable scheduling Paid vacation, sick time, and holidays for full-time...Full timeLocal areaFlexible hoursShift workNight shift$26.75 - $29.25 per hour
...regularly and respectfully with parents/guardians Manage a host site relationship (where applicable) Supply management and... ...pressure and able to calm those around you? Are you comfortable leading groups of kids on your own while still collaborating with a team...Hourly paySummer workShift work- .... We are committed to being America's best first job. Let's talk. Make your move. See a day in the life of a Guest Experience Lead at McDonald's Requirements: We believe in letting you do you. If you're looking for a part-time job that supports your full-...Full timePart timeLocal area
- Our Team Leads are the ones who "make it happen". You will be responsible for running shifts of 2-8 team members when the GM is not present... ...procedures and policies Coordinate and participate off site program customer visits and deliveries Requirements: Must...All shiftsFlexible hoursShift work
- ...recruitment efforts, working closely with HR and stakeholders, and necessitates 50% travel within the region. Ideal candidates should have 7-10 years of recruitment experience, particularly in engineering or construction, and hold a Bachelor’s Degree. #J-18808-Ljbffr...Remote work
$22.65 - $24 per hour
...Job Overview The Custodial Lead will be responsible for the cleanliness and sanitation of the areas assigned and provides some work direction to custodial staff. Roles & Responsibilities Perform janitorial duties Perform all duties listed on the daily schedule Provide...Hourly payImmediate startShift work$22.65 - $24 per hour
...Description Position at SBM Management SBM Management is currently looking to hire a Custodial Lead to join their team! The Custodial Lead has responsibilities for overseeing activities within the assigned program. This includes the company employees...Hourly payTemporary workCurrently hiringImmediate startShift work- Shipping And Logistics Manager Manage daily shipping and logistics operations to ensure on-time customer deliveries. Coordinate domestic and international shipments using multiple carriers and freight forwarders. Ensure compliance with export regulations, customs...
$165k - $200k
...are equally at home at a camp site, a job site, or on a Tuesday... ...ll do As a Staff Software Engineer (Android) on the Digital Products... ...consistency, performance, reliability, and maintainability across... ...a technical leader, you will lead large-scale, complex initiatives...Full timeWork at officeImmediate startFlexible hours- Quanta Manufacturing Fremont is seeking a skilled Project Manager to manage manufacturing projects of moderate complexity. The successful candidate will own project execution, drive cross-functional coordination, and ensure that quality, schedule, and delivery commitments...
- ...Tesla, Waymo, and Zipline). About the Role As Human Resources Lead, you will be responsible for building the company’s people function... ...Team and Culture World-Class Team: Work alongside exceptional engineers, AI researchers, and healthcare experts Direct Access: Report...Shift work
$141k - $307k
...next big semiconductor breakthrough. We lead the way in one of the most critical and fast... ..., CA, US, 94538 Worker Category: On-site Flex The group you’ll be a part of The Office... ...product managers, AI governance, and engineering teams to take solutions from architecture...Work at officeLocal areaRemote workFlexible hours2 days per week3 days per week1 day per week- ...Woodgrain is hiring a (Lead) Merchandise Stocker for Fremont, CA, covering Union City among other locations. The role involves managing store inventory, ensuring excellent customer service, and requires a valid driver's license and heavy lifting ability. Benefits include...
- ...Lead Product Manager At NovaSensor, an Amphenol company, we are pioneers in high... ...vehicle safety and efficiency, or ensuring reliability in industrial systems. We are part of... ...developments. Collaborate with Engineering and the cross-functional project teams to...Work experience placementWorldwideFlexible hours
$119k - $200k
...various divisions within the company. We are looking for versatile engineers who are interested in architecting and implementing elegant... ..., but in the end, you know that what matters is delivering reliable solutions. (Our ultimate aim is to help people; the “right” solution...Full timeTemporary workFlexible hours$35 per hour
...members have diverse backgrounds in software engineers, design engineering, ML engineering and... ...user experiences. You will take the lead in creating innovative applications, implementing... ...elegant, maintainable, performant and reliable user-facing software applications You...Hourly payFull timeTemporary workInternshipFlexible hours- ...cross functional teams to design and develop software programs. Provide technical guidance and mentoring for more junior engineers. May visit customer site to provide support and have ability to travel (total is less than 10%). Bachelor's degree in Computer Engineering,...Full time
$49.92k - $58.24k
...Sanmina Corporation (Nasdaq: SANM) is a leading integrated manufacturing solutions provider serving the fastest-growing segments of... ...Proficiency with soldering (IPC-J-STD-001) and interpreting schematics/engineering drawings. Strong communication, leadership, and problem-...Hourly payPermanent employmentWork experience placementDay shift$154k - $211.75k
...Leading the future in luxury electric and mobility At Lucid, we set out to introduce... ...Responsibilities As a key member of our engineering team, you would be designing,... ...software components and outgoing traffic is reliably transported to reach their intended destinations...Full timeImmediate start- Partner with product management and engineering leadership to define long-term platform strategy. Design and build scalable, fault tolerant microservices using Python and modern cloud patterns. Architect systems using Kubernetes, container orchestration, and infrastructure...Full timeShift work
$99k - $220k
...n## The impact you\u2019ll make\n\nJoin our team as a software engineer to build the Dextro Software platform. Dextro\u2122 - Lam Research... ...from concept through completion.\n * May visit customer sites to provide support; travel is expected to be less than 10%.\n\n...Full timeLocal areaRemote workFlexible hours2 days per week3 days per week1 day per week- ...Reliability Engineer – (Pivotal Systems) The Reliability Engineer is responsible for developing and executing reliability programs for semiconductor... ...analysis to predict product life and performance. Lead or support Failure Modes and Effects Analysis (FMEA), reliability...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- construction site safety Fremont, CA
- on-site clinical research associate (traveling/remote) Fremont, CA
- site safety Fremont, CA
- historic site Fremont, CA
- IT site lead Fremont, CA
- site leader Fremont, CA
- junior website developer Fremont, CA
- official site Fremont, CA
- site services specialist Fremont, CA
- lead backend developer





