Senior Manager, Reliability Engineering & AIOps
$137k - $287kLam Research Corporation
The group you’ll be a part ofYou will join the Reliability Engineering team within Infrastructure Platform Engineering. The group keeps Lam's global infrastructure estate available and recoverable across Azure, AWS, GCP, compute, storage, network, and high-performance computing, supporting engineering and operations teams in the US, Japan, Singapore, Malaysia, India, and Korea.The impact you’ll makeAs Senior Manager of Reliability Engineering & AIOps, you lead the team that keeps critical infrastructure running and proves it is ready for the next failure. In this role, you will directly contribute to the availability of the systems Lam's engineering, manufacturing, and business teams depend on every day, and you will build the automation that makes outages rare, short, and unremarkable. What you’ll doLead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on. Set the reliability strategy: define the service level objective program, publish an error-budget policy, and drive adoption across platform and service teams. Build and run a follow-the-sun on-call and response model across six regions, with clean handoffs and one consistent set of runbooks and severity definitions worldwide. Own the incident management and paging platform end to end, including services, schedules, escalation policies, and routing, configured as code and tuned so alerts fire on real risk rather than noise. Serve as incident commander on major incidents, own executive and stakeholder communications, and lead blameless postmortems with tracked follow-up. Own disaster recovery strategy and execution across Azure, AWS, GCP, and core infrastructure platforms, including service-tier recovery objectives, backup and restore validation, failover readiness, DR certification, runbook governance, and recurring exercises measured against RTO and RPO targets. Lead capacity planning and performance engineering across Azure, AWS, GCP, compute, storage, network, and HPC platforms, using demand forecasting, utilization trends, growth modeling, and automation to prevent capacity risk and reduce manual operational work. Define and drive AI Ops requirements for reliability engineering across Azure, AWS, and GCP, including Microsoft Copilot, Cursor, GitHub Copilot, and LLM-based operational workflows for incident triage, runbook generation, knowledge retrieval, root-cause analysis, and safe remediation recommendations. This is a full-time role on a standard schedule, with participation in a global on-call rotationWho we’re looking forBachelor's degree in Computer Science, Engineering, or a related field with 10 years of related experience; or a Master's degree with 8 years of experience; or equivalent experience. Experience leading or mentoring a reliability or operations team and setting technical direction. Proven incident command on major outages, plus ownership of a postmortem process. Strong background in disaster recovery planning across Azure, AWS, GCP, and core infrastructure platforms, including restore validation, failover testing, recovery-objective definition, and corrective action tracking after DR exercises or production incidents. Hands-on ownership of an incident management and paging platform at scale, such as PagerDuty. Experience with capacity planning, performance trending, utilization analysis, and infrastructure demand forecasting for globally distributed production environments across Azure, AWS, GCP, and on-premises platforms. Track record of defining and defending service level objectives and error budgets in production. Working depth in observability tooling (Prometheus, Grafana, Loki, Tempo or equivalent), infrastructure as code (Terraform), and Python or Go. Practical experience applying AI-assisted engineering and operations tools such as Microsoft Copilot, Cursor, GitHub Copilot, or enterprise LLM platforms to improve troubleshooting, automation, documentation, and engineering productivity across Azure, AWS, GCP, and hybrid infrastructure, with clear guardrails for security, privacy, auditability, and production safety. Preferred qualificationsExperience running global, follow-the-sun operations across multiple regions and time zones. Capacity and performance engineering at multi-region scale, including Azure, AWS, GCP, high-performance computing, large storage estates, hybrid cloud infrastructure, and proactive capacity governance. Policy as code, progressive delivery, and chaos engineering in practice. Experience building or operating AI and agent-assisted automation in operations, with a clear view of its failure modes. Experience designing or operating AI Ops capabilities across Azure, AWS, GCP, and hybrid environments, including LLM-grounded knowledge bases, agent-assisted incident workflows, prompt and evaluation practices, and supervised automation that can recommend or propose operational changes before execution. Our commitmentWe believe it is important for every person to feel valued, included, and empowered to achieve their full potential. By bringing unique individuals and viewpoints together, we achieve extraordinary results.Lam Research ("Lam" or the "Company") is an equal opportunity employer. Lam is committed to and reaffirms support of equal opportunity in employment and non-discrimination in employment policies, practices and procedures on the basis of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex (including pregnancy, childbirth and related medical conditions), gender, gender identity, gender expression, age, sexual orientation, or military and veteran status or any other category protected by applicable federal, state, or local laws. It is the Company's intention to comply with all applicable laws and regulations. Company policy prohibits unlawful discrimination against applicants or employees.Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on-site collaboration with colleagues and the flexibility to work remotely and fall into two categories – On-site Flex and Virtual Flex. ‘On-site Flex’ you’ll work 3+ days per week on-site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. ‘Virtual Flex’ you’ll work 1-2 days per week on-site at a Lam or customer/supplier location, and remotely the rest of the time.#LI-DM1SalaryCA San Francisco Bay Area Salary Range for this position: $137,000.00 - $287,000.00.The above salary range for this position is relevant to applicants that reside or work onsite in the California, San Francisco Bay Area only. Salary offers will depend on factors that include the location you work from, your level, education, training, specific skills, years of experience and comparison to other employees already in this role. Actual salary may vary from salary offered due to numerous factors including but not limited to unpaid time off, unpaid leave, company mandated shutdown, and other relevant factors.Our Perks and BenefitsAt Lam, our people make amazing things possible. That’s why we invest in you throughout the phases of your life with a comprehensive set of outstanding benefits.Department:Information Systems
$170k - $220k
...Senior Engineer – OpenSearch DTEX is seeking a highly skilled and experienced Senior Engineer – OpenSearch to join our engineering team... ...-scale OpenSearch clusters, contribute to our Insider Risk Management (IRM) and Data Loss Prevention (DLP) products, and actively...SeniorRemote workWorldwide$155k - $175k
Tracking Code4155Job DescriptionPosition Summary Watch your engineering designs come to life! As a Senior Associate Principal Engineer at Southland Industries, you'll experience the excitement of working on design-build and design-assist projects in healthcare, education...SeniorWork at officeImmediate startFlexible hours- ...Senior Director Of Engineering InterSources Inc seeks a Senior Director of Engineering, reporting directly to Vice President for our Professional... ...tools that support a customer and related Network Management Systems Troubleshoot and resolve highly complex customer...Senior
- ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available database... ...environments in production and cloud environments. Manage PostgreSQL deployments on Kubernetes and Amazon RDS....Senior
$125k - $270k
LAM RESEARCH Corporation is seeking a Supervisor for the Fremont Lab Operations Group to manage a team of 30 technicians and engineers. This role involves overseeing technical tests, setting expectations, and ensuring efficient workflow in a clean room environment. The...Senior$160k - $240k
...most difficult yield, device performance, quality, and reliability issues. Onto Innovation strives to optimize customers’... ...are seeking a highly skilled and innovative Senior Manager, Mechanical Engineering to join our team in leading and developing advanced hardware...SeniorPermanent employmentFull time- ...innovative and growing medical technology company is seeking a Reliability Quality Engineer to ensure the product reliability and quality of a... ...ensure proper linkage between design requirements, risk management, reliability verification, and validation activities....SeniorPermanent employmentFull time
$81.1k - $187k
...professionals for this role. Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production... ...Key Responsibilities Capacity Ingestion and Management: -Takes proactive steps to design and architect infrastructure...SeniorTemporary workImmediate startFlexible hoursShift work- A leading IT services company in Pleasanton is seeking a Senior Site Reliability Engineer / DevOps Engineer to manage AWS infrastructure and automation. The ideal candidate will have extensive experience in AWS cloud environments, infrastructure as code, and strong Linux...SeniorFull time
- ...activities in collaboration with foundries, internal reliability teams, and product engineering organizations. Track, analyze, and resolve wafer-level... ...Statistical Process Control (SPC) implementation, excursion management procedures, change control practices, and corrective/...Senior
$150k - $225k
...as part of a team. If that's you and this role fits, we want to hear from you. Join the power movement as Senior Manager, Process Engineering We're looking for a Senior Manager, Process Engineering to join our team. The Process Engineering Manager will...SeniorFull timeFlexible hoursShift workNight shiftWeekend work$167.3k - $284.4k
...expert teams of physicists, engineers, data scientists and problem-... ...application development engineers, and senior product technology process... .../Preferred QualificationsSr. Reliability Engineer - SEM SystemsJoin a... ...conferencing with our hiring managers. If you are concerned that a...SeniorMinimum wageFull timeWork experience placementFlexible hours$166k - $350k
...Products Group, we are dedicated to excellence in the design and engineering of Lam's etch and deposition products. We drive innovation to... ...focused on creating test code for a large software system.Manages the software lifecycle for the Test Farms, including the management...SeniorLocal areaRemote workFlexible hours2 days per week3 days per week1 day per week- ...Reliability Engineer Play a critical role in the development of new life science technology as a reliability engineer in the R&D group. Provide... ...improve product reliability throughout the design process. Manage reliability projects by implementing tests, analyzing data,...Senior
- ...Reliability Engineer – (Pivotal Systems) The Reliability Engineer is responsible for developing and executing reliability programs for... ...Generate reliability reports, metrics, and recommendations for management and engineering teams. Participate in change control...
$209.6k - $314.4k
...and a leading AI platform for managing people, money, and agents, we... ...built. We’re forming small, senior, cross-functional AI teams that... ...together product leaders, AI engineers, and full-stack builders to... ...reliabilityEnsure the scalability, reliability, and security of AI agent...SeniorFull timeWork at officeRemote workHome officeFlexible hours$125k - $155k
Kimley-Horn's Pleasanton office is looking for a Civil Engineer with over 6 years of experience. The role requires expertise in site development engineering and project management. Key responsibilities include overseeing land development projects, managing project budgets...SeniorWork at office- ...Job Description Job Description Job Title - Reliability Equipment Engineer, Fremont, CA Role Description We are seeking a Reliability... ...create dashboards. ● Hardware Failure Isolation & Vendor Management: Perform initial hardware-level troubleshooting to...Temporary work
$274.46k - $296.8k
...around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so... ...GA; and Dublin, Ireland, the Data Platform and Observability Engineering team is vital to enabling real‑time insights across Workday’s...SeniorWork at officeRemote workFlexible hours- Google is seeking a Senior CMOS Test and Validation Lead in Fremont, CA. This role involves... ...-Signal ASICs. You will lead a team of engineers and ensure zero-defect quality. The... ...experience in test engineering, CMOS, and vendor management. A Bachelor's degree in a relevant...Senior
- ...Roles & Responsibilities Lead and monitor multiple reliability projects Provide reliability risk assessments to cross-functional... ...action identification, implementation, and validation Provide engineering input for new designs to ensure reliable, robust products...SeniorFull time
- ...Senior - Principal Engineer - Wastewater Collection System Planning The Senior - Principal Engineer role will contribute to the development... ...their career toward technical leadership, client service management, and/or regional/sector management are ideal. Key Responsibilities...SeniorTemporary workWork experience placementWork at officeLocal areaRemote workFlexible hoursNight shiftAfternoon shift
- ...cloud-native systems. As a Staff Platform Engineer, you will play a critical role in... ...technical leadership role. You will own reliability for major platform domains, design scalable... ...applications Architect, implement, and manage highly available and scalable Kubernetes...Senior
$157k - $271.4k
...searching for the best talent to join our Vision team as a Senior Manager, R&D Software Engineering located in Milpitas, CA.Fueled by innovation at the... ...ability in software / firmware development delivering reliable, testable and maintainable code for embedded systemExperience...SeniorFull timeLocal areaImmediate start$150k - $195k
Fremont, CAUS Engineering - Hardware Engineering /Full-time /On-sitePlusAI is a Physical AI company pioneering AI-based virtual driver... ...your work is performed in accordance with the company’s Quality Management System (QMS) requirements and contribute to continuous...SeniorFull time- ...technology company in Fremont, California, is seeking a Systems & Controls Engineering Lead. This is a high-impact leadership role that combines technical expertise in systems engineering with team management responsibilities. The ideal candidate will have a Master's degree in...Senior
- ...Design and develop fiber optic transceivers, CPO and NPO Optical Engines. Drive optical engine product design activities from initial... ...drawings, BOMS and ECO to document design. Interact with program manager, buyer/planner, suppliers to ensure designed parts are...SeniorFull timeWork at office
- Electrical Engineer / Electrical SpecialistFacilities Engineering - Manufacturing PlantPosition... ...Engineering team by ensuring the safe, reliable, and efficient operation of electrical... ...power distribution systemsExperience managing or overseeing electrical contractorsPLC,...SeniorFull timeFor contractorsWork at officeShift work
- ...Hi Friends, I am sending requirement, kindly get back to me if the job description suits you. Position: Senior Mechanical Engineer Client: Hyzon Motors Location: Troy, MI Pay Rate: Open to Market Rate Visa: USC an GC Only (Open to relocate, but...SeniorRelocation
$90k - $130k
...passionate and forward-thinking experts. We’re one of the largest engineering and system integration firms in the United States providing... ...at that site. We are seeking an enthusiastic and experienced Senior Controls Engineer to be responsible for leading or contributing...SeniorLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Manager, Reliability Engineering & AIOps. Be the first to apply!
- srs distribution Fremont, CA
- senior dynamics crm developer Fremont, CA
- senior application security Fremont, CA
- senior advisor Fremont, CA
- senior cloud data engineer Fremont, CA
- senior Fremont, CA
- senior customer success engineer Fremont, CA
- senior property accountant Fremont, CA
- senior business development Fremont, CA
- senior operations technician Fremont, CA


