Senior Site Reliability Engineer
JUUL Labs
Senior Site Reliability Engineer
The company: Juul Labs's mission is to transition the world's billion adult smokers away from combustible cigarettes, eliminate their use, and combat underage usage of our products. We have the opportunity to address one of the world's most intractable challenges through a commitment to exceptional quality, research, design, and innovation. Backed by leading technology investors, we are committed to the same excellence when it comes to hiring great talent.
We are a diverse team that is united by this common purpose and we are hiring the world's best engineers, scientists, designers, product managers, operations experts, and customer service and business professionals. If the opportunity to build your career is compelling, read on for more details.
Must live in either Mountain View, CA, Washington D.C., Austin, TX, or Durham, NC.
Role and Responsibilities
A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance of Juul's hybrid cloud infrastructure (Nutanix, AWS/GCP). This involves leading automation efforts, architecting for reliability, and acting as the final escalation point for critical incidents to ensure the platform is scalable and efficient.
Nutanix Platform Management
- Design, deploy, and maintain enterprise-scale Nutanix AHV clusters and Prism Central for multi-cluster management
- Expert-level proficiency with Nutanix CLI (nCLI and acli) for advanced operations, troubleshooting, and automation
- Develop automation scripts using Nutanix REST APIs, Python SDK, PowerShell, and Terraform for infrastructure-as-code
- Create and manage VM templates, golden images, and standardized deployment catalogs for consistent provisioning
- Design disaster recovery solutions using Leap, Protection Domains, cross-cluster replication, and metro clustering
- Implement network micro-segmentation using Nutanix Flow and configure RBAC, encryption, and security hardening
- Lead L3 troubleshooting using advanced diagnostics, log analysis (CVM, Genesis), NCC health checks, and cluster service resolution
- Configure high availability, VM affinity rules, QoS policies, and optimize performance for mission-critical workloads
- Manage AHV networking with OVS bridges, VLANs, bonds, LACP and implement resource reservations and workload balance.
- Design, deploy, and maintain hybrid cloud infrastructure across Nutanix HCI, AWS, and GCP platforms
- Architect and implement multi-cloud solutions ensuring high availability, scalability, and disaster recovery
Cloud Platform Engineering
- Architect and deploy enterprise-scale, highly available multi-cloud solutions across AWS and GCP with multi-region/multi-account strategies
- Expert-level proficiency with AWS CLI, GCP CLI, SDK, boto3, and Python for advanced automation and infrastructure orchestration
- Design AWS Organizations and GCP Organization hierarchies with consolidated billing, IAM policies, and centralized governance
- Configure and manage AWS Systems Manager (SSM) including Session Manager, Run Command, State Manager, and Automation for centralized fleet operations
- Implement centralized logging using CloudWatch/CloudTrail and GCP Cloud Logging with S3/Cloud Storage aggregation
- Integrate AWS and GCP with Splunk using HEC, CloudWatch subscriptions, Pub/Sub, Dataflow, and cloud-specific add-ons for SIEM correlation
- Design and deploy advanced load balancing solutions with AWS ALB/NLB/ELB and GCP Cloud Load Balancing including SSL termination and auto-scaling
- Develop infrastructure-as-code using Terraform, CloudFormation, CDK for repeatable multi-cloud deployments and CI/CD pipelines
- Configure AWS SSO, cross-account IAM roles, GCP Workload Identity, and federated access for centralized identity management
- Design VPC architectures with AWS Transit Gateway/PrivateLink and GCP Shared VPC/VPC peering for hybrid connectivity
- Manage containerized workloads using EKS, GKE, ECS, Cloud Run with service mesh, observability, and security best practices
- Implement disaster recovery using AWS Backup, Cross-Region Replication, GCP snapshots, and multi-region failover strategies
- Lead L3 troubleshooting using CloudWatch Insights, GCP Cloud Trace, VPC Flow Logs, X-Ray, and vendor support escalation
- Perform cost optimization through Reserved Instances, Committed Use Discounts, rightsizing, and automated resource lifecycle management
System Administration
- Administer and support Windows Server and Unix/Linux environments in production and non-production settings
- Perform OS-level hardening, patch management, and security compliance across heterogeneous systems
- Automate routine administrative tasks using PowerShell, Bash, Python, or similar scripting languages
- Manage GitHub organization settings, user permissions, repository access controls, and monitor GitHub Actions workflows and repository health across multiple teams
- Configure Splunk forwarders, heavy forwarders and other integrations for data ingestion from cloud and on-premises sources
Personal and Professional Qualifications
- 8-12+ years infrastructure experience with 8+ years in Nutanix HCI and enterprise cloud AWS/GCP)
- Expert-level skills in Python, PowerShell, Bash scripting, infrastructure-as-code (Terraform/CloudFormation), and container orchestration (Kubernetes, EKS/GKE)
- Proven experience managing enterprise-scale environments, hybrid cloud migrations, disaster recovery, and L3 critical incident management
- Strong networking knowledge (TCP/IP, VLANs, routing, VPN), security hardening, and compliance frameworks (ITIL)
- Strategic thinker with exceptional analytical and troubleshooting abilities for complex multi-layer infrastructure issues
- Excellent communication skills to translate technical concepts to executives and non-technical stakeholders
- Calm under pressure during critical outages with meticulous attention to security, compliance, and configuration management
- Self-motivated continuous learner committed to staying current with evolving cloud technologies and automation opportunities
- Available for on-call rotations with strong documentation skills and customer service orientation
- Certifications (plus): Nutanix NCP/NCAP, AWS Solutions Architect Professional, AWS DevOps
- Professional, GCP Professional Cloud Architect, Terraform
Education
- Bachelor's or master's degree in computer science/IT
Juul Labs Perks & Benefits
- People. Work with talented, committed and supportive teammates
- Equity and performance bonuses. Every employee is a stakeholder in our success
- Cell phone subsidy, commuter benefits and discounts on JUUL products
- Excellent medical, dental and vision, disability, and life insurance, plus family support, wellness, legal, and employee assistance program benefits
- 401(k) plan with company matching
- Plus biannual discretionary performance bonuses
$160k - $240k
...millions of times a day - quickly, reliably, and securely. Any time you... ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our... ...operations or DevOps at a mid-to-senior level.Strong shell scripting...SeniorFull time$90k - $180k
...generic medicines. Our 115,000 colleagues serve people in more than 160 countries.JOB DESCRIPTION:About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are...SeniorRemote workShift work$262k - $364k
...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with... ...capacity and performance.Build creative engineering solutions to operations and... ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software...Senior$222k - $300.5k
...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational... .... The Fintech Platform Systems Engineering team builds and operates the AWS-based... ...negotiable.The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to...SeniorWorldwideShift work$214.1k - $309.8k
...Minimum Qualifications: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...SeniorFull timeTemporary workLocal areaFlexible hours- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms...Senior
$174k - $253k
...MINIMUM QUALIFICATIONS: Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5... ...s degree in Computer Science or Engineering. ABOUT THE JOB: Site Reliability Engineering (SRE) is what you get when you treat operations...Senior- Google is hiring Site Reliability Engineers (SRE) in Sunnyvale, CA, to ensure reliability and performance across Google’s services. The role blends software and systems engineering, allowing code fixes to improve systems while maintaining production reliability at scale...Senior
$174k - $252k
Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5 years of experience with... ...s degree in Computer Science or Engineering. About The Job Site Reliability Engineering (SRE) is what you get when you treat operations...Senior$174k - $252k
Senior Software Engineer, Site Reliability Engineering X Applicants in San Francisco: Qualified applications with arrest or conviction records will be considered for employment in accordance with the San Francisco Fair Chance Ordinance for Employers and the California...SeniorFull time$145k - $165k
A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...Senior- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...SeniorFlexible hours
- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...SeniorWork at officeLocal areaWork from homeFlexible hours
- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...SeniorFull time- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SeniorFull timeWork at office2 days per week
$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...SeniorFull time$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer NetworkingSenior$128k - $216k
...another millions of times a day - quickly, reliably, and securely. Any time you swipe... ...a difference at Fiserv.Job TitleSr. Site Reliability EngineerAbout... ...with confidence.What does a successful Senior Site Reliability Engineer do at Fiserv?As a Senior Site Reliability...SeniorFull timeWorldwide$165k - $280k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...SeniorPermanent employmentTemporary workWorldwideWeekend work$192.4k - $275.8k
...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines... ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the...SeniorFull timeTemporary workLocal areaFlexible hours- ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response...SeniorContract work
$200k - $322k
...best work.We are seeking a highly skilled Senior Staff SRE to join our dynamic team. Our... ...includes building for performance and reliability at global scale, covering automation, monitoring... ...with NVIDIA leadership, senior engineers, program managers, and product managers...SeniorFull timeRemote work- ...Senior Sre For Gpu Infrastructure You'll own the GPU infrastructure Luma's research... ...you keep training and inference clusters reliable and fast, and you help redesign them for... ...metal role for a first-principles Linux engineer. You'll be the final escalation for the...SeniorWork experience placement
- ...the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity... ...We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background...SeniorWork experience placementImmediate start
- ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software...SeniorImmediate startRemote workWorldwide
- ...Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base...SeniorWork at officeLocal areaWork from homeFlexible hours
$187.04k - $359.72k
...communication skills. Responsibilities include supporting services from design to execution, ensuring system health, and improving reliability through automation. This role offers a competitive salary range of $187,040 - $359,720 annually, alongside comprehensive benefits...Senior$166k - $244k
A leading technology company located in Sunnyvale, California, is seeking a Site Reliability Engineer responsible for building and maintaining large-scale systems. The ideal candidate should possess a degree in Computer Science and have significant experience in programming...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Mountain View, CA
- site reliability engineer sre Mountain View, CA
- senior associate architect Mountain View, CA
- senior dynamics crm developer Mountain View, CA
- senior application security Mountain View, CA
- senior account director Mountain View, CA
- sr hr business partner Mountain View, CA
- senior plumbing designer Mountain View, CA
- senior advisor Mountain View, CA
- senior ux designer Mountain View, CA

