Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

LeanData

Senior Site Reliability Engineer

LeanData helps the world's fastest-growing companies automate, simplify, and accelerate revenue.

We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is designed for a builder - someone who wants to move beyond maintenance and into the realm of architectural transformation.

You will have the autonomy to evaluate our existing AWS footprint and lead the charge in modernizing our environment. Your mission is to take a high-velocity system and implement the best practices, guardrails, and automated architectures that will support our next 10x of scale. You will be the primary authority on reliability, performance, and infrastructure security.

This is a hybrid role based in our Santa Clara, CA office, with an in-office schedule of two days per week – Monday and Wednesday.

Key Responsibilities
  • Architectural Modernization: Lead the design and implementation of a scalable, "Cloud-First" AWS architecture. You will drive the transition toward fully automated, state-of-the-art Infrastructure as Code (Terraform).

  • High Availability & Resilience: Design and implement robust Disaster Recovery (DR) and Business Continuity plans, moving our services toward a zero-downtime deployment model.

  • Performance & Capacity Engineering: Own the strategy for capacity planning and autoscaling. You will optimize our compute resources (EC2, Lambda) to handle bursty traffic patterns with precision and cost-efficiency.

  • Advanced Observability: Define our monitoring and alerting philosophy using New Relic for deep APM and system insights. Partner this with IncidentIO to ensure we catch and resolve issues before they impact customers.

  • Streamlined CI/CD: Partner with feature teams to refine Change Management and CI/CD pipelines, ensuring code moves from "commit" to "production" safely and predictably.

  • Cloud Security: Harden our network architecture and application security posture, including WAF management and secure service-to-service communication.

The Tech Stack
  • Cloud Infrastructure: AWS (EC2, Lambda, SQS, SNS, ALB, API Gateway, S3, WAF).

  • Observability & Incident Response: New Relic (APM/Infrastructure), IncidentIO.

  • Automation & Tools: Terraform, Redis/Elasticache, Shell Scripting, NPM/PM2.

  • Application Ecosystem: NodeJS, Python, C#, Angular, Apex.

  • Integration: Salesforce Managed Packages, MSFT Dynamics365.

Who You Are
  • Experienced Architect: 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record of managing complex AWS environments.

  • Proven Incident Commander: You demonstrate calm, decisive leadership during high-pressure outages. You have extensive experience running blameless postmortems and, crucially, driving the remediation work needed to prevent recurrence.

  • Observability Pro: You have deep experience configuring New Relic (or similar platforms) to create meaningful dashboards, SLIs, and SLOs.

  • Automation Advocate: You believe that manual intervention is a bug. You have deep experience with Terraform and a "Code-First" approach to infrastructure.

  • Strategic Problem Solver: You can look at a complex, "needs-based" architecture and formulate a clear, prioritized roadmap to move it toward industry best practices.

  • Collaborative Leader: You enjoy working with feature engineers to help them build "reliability-by-design" into their services.

  • Education: A Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent professional experience).

Why work at LeanData:

  • LeanData covers employee insurance premiums up to 90%

  • Stock options in LeanData for all full-time employees

  • Flexible PTO

  • 401K plan

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Santa Clara, CA vacancy
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response... 
    Senior
    Contract work

    VDart

    Santa Clara, CA
    2 days ago
  • $160k - $240k

     ...consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay...  ...scale, come make a difference at Fiserv. Job Title Senior Site Reliability Engineer What does a successful Site Reliability Engineer do at... 
    Senior

    Fiserv

    Sunnyvale, CA
    4 days ago
  • $132.6k - $214.5k

     ...you will collaborate closely with our engineering teams to develop innovative solutions that...  ...' performance and health. As a Senior Staff SRE with the Cortex Observability...  ...operability of the product and ensure the reliability and availability of our services. Qualifications... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $146.7k

     ...Senior Lead Site Reliability Engineer Immigration sponsorship is not available for this position. What you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid systems across the globe. You will be responsible... 
    Senior
    Casual work
    Work at office
    Remote work
    Worldwide
    Shift work

    Zoom Video Communications

    San Jose, CA
    14 hours ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Senior
    Part time
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    17 days ago
  •  ...the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity...  ...We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background... 
    Senior
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    23 hours ago
  • $128k - $216k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe your...  ...make a difference at Fiserv. Sr. Site Reliability Engineer About Clover Clover is a pioneer...  ...confidence. What Does A Successful Senior Site Reliability Engineer Do At Fiserv... 
    Senior
    Worldwide

    BentoBox

    Sunnyvale, CA
    4 days ago
  •  ...Senior Lead Site Reliability Engineer Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan... 
    Senior

    Chase

    Palo Alto, CA
    3 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    6 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    2 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    2 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $152k - $241.5k

     ...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and...  ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-...  ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $192.4k - $275.8k

     ...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines...  ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    4 days ago
  • $200k - $322k

     ...best work.We are seeking a highly skilled Senior Staff SRE to join our dynamic team. Our...  ...includes building for performance and reliability at global scale, covering automation, monitoring...  ...with NVIDIA leadership, senior engineers, program managers, and product managers... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $149.8k - $224.6k

     ...This hybrid role combines the hands-on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE). The ideal candidate has a strong technical foundation, thrives... 
    Local area

    F5

    San Jose, CA
    2 days ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...things have always been done. Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights... 
    Night shift

    Forward

    Santa Clara, CA
    3 days ago
  •  ...Site Reliability Engineer (SRE) Share Contractual Sunnyvale, CA PDT - 8450 8-10 Overview: *Must have Apple experience* • At least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Advanced experience with programming languages (Python... 

    Purple Drive

    Sunnyvale, CA
    3 days ago
  •  ...Technical Support Engineer/Site Reliability Engineer At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are... 

    F5

    San Jose, CA
    1 day ago
  •  ...Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job Description Deploy, support and monitor new and existing services, platforms, and application stacks. Use scale testing to measure, tune... 

    Zortech Solutions

    Cupertino, CA
    4 days ago
  •  ...Senior Site Reliability Engineer Location: Remote Duration: 12 month contract to start IV Process: 1-3 Round IV process International Tech Top Skills: Java Python NodeJS -DevOps Engineer should work here too Main Responsibilities... 
    Contract work
    Local area
    Remote work

    My3Tech Inc

    Sunnyvale, CA
    2 days ago
  •  ...Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots. FII provides customers with intelligent... 
    Permanent employment
    Full time
    Work at office
    Local area

    Foxconn Industrial Internet

    San Jose, CA
    23 hours ago
  •  ...AWS Infra SRE/DevOps Engineer AWS Infra SRE/DevOps engineer with proven work experience ensuring reliability, availability and performance of cloud infra and platform. Specialist on Cisco Cloud run-on for infrastructure management, who can install, run, and maintain... 
    Work experience placement

    The Dignify Solutions, LLC

    San Jose, CA
    2 days ago
  •  ...Eyes on glass. Hands on the pipeline. Real ownership from day one. This isn't a watch-and-wait monitoring seat. Our client needs engineers who can read a Kibana query at 3am, know the difference between a blip and a breach, and act on it, on a FedRAMP-authorised cloud... 
    Hourly pay
    For contractors
    Shift work
    Night shift
    Weekend work

    C-Serv

    San Jose, CA
    23 hours ago
  •  ...Qualifications: 8+ years of software engineering experience, or equivalent demonstrated through...  ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform...  ...vendor resources. Willingness to work on-site at stated location in the job opening.... 
    For contractors
    Work experience placement

    Cedent Life Talent

    San Jose, CA
    3 days ago
  • $65 - $85 per hour

     ...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    1 day ago
  • $170k - $200k

     ...Site Reliability Engineer We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high... 
    Full time
    Worldwide

    Edelman

    Sunnyvale, CA
    2 days ago
  • $262k - $364k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...or Engineering, or a related field.Site Reliability Engineering (SRE) combines software and...  ...Software Engineer chose to join SRE.As the Senior Engineering Manager for Collaboration... 
    Senior

    Google

    Sunnyvale, CA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!