Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

LeanData

LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is designed for a builder - someone who wants to move beyond maintenance and into the realm of architectural transformation.You will have the autonomy to evaluate our existing AWS footprint and lead the charge in modernizing our environment. Your mission is to take a high-velocity system and implement the best practices, guardrails, and automated architectures that will support our next 10x of scale. You will be the primary authority on reliability, performance, and infrastructure security.Please note: This is a hybrid role based in our Santa Clara, CA office, with an in-office schedule of two days per week – Monday and Wednesday.Key ResponsibilitiesArchitectural Modernization: Lead the design and implementation of a scalable, "Cloud-First" AWS architecture. You will drive the transition toward fully automated, state-of-the-art Infrastructure as Code (Terraform).High Availability & Resilience: Design and implement robust Disaster Recovery (DR) and Business Continuity plans, moving our services toward a zero-downtime deployment model.Performance & Capacity Engineering: Own the strategy for capacity planning and autoscaling. You will optimize our compute resources (EC2, Lambda) to handle bursty traffic patterns with precision and cost-efficiency.Advanced Observability: Define our monitoring and alerting philosophy using New Relic for deep APM and system insights. Partner this with IncidentIO to ensure we catch and resolve issues before they impact customers.Streamlined CI/CD: Partner with feature teams to refine Change Management and CI/CD pipelines, ensuring code moves from "commit" to "production" safely and predictably.Cloud Security: Harden our network architecture and application security posture, including WAF management and secure service-to-service communication.The Tech StackCloud Infrastructure: AWS (EC2, Lambda, SQS, SNS, ALB, API Gateway, S3, WAF).Observability & Incident Response:New Relic (APM/Infrastructure), IncidentIO.Automation & Tools: Terraform, Redis/Elasticache, Shell Scripting, NPM/PM2.Application Ecosystem: NodeJS, Python, C#, Angular, Apex.Integration: Salesforce Managed Packages, MSFT Dynamics365.Who You AreExperienced Architect: 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record of managing complex AWS environments.Proven Incident Commander: You demonstrate calm, decisive leadership during high-pressure outages. You have extensive experience running blameless postmortems and, crucially, driving the remediation work needed to prevent recurrence.Observability Pro: You have deep experience configuring New Relic (or similar platforms) to create meaningful dashboards, SLIs, and SLOs.Automation Advocate: You believe that manual intervention is a bug. You have deep experience with Terraform and a "Code-First" approach to infrastructure.Strategic Problem Solver: You can look at a complex, "needs-based" architecture and formulate a clear, prioritized roadmap to move it toward industry best practices.Collaborative Leader: You enjoy working with feature engineers to help them build "reliability-by-design" into their services.Education: A Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent professional experience).Why work at LeanData:LeanData covers employee insurance premiums up to 90%Stock options in LeanData for all full-time employeesFlexible PTO401K planCompensation Range: $140K - $180KLocationSanta ClaraEmployment TypeFull timeDepartmentEngineeringCompensationTier 1Estimated Base Salary $140K – $180KWe value providing opportunities for professional growth, and hence, we typically refrain from offering at the upper limit of the salary range. This approach allows for future development within the role, ensuring that compensation can be adjusted to reflect an individual's advancing skills and contributions to the team while maintaining internal fairness.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Santa Clara, CA vacancy
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    20 hours ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 days ago
  • $174k - $253k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Sunnyvale, CA
    20 hours ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    20 hours ago
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We... 
    Senior
    Remote work

    Abbott

    Sunnyvale, CA
    5 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $159.2k - $301.6k

     ...for running Graphs on the cloud. In this reliability-focused role, you will own the...  ...infrastructure. You'll partner with the backend engineers building these APIs to make sure the system...  ...Science. 5-10 years of experience in site reliability engineering, infrastructure,... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    4 days ago
  • $152k - $241.5k

     ...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and...  ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-...  ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $262k - $365k

     ...automation, and evolve systems by pushing for changes that improve reliability and velocity.Practice sustainable incident response and...  ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems... 
    Senior

    Google

    San Jose, CA
    20 hours ago
  • $210.6k - $305.1k

     ...Minimum Qualifications:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure...  ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    5 days ago
  •  ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 
    Senior

    TechDigital Group

    Santa Clara, CA
    3 days ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    3 days ago
  • $262k - $365k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and...  ...Software Engineer chose to join SRE.As the Senior Engineering Manager for Collaboration... 
    Senior

    Google

    Sunnyvale, CA
    5 days ago
  •  ...Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $187.04k - $359.72k

     ...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum...  ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas....  ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company... 
    Senior
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    3 days ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    Palo Alto, CA
    3 days ago
  • $207.4k - $259.2k

     ...differences, and supports and celebrates all of our team members.We are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you will be responsible for the reliability, scalability,... 
    Senior
    Permanent employment
    Local area

    Archer Aviation

    San Jose, CA
    2 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms... 
    Senior

    JP Morgan Chase

    Palo Alto, CA
    3 days ago
  •  ...professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of...  ...professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the... 
    Senior

    J.P. Morgan

    Palo Alto, CA
    5 days ago
  •  ...Job Description Job Description About the Role Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform... 
    Senior

    Kody

    Sunnyvale, CA
    21 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    5 days ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    3 days ago
  • $147k - $211k

     ...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style...  ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations... 

    Google

    Sunnyvale, CA
    20 hours ago
  • $184k

     ...best work. Come join the team and see how you can make a lasting impact on the world. We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely positioned to drive... 
    Senior
    Full time

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $207k - $301k

     ...implementation of solutions to enhance the reliability of systems that support F1.Scale systems...  ...for multiple teams.Engage in software engineering on services written in Java, C++, and Go...  ...related technical field.Experience in a Site Reliability Engineering role.Experience... 

    Google

    San Jose, CA
    4 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $122.5k - $175k

     ...we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief Architect... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    4 days ago
  • $184k - $287.5k

    NVIDIA is searching for a creative and highly motivated engineer with expertise in systems software to join the GPU Software team. You will design key aspects of our production GPU kernel drivers and embedded SW that impacts our products both in the datacenter and in gaming... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    Our Autonomous Vehicles Platform team is searching for engineers to develop and bring NVIDIA's automotive platform out to the world. You will participate in a focused effort to develop and productize ground-breaking solutions that will revolutionize the world of transportation... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    21 hours ago
  • $152k - $241.5k

    The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous vehicle platform to new markets! This role involves developing and productizing innovative solutions that will transform transportation and the field of... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!