Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Neshent Technologies

We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available database and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and infrastructure automation.

The ideal candidate will have strong production experience with PostgreSQL, AWS, Kubernetes, infrastructure as code, and distributed data systems and will work closely with SRE, application, security, and infrastructure teams.

Roles & Responsibilities
  • Design, administer, maintain, and secure PostgreSQL environments in production and cloud environments.
  • Manage PostgreSQL deployments on Kubernetes and Amazon RDS.
  • Perform database administration activities including installation, upgrades, patching, backup/recovery, monitoring, capacity planning, and performance tuning.
  • Troubleshoot production issues across application, database, operating system, storage, and network layers.
  • Design and maintain ETL pipelines and develop scripts and procedures for data migration.
  • Build and automate database and infrastructure operations using Terraform, Ansible, Chef, or Puppet.
  • Develop operational tooling and automation using Python, Bash, Go, Ruby, or Perl.
  • Monitor database health, performance, availability, and capacity and implement improvements.
  • Participate in on-call rotations, incident response, alerting, and post-incident reviews.
  • Improve reliability practices through SLOs, disaster recovery, backup validation, and failover testing.
  • Build database platform tooling and self-service workflows that enable application teams to use data services safely and efficiently.
  • Operate and support distributed data systems such as Kafka/MSK, ClickHouse, Redis, MySQL, Cassandra, or Elasticsearch.
  • Collaborate with SRE, application engineering, security, and infrastructure teams to drive reliable architectural changes.
  • Create technical documentation and design proposals and communicate solutions effectively to engineering stakeholders.
Required Qualifications
  • 5+ years of experience designing, operating, and troubleshooting PostgreSQL in production environments.
  • 5+ years managing production databases or distributed data systems across application, database, OS, storage, and network layers.
  • Hands-on experience with AWS, Amazon RDS, Kubernetes, Terraform, service discovery, and secrets management.
  • 3+ years of Linux systems engineering experience, including performance tuning, memory management, I/O tuning, security, configuration, and networking.
  • Experience automating infrastructure or database operations using Terraform, Ansible, Chef, or Puppet.
  • 2+ years of scripting or programming experience with Python, Bash, Go, Ruby, or Perl.
  • Experience working with at least one non-PostgreSQL data platform such as Kafka/MSK, ClickHouse, Redis, MySQL, Cassandra, or Elasticsearch.
  • Strong understanding of PostgreSQL internals, including replication, concurrency, transactions, indexing, maintenance, backup/recovery, and query performance.
  • Strong troubleshooting, problem-solving, communication, and documentation skills.
Preferred Skills
  • Experience building database platform tooling, self-service workflows, and standardized operational processes.
  • Experience operating Kafka/MSK, ClickHouse, or Redis at scale, including clustering, replication, partitioning/sharding, retention, capacity planning, and workload tuning.
  • Experience defining and improving reliability practices for production data systems.
  • Strong knowledge of SLOs, disaster recovery, backup validation, failover, and resilience testing.
  • Experience with cloud-native database architectures and distributed systems.
  • Ability to create technical design documents and communicate complex database and reliability concepts effectively.
  • Ability to work effectively in a fast-paced, collaborative engineering environment.
Vacancy posted 10 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Los Gatos, CA vacancy
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    4 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    22 hours ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    1 day ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $160k - $240k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our...  ...operations or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    1 day ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    4 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $152k - $241.5k

     ...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and...  ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-...  ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $262k - $364k

     ...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with...  ...capacity and performance.Build creative engineering solutions to operations and...  ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software... 
    Senior

    Google

    Mountain View, CA
    1 day ago
  • $222k - $300.5k

     ...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational...  .... The Fintech Platform Systems Engineering team builds and operates the AWS-based...  ...negotiable.The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to... 
    Senior
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    4 days ago
  • $192.4k - $275.8k

     ...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines...  ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    1 day ago
  •  ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation... 
    Senior

    Oracle

    Santa Clara, CA
    1 day ago
  • $200k - $322k

     ...best work.We are seeking a highly skilled Senior Staff SRE to join our dynamic team. Our...  ...includes building for performance and reliability at global scale, covering automation, monitoring...  ...with NVIDIA leadership, senior engineers, program managers, and product managers... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $262k - $364k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...or Engineering, or a related field.Site Reliability Engineering (SRE) combines software and...  ...Software Engineer chose to join SRE.As the Senior Engineering Manager for Collaboration... 
    Senior

    Google

    Sunnyvale, CA
    3 days ago
  • $187.04k - $359.72k

     ...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum...  ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas....  ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company... 
    Senior
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    1 day ago
  • $207.4k - $259.2k

     ...built specifically for aviation. We’re seeking exceptional engineers, operators and builders to join us on our mission to build the...  ...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role... 
    Senior
    Permanent employment
    Local area
    Visa sponsorship
    Night shift

    Archer Aviation

    San Jose, CA
    4 days ago
  • $125k - $160k

     ...Docker and Kubernetes.Collaborate with Product Management and engineering peers from concept through delivery.Maintain high engineering...  ....Continuously improve system performance, scalability, and reliability.Qualifications:7+ years of professional experience with Java.... 
    Senior
    Full time
    Remote work
    Flexible hours

    Centric Software

    Campbell, CA
    1 day ago
  • $100k - $170k

     ...enhance our overall quality of life. We are seeking a  DevOps Engineer who is eager to have an immediate impact in establishing and...  ...manage containerized applications on Amazon EKS with focus on reliability and performance Build and maintain Infrastructure as Code... 
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    17 days ago
  •  ...part of something exceptional.Position: Senior DevOps EngineerLocation: Campbell, CA /...  ...OverviewWe are seeking a Senior DevOps Engineer to partner closely with development and...  ...cause analysis and continuously improve reliability, scalability, security, performance, and... 
    Senior
    Full time
    Remote work

    Centric Software

    Campbell, CA
    2 days ago
  •  ...Software EngineerWe are looking for a few exceptional software engineers to work on our cloud based B2B e-commerce, renewals and subscriptions platform.As a member of the engineering team, you will work with product management and other team members to design and implement... 
    Senior

    Rainmaker Systems

    Campbell, CA
    1 day ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems...  ..., analytics, and automated anomaly detection.Partner with senior leaders and engineers across Cloud, Platform, Security,... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $130k - $200k

     ...Senior Software EngineerReady to make connectivity from space universally accessible, secure and actionable? Then you've come to the...  ...intelligence.What is the role?E-Space is looking for a Senior Software Engineer to join our Ground Software team. You will collaborate with... 
    Senior
    Immediate start

    eSpace

    Saratoga, CA
    22 hours ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    1 day ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    3 days ago
  •  ...The RoleThis hybrid role combines the hands-on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE).The ideal candidate has a strong technical foundation, thrives in a... 
    Full time
    Local area

    F5 Networks

    San Jose, CA
    2 days ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Full time

    Google

    Sunnyvale, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!