Senior Site Reliability Engineer
Neshent Technologies
We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available database and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and infrastructure automation.
The ideal candidate will have strong production experience with PostgreSQL, AWS, Kubernetes, infrastructure as code, and distributed data systems and will work closely with SRE, application, security, and infrastructure teams.
Roles & Responsibilities
- Design, administer, maintain, and secure PostgreSQL environments in production and cloud environments.
- Manage PostgreSQL deployments on Kubernetes and Amazon RDS.
- Perform database administration activities including installation, upgrades, patching, backup/recovery, monitoring, capacity planning, and performance tuning.
- Troubleshoot production issues across application, database, operating system, storage, and network layers.
- Design and maintain ETL pipelines and develop scripts and procedures for data migration.
- Build and automate database and infrastructure operations using Terraform, Ansible, Chef, or Puppet.
- Develop operational tooling and automation using Python, Bash, Go, Ruby, or Perl.
- Monitor database health, performance, availability, and capacity and implement improvements.
- Participate in on-call rotations, incident response, alerting, and post-incident reviews.
- Improve reliability practices through SLOs, disaster recovery, backup validation, and failover testing.
- Build database platform tooling and self-service workflows that enable application teams to use data services safely and efficiently.
- Operate and support distributed data systems such as Kafka/MSK, ClickHouse, Redis, MySQL, Cassandra, or Elasticsearch.
- Collaborate with SRE, application engineering, security, and infrastructure teams to drive reliable architectural changes.
- Create technical documentation and design proposals and communicate solutions effectively to engineering stakeholders.
Required Qualifications
- 5+ years of experience designing, operating, and troubleshooting PostgreSQL in production environments.
- 5+ years managing production databases or distributed data systems across application, database, OS, storage, and network layers.
- Hands-on experience with AWS, Amazon RDS, Kubernetes, Terraform, service discovery, and secrets management.
- 3+ years of Linux systems engineering experience, including performance tuning, memory management, I/O tuning, security, configuration, and networking.
- Experience automating infrastructure or database operations using Terraform, Ansible, Chef, or Puppet.
- 2+ years of scripting or programming experience with Python, Bash, Go, Ruby, or Perl.
- Experience working with at least one non-PostgreSQL data platform such as Kafka/MSK, ClickHouse, Redis, MySQL, Cassandra, or Elasticsearch.
- Strong understanding of PostgreSQL internals, including replication, concurrency, transactions, indexing, maintenance, backup/recovery, and query performance.
- Strong troubleshooting, problem-solving, communication, and documentation skills.
Preferred Skills
- Experience building database platform tooling, self-service workflows, and standardized operational processes.
- Experience operating Kafka/MSK, ClickHouse, or Redis at scale, including clustering, replication, partitioning/sharding, retention, capacity planning, and workload tuning.
- Experience defining and improving reliability practices for production data systems.
- Strong knowledge of SLOs, disaster recovery, backup validation, failover, and resilience testing.
- Experience with cloud-native database architectures and distributed systems.
- Ability to create technical design documents and communicate complex database and reliability concepts effectively.
- Ability to work effectively in a fast-paced, collaborative engineering environment.
- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...SeniorFlexible hours
- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...SeniorWork at officeLocal areaWork from homeFlexible hours
- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SeniorFull timeWork at office2 days per week
$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...SeniorFull time$160k - $240k
...millions of times a day - quickly, reliably, and securely. Any time you... ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our... ...operations or DevOps at a mid-to-senior level.Strong shell scripting...SeniorFull time$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer NetworkingSenior$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...SeniorFull time$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours$152k - $241.5k
...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and... ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-... ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through...SeniorFull time$262k - $364k
...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with... ...capacity and performance.Build creative engineering solutions to operations and... ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software...Senior$222k - $300.5k
...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational... .... The Fintech Platform Systems Engineering team builds and operates the AWS-based... ...negotiable.The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to...SeniorWorldwideShift work$192.4k - $275.8k
...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines... ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the...SeniorFull timeTemporary workLocal areaFlexible hours$145k - $165k
...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...Senior- ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation...Senior
$200k - $322k
...best work.We are seeking a highly skilled Senior Staff SRE to join our dynamic team. Our... ...includes building for performance and reliability at global scale, covering automation, monitoring... ...with NVIDIA leadership, senior engineers, program managers, and product managers...SeniorFull timeRemote work- ...Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base...SeniorWork at officeLocal areaWork from homeFlexible hours
$262k - $364k
Lead a team of Software/Systems Engineers on projects for users and be directly responsible... ...or Engineering, or a related field.Site Reliability Engineering (SRE) combines software and... ...Software Engineer chose to join SRE.As the Senior Engineering Manager for Collaboration...Senior$187.04k - $359.72k
...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum... ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas.... ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company...SeniorTemporary workLocal areaOverseasShift work$207.4k - $259.2k
...built specifically for aviation. We’re seeking exceptional engineers, operators and builders to join us on our mission to build the... ...are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role...SeniorPermanent employmentLocal areaVisa sponsorshipNight shift$125k - $160k
...Docker and Kubernetes.Collaborate with Product Management and engineering peers from concept through delivery.Maintain high engineering... ....Continuously improve system performance, scalability, and reliability.Qualifications:7+ years of professional experience with Java....SeniorFull timeRemote workFlexible hours$100k - $170k
...enhance our overall quality of life. We are seeking a DevOps Engineer who is eager to have an immediate impact in establishing and... ...manage containerized applications on Amazon EKS with focus on reliability and performance Build and maintain Infrastructure as Code...Full timeWork at officeImmediate startVisa sponsorshipNight shift- ...part of something exceptional.Position: Senior DevOps EngineerLocation: Campbell, CA /... ...OverviewWe are seeking a Senior DevOps Engineer to partner closely with development and... ...cause analysis and continuously improve reliability, scalability, security, performance, and...SeniorFull timeRemote work
- ...Software EngineerWe are looking for a few exceptional software engineers to work on our cloud based B2B e-commerce, renewals and subscriptions platform.As a member of the engineering team, you will work with product management and other team members to design and implement...Senior
$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems... ..., analytics, and automated anomaly detection.Partner with senior leaders and engineers across Cloud, Platform, Security,...Full time$130k - $200k
...Senior Software EngineerReady to make connectivity from space universally accessible, secure and actionable? Then you've come to the... ...intelligence.What is the role?E-Space is looking for a Senior Software Engineer to join our Ground Software team. You will collaborate with...SeniorImmediate start$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...Night shift$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...Full timeWorldwide- ...The RoleThis hybrid role combines the hands-on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE).The ideal candidate has a strong technical foundation, thrives in a...Full timeLocal area
$255.7k - $300k
...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system... ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- senior implementation engineer Los Gatos, CA
- senior Los Gatos, CA
- senior manager legal Los Gatos, CA
- senior software engineer remote Los Gatos, CA
- senior tech Los Gatos, CA
- senior cloud network engineer Los Gatos, CA
- senior designer remote Los Gatos, CA
- senior medical science liaison Los Gatos, CA
- senior level Los Gatos, CA
- senior performance tester Los Gatos, CA


