Senior Site Reliability Engineer
Nexla
Nexla is the leading Integration platform, built with AI, for AI. Nexla takes a metadata driven approach to converge diverse integrations across Data, Documents, Agents, Applications, and APIs into a single design pattern. We accelerate the development of solutions for GenAI, Analytics, and Inter-company data. Nexla makes data users and developers up to 10x more productive by delivering a true blend of no-code, low-code, and pro-code interfaces. Leading companies including DoorDash, LinkedIn, Johnson & Johnson, and LiveRamp trust Nexla for mission-critical data. Named in the 2022, 2023, and 2024 Gartner Magic Quadrant™ for Data Integration Tools and top-rated by customers on Gartner Peer Insights, headquartered in San Mateo, California. At Nexla, our culture is built around our core values: Have Empathy , Be Curious , Be Intellectually Honest , Achieve Excellence , and Remember to Relax . We put our customers at the heart of everything we do, foster a data-driven mindset, take ownership of our work, and believe in the power of teamwork to achieve ambitious goals. Role You will own the reliability of the distributed data systems at the heart of Nexla - the streaming runtime and processing engines that move hundreds of billions of rows per day for top-tier enterprises. This is an SRE role for our big data stack: Kafka, Spark, Flink, Ray, Redis, and data warehouses, all running on Kubernetes. This is not a cloud-provisioning role. We are looking for someone who has lived inside stateful, high-throughput systems in production who has chased down a broker outage, a checkpoint stall, a crashlooping cache, and a sink that silently stopped writing, and who fixes the architecture rather than the symptom. If keeping a large, busy data platform alive and fast is the kind of problem you find satisfying, you will have a lot of fun working with us. This is a unique opportunity to shape the foundation of a product that is defining the next wave of intelligent, context-aware data movement. Responsibilities Streaming & Data Plane Reliability: Own the health of our Kafka-based runtime (managed via Strimzi on Kubernetes) – broker health, topic lifecycle and count management, partition and throughput tuning, certificate/secret rotation, and version upgrades – at a scale of hundreds of thousands of topics and hundreds of billions of rows per day. Distributed Processing Engines: Operate and tune distributed system workloads in production in collaboration with backend teams, resource allocation, autoscaling, checkpointing, backpressure, and failure recovery for both batch and streaming jobs. Stateful Services: Run Redis clusters and other stateful systems reliably – failover, persistence, liveness/readiness tuning, and capacity planning under heavy and bursty load. Kubernetes & Operators: Take end-to-end ownership of Amazon EKS, Google GKE and the operators (Strimzi and others) running our stateful data workloads – cluster lifecycle, scaling, version upgrades, and resource governance. Observability: Build deep, data-aware monitoring – consumer lag, throughput, partition skew, job latency, error rates – not just host and CPU metrics. Make the data plane's behavior legible before it breaks. Incident Management: Lead root-cause analysis for distributed-systems failures (broker outages, crashloops, sink decommissions, control-plane race conditions) and drive durable fixes. Mitigate fast, but design out the recurrence. Infrastructure as Code & Automation: Provision and manage cloud infrastructure with Terraform; build operational runbooks and automation, including for air-gapped / private enterprise installs (pre-staged images, operator-facing procedures). Collaboration: Partner with platform, runtime, and connector engineering – and with SREs and support – to ship and scale new data-movement features reliably in a large-scale Linux environment. Qualifications Experience: 8+ years in infrastructure, SRE, or DevOps, with significant time spent operating production distributed data systems (not just application/cloud infra). Kafka: Deep, hands-on operational experience running Kafka at scale in production – ideally on Kubernetes via Strimzi – including upgrades, topic/partition management, performance tuning, and TLS/secret rotation. Distributed Processing (Strong Plus): Production experience operating one or more of Spark, Flink, or Ray – resource tuning, checkpointing, failure recovery. Stateful Systems (Must Have): Production experience with Redis (clustering, persistence, failover) and a solid understanding of operating stateful workloads on Kubernetes (StatefulSets, PVCs, probes, operators). Data Warehouses: Familiarity operating against Snowflake, BigQuery, or similar, and an understanding of JDBC connectivity and sink reliability. Kubernetes & EKS: Strong hands-on EKS – cluster creation, scaling, version upgrades, and operator management. Infrastructure as Code: Advanced proficiency with Terraform. Programming: Proficiency in Python (or similar) for automation and tooling. Comfort reading and debugging JVM-based systems is a strong plus. Reliability Mindset: Demonstrated ownership of incident management, RCA, capacity planning, and performance tuning for high-throughput systems. CI/CD: Solid understanding of CI/CD methodology (Jenkins, GitHub Actions, or GitLab CI) for containerized and non-containerized apps. Supporting, not the core of the role. Nice to Have: Configuration management (Ansible preferred); broader AWS services (IAM, VPC, EC2, S3, Lambda); AWS CloudFormation. Soft Skills: Excellent communication and organizational skills; ability to coordinate effectively within a team and with customers. Why This Might Be Worth It You own the hard part. The stateful, distributed systems that move billions of rows are the platform's most demanding reliability problems – and they'd be yours. Impact at scale from day one. Your work keeps mission-critical data flowing for companies like DoorDash and LinkedIn. The AI wave is real for us. We're not bolting AI onto a legacy product. Intelligent connectors, context-aware data movement, and agentic workflows are the core of what we're building next – on top of the runtime you'd run. Small team, big problems. Direct access to the CTO, real influence over product direction, and the autonomy to make significant technical bets. Recognized platform, startup energy. Enterprise validation with the speed and ownership of an early-stage company. #J-18808-Ljbffr Nexla
$130k - $200k
IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SeniorFull timeWork at officeImmediate start- ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software...SeniorImmediate startRemote workWorldwide
$137.77k - $194.59k
...distributed team of roughly 80 scientists and engineers building and operating Rubin's petascale... ...Your role: \n You will own the reliability and robustness of Rubin Observatory's... ...nature of this position, SLAC is open to on-site, hybrid, and remote work options. \n \...SeniorRemote workFlexible hoursNight shift- ...strong customer and partner networks. About the Role This role leads reliability strategy and architectural improvements across infrastructure, GPU systems, observability, ML Ops and IT Ops. Mentor engineers, manage high‑severity incidents, and drive SLO governance. You...SeniorFull time
- ...Site Reliability Engineer (SRE) Xona is the navigational intelligence company bringing real-time, centimeter-level certainty to any device, anywhere on Earth. With Pulsar – the world's most advanced PNT satellite infrastructure in Low Earth Orbit – Xona will offer a...SuggestedPermanent employment
- ...Site Reliability Engineer As a Site Reliability Engineer, you have a mindset to maximize system availability through both proactive and reactive means: you build robust technical support and automation to eliminate or minimize incidents, as well as investigate and resolve...Live in
- Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development and operation of our autonomous vehicles. In this role, you will own the full lifecycle of our services—from designing fault...
$150k - $170k
A tech education startup in California seeks a Senior Software Engineer to build innovative user features and work with cutting-edge AI technologies. This role focuses on enhancing the online learning experience and offers a competitive salary of $150,000 to $170,000,...SeniorRemote work- ...Senior Platform Software Engineer Full Time opportunity, Office located in Belmont, CA (Hybrid) As a Senior Platform Software Engineer, you will... ...Python-based backend systems, ensuring high performance, reliability, and cost-efficiency. Join us to shape the future of real-...SeniorFull timeWork at office
- ...Design to make sure what gets built actually solves the right problems — and stay close to customer outcomes. Mentorship: Help junior engineers grow by sharing context, giving direct feedback, and modeling strong engineering habits. Engineering initiatives: Contribute to...Senior
$277.17k - $343.34k
...We're looking for a Principal Platform Engineer who treats platform as a product: someone... ...and oversee development of high scale and reliable infrastructure systems, with clear SLOs... ...and practices. Mentor junior and senior engineers, lead design reviews, and drive...SeniorFull timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday- ...transportation and ground-up build autonomous robotaxis that are safe, reliable, clean, and enjoyable for everyone. We are still in the early... ...Science, Collision Avoidance, etc. and our Advanced Hardware Engineering group and have the opportunity to significantly push the...SeniorTemporary workRelocation package
- ...About the Team We are a team of engineers, scientists, and domain experts dedicated to making housing and building development more... ...pipelines to ensure high velocity without sacrificing reliability. Data Orchestration: Contribute to our data orchestration...Senior
$242.1k - $293.8k
...solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Engine DataModel team, you will own and innovate on the foundational components that form the backbone of the...SeniorFull timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday3 days per week$160k - $180k
...Senior Software Engineer – Platform Team We're looking for a self‑motivated Senior Software Engineer to join the Platform team that powers all... ...production rollout and monitoring Improve platform reliability, performance, and scalability — across services, data layer...SeniorFull timeWork at office2 days per week- About the Role You will work directly with our leadership team to build new features, enhance existing ones, and ensure the overall stability and performance of our software. This is a unique opportunity to join a growing startup at an early stage and have a significant...Senior
- ...San Mateo is building a secure, scalable cloud platform and embracing AI to drive innovation for leading insurers. As a Software Engineer III, you will develop the Guidewire Cloud Platform, contribute to AI adoption, and craft high-impact cloud solutions in a collaborative...Senior
$150k - $170k
...Base pay range $150,000.00/yr - $170,000.00/yr Direct message the job poster from Hewitt Banks Position Senior Software Engineer – AI + EdTech at Hewitt Banks We’re a cross-functional founding team with 20+ years of experience across education, engineering...SeniorFull timeRemote workFlexible hours- ...Senior Software Engineer @ Anatomy Financial Anatomy is on a mission to automate financial operations for healthcare and enable providers to focus on quality patient care. Financial operations in healthcare are unique. Anatomy uses the latest advancements in AI and...SeniorLocal area
- ...Senior Software Engineer Remote (United States continental 48) / San Mateo, California, United States JetInsight's best-in-class quoting and fleet management software helps aircraft charter operators run and grow safer, more efficient, and more profitable businesses...SeniorRemote work
$130k - $200k
...Senior Software Engineer IXL Learning, developer of personalized learning products used by millions of people globally, is seeking senior software engineers who have a passion for technology and education to help us add new features to our extremely successful educational...SeniorFull timeWork at office$243.29k - $295.25k
...technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Sr. Studio Software Engineer, you will be a key contributor to the evolution of Roblox Studio, the primary IDE for making massive multiplayer online games on the...SeniorFull timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$193.3k - $289.9k
...SIMD optimizations on real‑time systems. Contributions to open‑source projects in multimedia, systems programming, or performance engineering. Estimated Base Pay Range: $193,300 – $289,900 USD This role is eligible for SIE’s top‑tier benefits package, which includes...Senior- Dormont Manufacturing Co is hiring a Senior Software Engineer for its Cloud Engineering team in San Mateo, California. You will lead projects to build scalable and resilient infrastructures, automate complex tasks, and mentor junior colleagues. The ideal candidate has...Senior
$193.3k - $289.9k
...and nurture the experiences under the PlayStation brand, a name synonymous with entertainment excellence and creativity. Senior Software Engineer San Mateo, CA Role Overview Sony Interactive Entertainment LLC seeks a Senior Software Engineer in San Mateo, CA to lead rapid...SeniorRemote workWork from home$243.29k - $295.25k
...hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on the Cache team (part of the Infra Storage org), you will innovate and operate large-scale, in-house distributed systems to...SeniorFull timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$190k - $260k
...Senior Software Engineer Alluxio powers the data layer for modern AI and analytics. Proven in production at eight of the top ten internet... ...logic, and metadata scalability to increase performance and reliability. Data path optimization - refine I/O pipelines for S3/...SeniorFull time- ...Role Summary We are looking for a highly skilled Sr. Fullstack Engineer to join our dynamic engineering team. In this role, you will... ...continuous integration/continuous deployment (CI/CD) to ensure high reliability and maintainability of our software. Troubleshoot and resolve...SeniorWork experience placementWork at officeRemote work3 days per week
$195.78k - $242.1k
...solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As Senior Software Engineer on the Consumer Frontend team, you will leverage the Roblox tech stack and tools to build groundbreaking experiences that...SeniorFull timeWork experience placementWork at officeLocal areaMonday to Friday- ...veterans from leading edge companies such as Facebook, LinkedIn, and Microsoft. Job Description We are looking for exceptional software engineers to scale our big data infrastructure and build innovative web and mobile products. Responsibilities Create elegant engineering...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- senior magento developer San Mateo, CA
- sr technical product manager San Mateo, CA
- senior manager pmo San Mateo, CA
- senior accountant part time San Mateo, CA
- senior application support engineer San Mateo, CA
- senior tax San Mateo, CA
- senior manager legal San Mateo, CA
- senior human factors engineer San Mateo, CA
- senior brand strategist San Mateo, CA
- senior aws cloud engineer San Mateo, CA

