Senior Site Reliability Engineer
TechChain Talent
About the job Senior Site Reliability Engineer
About the Company
Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient than most blockchain-based systems. It's designed so Stellar's ecosystem can make a real-world, lasting impact.
About the Role
SDF is looking for a Senior Site Reliability Engineer to help build and operate the foundation that powers our engineering teams. You'll ensure the reliability and scalability of our systems, design and improve the infrastructure behind our production environments, and automate operational work so developers can focus on building great products.
Key Responsibilities
- Maintain, improve, scale and secure our AWS/GCP infrastructure and Linux systems.
- Assist our development teams in running, packaging, deploying and troubleshooting applications
- Work with developers on streamlining deployment processes with Jenkins and other CI/CD tooling.
- Build, maintain, monitor and improve our Kubernetes clusters.
- Work with development teams on migrating applications to Kubernetes.
- Be responsible for maintenance and improvements to multiple internal services, for example Kubernetes, Prometheus, ELK.
- Monitor, triage and respond to alerts in our high availability environments.
- Participate in design and code reviews, and ensure that the foundation for our services is best in class.
- Evaluate new technologies, design and implement as appropriate.
- Identify automation opportunities and implement by creating custom or by using off the shelf solutions.
- Ability to understand Go, Rust, C++ and TypeScript source code
- Experience experimenting with AI-driven approaches to operations
- Comfortable with participating in on-call rotations and conducting thorough root cause analyses to keep systems running smoothly.
- Experienced in managing production workloads and skilled in using monitoring tools to detect issues early.
- A strong understanding of computer networking, TCP/UDP, load balancing, distributed computing, web services, and the fundamental protocols used by the internet ( DNS, etc.).
- No blockchain needed
- Experience using AI is a plus
- ...founders with PhDs in AI, Math, and Computer Science - is poised to redefine computing. About the Role We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and...Senior
- ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it...SeniorTemporary workWork experience placement
$210k - $240k
Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $...SeniorFull time$210.8k - $272.8k
About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user...SeniorLocal area$175k - $250k
...00.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance... ...scalability, performance, and reliability across environments. What You’ll Do...SeniorFull timeRemote workRelocationRelocation package- ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...platform capabilities in partnership with architects and product engineering Build a world-class observability platform and monitoring...Senior
$174.92k - $209.91k
...: to make access to data as simple and reliable as electricity. With Fivetran, customer... ...canonical and ready to query, with no engineering or maintenance required. We’re proud that... ...integrate our teams, systems, and career sites. About the Role Fivetran is building...SeniorFull timeWork at officeRemote work- ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will...Senior
$300k
...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring...SeniorPermanent employment$250k
...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves...SeniorPermanent employmentRemote work- ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems... ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes...Senior
- ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge...Work at officeWeekend work
$200k - $300k
...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round - Healthcare, AI Office Type: Onsite Salary: $200K-$300K Company Description We're representing a dynamic...Work at office- ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built...Relocation package
$100k - $170k
...Site Reliability Engineer Houston; San Francisco; Seattle About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services...Flexible hoursShift work$150k
...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security...$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering...Work at office$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description...Work at officeVisa sponsorshipFlexible hours$86k - $105k
...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best... ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,...Work at officeLocal area1 day per week- ...mission-critical industries, helping partners move more quickly and reliably from algorithm to silicon. Our platform accelerates deployment... .... The Roles We are looking for an experienced software engineer to help us build a new generation of transpilation tools...SeniorFull timeRemote workRelocation packageFlexible hours
- ...human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to... ...The Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible...Remote work1 day per week
- The company The future of data lies in decentralization, and the concept of a data mesh is the proven approach for implementing this at Enterprise scale. We’re here to make it a reality. Nextdata OS is a data-mesh-native platform built to meet the challenge of decentralizing...SeniorFull time
$204k - $281k
.... This is an opportunity to do career-defining work. We’re all in on this mission. If you are too, let’s talk. Manager, Site Reliability Engineering San Francisco, California Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on...Permanent employmentWorldwideFlexible hours$205k - $305k
...Director Of Site Reliability Engineering Interested in working on cutting-edge blockchain technology and creating equitable access to the global... ..., operate, and improve production services. This is a senior engineering leadership role reporting to the CTO. You will...Temporary workWork at officeLocal areaWorldwideFlexible hours- ...focused on building native actions — an agentic framework that understands you, and works reliably. We’re a team of AI researchers, designers, growth experts, and engineers rethinking human-computer interaction from the ground up. We value high-agency teammates who...SeniorFull time
- ...General Intelligence (AGI). Safety is more important to us than unfettered growth. About the Role We are looking for a senior software engineer to build the foundational platform for identity across all OpenAI products. This involves building authentication,...SeniorFull time
- ...company valued at $10 billion. We work in‑person five days a week in our new San Francisco headquarters. About the Role As a Site Reliability Engineer (SRE) at Mercor, you’ll own production reliability across our most critical systems, partnering directly with...
- ...excellence and quality. Set and uphold coding standards, architectural patterns, and review processes; mentor engineers. Co-own our culture of practical, reliable, and test-driven development that rapidly addresses business needs. Build for Scale. Contribute to...SeniorFull time
- ...What we do Idler builds reinforcement learning environments that teach AI models to code like 0.01% engineers. Our training environments are based on real-world coding scenarios that frontier models will actually encounter. We've closed a multimillion-dollar contract...SeniorFull timeContract workRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre San Francisco, CA
- site reliability engineer San Francisco, CA
- senior vice president communications San Francisco, CA
- senior manager quality engineering San Francisco, CA
- senior device engineer San Francisco, CA
- sr operations manager San Francisco, CA
- senior supervisor San Francisco, CA
- senior client services manager San Francisco, CA
- senior recruitment consultant San Francisco, CA
- senior platform engineer San Francisco, CA


