Site Reliability Engineer (SRE)
$350kThinking Machines Lab
The mission of Thinking Machines is to build AI that extends human will and judgment. About Tinker Tinker is our fine-tuning API that empowers researchers and developers to customize frontier AI to their needs — opening access to capabilities that have previously been concentrated in a handful of labs. We manage the infrastructure while allowing Tinkerers full flexibility in training open weights models with their own data, algorithms, and for their own needs. Tinker is rapidly adding new customers, features, and novel use-cases. We’re hiring to grow the platform alongside the Tinker community. About the Role We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research teams to make every layer of the system more robust and resilient. What You’ll Do Define and own end-to-end reliability, from CI/CD flows to production observability and incident response. Develop appropriate Service Level Objectives for distributed training systems, balancing job completion reliability and scheduling latency with development velocity. Design and implement monitoring and observability across the full training path. Drive incident response for Tinker platform issues, ensuring rapid recovery, thorough incident reviews, and systematic improvements that prevent recurrence. Harden multi-tenant isolation and resource scheduling so that LoRA-based workload co-scheduling maximizes utilization without compromising reliability or data separation Collaborate with security teams to address production vulnerabilities Skills and Qualifications Minimum qualifications: Bachelor's degree or equivalent experience in computer science, engineering, or similar. Experience in distributed systems, cloud infrastructure, or site reliability engineering. Proficiency writing software to solve reliability problems, including building tooling and automation. Experience with production incident response, postmortems, and systematic reliability improvement. Strong communication skills and track record of coordination across engineering and research teams. Preferred qualifications — we encourage you to apply if you meet some but not all of these: Deep experience operating production cloud services at scale (e.g., public cloud platforms, internal cloud services) Background in distributed training frameworks and how infrastructure failures surface in training behavior. Track record building checkpoint and recovery systems for long-running distributed jobs. Expertise in Kubernetes at scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads. Logistics Location: This role is based in San Francisco, California. Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 – $475,000 USD. Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together. Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description...SuggestedWork at officeVisa sponsorshipFlexible hours- ...globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You...SuggestedTemporary workWorldwide
$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,...SuggestedWork at officeLocal area1 day per week- ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge...SuggestedWork at officeWeekend work
$163.71k - $306k
...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system... ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially...Suggested$300k
...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and... .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting...Permanent employment$165k - $241.4k
...portfolios Your ImpactThe FedRAMP SRE team is focused on our... ...effective.We’re looking for talented engineers with a software or operations... ...teams to ensure the reliability, performance and security of... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...OpportunityTo achieve our ambitious goals, we’re looking for an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth....WorldwideHome officeFlexible hours
$165k - $225.6k
...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability... ...engineering teams to champion DevOps and SRE best practices, deliver excellent internal...Permanent employmentLocal areaWorldwideFlexible hours- ...what’s next.About the teamThe Engineering team at Airwallex is a diverse... ...working together to build scalable, reliable, and secure products that... ...to grow without borders.Our SRE team is breaking new engineering... ....What you’ll doAs a Senior Site Reliability Engineer, you’ll work...Temporary workLocal areaWorldwide
$148.5k - $223.9k
...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts... ...and our customers protected. The ExperienceAs an SRE, you will be a technical leader of the team driving...Full timeWorldwideWeekend work- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the...
$113.4k - $162k
...conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,... ...practices. Contribute to the design and implementation of new SRE best practices.You'll be a great fit if you have:Experienced...Temporary work$152.5k - $205k
...is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate... ...public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed...Flexible hours$117k - $209.33k
...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...cloud services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique...Full timeFor contractors$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support... ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and...Work at officeLocal areaRemote workWorldwideFlexible hours$140k - $205k
Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The... ...Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability,...Full timeTemporary workWork at officeFlexible hoursWeekend work- ...builds the platforms and tooling that help engineering teams develop, deploy, and operate... ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll... ...experience in backend systems, SRE, or platform engineering roles.Proven track...Permanent employmentWork experience placementWork at officeLocal area
- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools... ...of working in cloud-based systems operations, as a SRE or DevOps engineer. ~ First-hand experience with configuration...
$166.9k - $225.9k
...and career news. Job Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE... ...you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering,...Work at officeImmediate startWorldwideMonday to FridayFlexible hours- ...would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet,... ...Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability,...Remote work1 day per week
- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services... ...define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes...
- ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring... ...Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers... ...infrastructure, platform engineering, SRE, or DevOps. ~ Hands-on ownership of Kubernetes...Work at officeLocal areaFlexible hours
$155k - $222.6k
...soil. Meet the Team The SRE Fleet team is responsible for... ...platform. As a team of six engineers distributed across the US,... ...strong focus on automation, reliability, and operational excellence.... ...~2+ years of experience in Site Reliability Engineering, DevOps...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours- ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform... ...teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally...Contract workLocal area
$100k - $170k
...Site Reliability Engineer Houston; San Francisco; Seattle About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost... ...that makes AI work. The Role This is a career-level SRE role for someone who wants to own systems, not just watch them...Flexible hoursShift work- ...DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will... ...Skill Requirements: - Engineering Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production...
- ...SRE Location: San Francisco, CA (5 Days In-Office) You are the infrastructure... ...treatment. What We Look for in a Great Engineer You have the intensity and technical... ...feature release while maintaining the highest reliability. DevX Support: Support Developer...Work at office
$260k - $300k
...makers of Devin, the first AI software engineer. Our team is extremely talent-dense... .... You will own both the production reliability of our user-facing products and the... ...Strong software engineering fundamentals; SRE at Cognition means writing real code, not...$150k
...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- site reliability engineer San Francisco, CA
- site reliability engineer sre San Francisco, CA
- site reliability engineer remote San Francisco, CA
- site services specialist San Francisco, CA
- construction site safety San Francisco, CA
- site leader San Francisco, CA
- official site San Francisco, CA
- website content developer San Francisco, CA
- on site coordinator San Francisco, CA
- IT site lead San Francisco, CA

