Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

$350k
Full-time

Thinking Machines Lab

The mission of Thinking Machines is to build AI that extends human will and judgment. About Tinker Tinker is our fine-tuning API that empowers researchers and developers to customize frontier AI to their needs — opening access to capabilities that have previously been concentrated in a handful of labs. We manage the infrastructure while allowing Tinkerers full flexibility in training open weights models with their own data, algorithms, and for their own needs. Tinker is rapidly adding new customers, features, and novel use-cases. We’re hiring to grow the platform alongside the Tinker community. About the Role We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research teams to make every layer of the system more robust and resilient. What You’ll Do Define and own end-to-end reliability, from CI/CD flows to production observability and incident response. Develop appropriate Service Level Objectives for distributed training systems, balancing job completion reliability and scheduling latency with development velocity. Design and implement monitoring and observability across the full training path. Drive incident response for Tinker platform issues, ensuring rapid recovery, thorough incident reviews, and systematic improvements that prevent recurrence. Harden multi-tenant isolation and resource scheduling so that LoRA-based workload co-scheduling maximizes utilization without compromising reliability or data separation Collaborate with security teams to address production vulnerabilities Skills and Qualifications Minimum qualifications: Bachelor's degree or equivalent experience in computer science, engineering, or similar. Experience in distributed systems, cloud infrastructure, or site reliability engineering. Proficiency writing software to solve reliability problems, including building tooling and automation. Experience with production incident response, postmortems, and systematic reliability improvement. Strong communication skills and track record of coordination across engineering and research teams. Preferred qualifications — we encourage you to apply if you meet some but not all of these: Deep experience operating production cloud services at scale (e.g., public cloud platforms, internal cloud services) Background in distributed training frameworks and how infrastructure failures surface in training behavior. Track record building checkpoint and recovery systems for long-running distributed jobs. Expertise in Kubernetes at scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads. Logistics Location: This role is based in San Francisco, California. Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 – $475,000 USD. Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together. Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in San Francisco, CA vacancy
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description... 
    Suggested
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    San Francisco, CA
    23 hours ago
  •  ...globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You... 
    Suggested
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    23 hours ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Suggested
    Work at office
    Local area
    1 day per week

    Mithril

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Suggested
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    3 days ago
  • $163.71k - $306k

     ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 
    Suggested

    Retool

    San Francisco, CA
    4 days ago
  • $300k

     ...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and...  .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $165k - $241.4k

     ...portfolios Your ImpactThe FedRAMP SRE team is focused on our...  ...effective.We’re looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  •  ...OpportunityTo achieve our ambitious goals, we’re looking for an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth.... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    4 days ago
  • $165k - $225.6k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability...  ...engineering teams to champion DevOps and SRE best practices, deliver excellent internal... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  •  ...what’s next.About the teamThe Engineering team at Airwallex is a diverse...  ...working together to build scalable, reliable, and secure products that...  ...to grow without borders.Our SRE team is breaking new engineering...  ....What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts...  ...and our customers protected. The ExperienceAs an SRE, you will be a technical leader of the team driving... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 

    Alembic

    San Francisco, CA
    4 days ago
  • $113.4k - $162k

     ...conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,...  ...practices. Contribute to the design and implementation of new SRE best practices.You'll be a great fit if you have:Experienced... 
    Temporary work

    TextNow

    San Francisco, CA
    2 days ago
  • $152.5k - $205k

     ...is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate...  ...public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...cloud services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    19 hours ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support...  ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The...  ...Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability,... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    San Francisco, CA
    4 days ago
  •  ...builds the platforms and tooling that help engineering teams develop, deploy, and operate...  ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll...  ...experience in backend systems, SRE, or platform engineering roles.Proven track... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    3 days ago
  •  ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools...  ...of working in cloud-based systems operations, as a SRE or DevOps engineer. ~ First-hand experience with configuration... 

    TechChain Talent

    San Francisco, CA
    4 days ago
  • $166.9k - $225.9k

     ...and career news. Job Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE...  ...you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering,... 
    Work at office
    Immediate start
    Worldwide
    Monday to Friday
    Flexible hours

    Drata Inc

    San Francisco, CA
    23 hours ago
  •  ...would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet,...  ...Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability,... 
    Remote work
    1 day per week

    Runloop AI, Inc

    San Francisco, CA
    3 days ago
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services...  ...define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes... 

    Socket

    San Francisco, CA
    1 day ago
  •  ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring...  ...Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers...  ...infrastructure, platform engineering, SRE, or DevOps. ~ Hands-on ownership of Kubernetes... 
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    4 days ago
  • $155k - $222.6k

     ...soil. Meet the Team The SRE Fleet team is responsible for...  ...platform. As a team of six engineers distributed across the US,...  ...strong focus on automation, reliability, and operational excellence....  ...~2+ years of experience in Site Reliability Engineering, DevOps... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    San Francisco, CA
    23 hours ago
  •  ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform...  ...teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally... 
    Contract work
    Local area

    InterSources

    San Francisco, CA
    3 days ago
  • $100k - $170k

     ...Site Reliability Engineer Houston; San Francisco; Seattle About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost...  ...that makes AI work. The Role This is a career-level SRE role for someone who wants to own systems, not just watch them... 
    Flexible hours
    Shift work

    Nscale

    San Francisco, CA
    3 days ago
  •  ...DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will...  ...Skill Requirements: - Engineering Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production... 

    BayOne Solutions

    San Francisco, CA
    23 hours ago
  •  ...SRE Location: San Francisco, CA (5 Days In-Office) You are the infrastructure...  ...treatment. What We Look for in a Great Engineer You have the intensity and technical...  ...feature release while maintaining the highest reliability. DevX Support: Support Developer... 
    Work at office

    Latent

    San Francisco, CA
    5 hours ago
  • $260k - $300k

     ...makers of Devin, the first AI software engineer. Our team is extremely talent-dense...  .... You will own both the production reliability of our user-facing products and the...  ...Strong software engineering fundamentals; SRE at Cognition means writing real code, not... 

    Cognition Corp

    San Francisco, CA
    3 days ago
  • $150k

     ...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security... 

    VantageScore®

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!