AI Platform DevOps & SRE Lead
Reactor
Reactor is looking for a DevOps/SRE engineer in San Francisco to enhance the reliability and observability of their AI platform. The position requires running production Kubernetes clusters, strong CI/CD pipeline experience, and a solid understanding of GitOps and infrastructure as code. Successful candidates will triage production issues, manage secret infrastructure, and define SLOs. Benefits include a competitive salary, equity, and health coverage, with opportunities for relocation support. #J-18808-Ljbffr Reactor
- Qcells North America is seeking a Senior DevOps & SRE Manager - Platform Reliability & Global Operations to lead a blended DevOps and SRE organization, ensuring reliability, security, and scalable operations across a multi-platform ecosystem. You will manage global teams...Platform
- ...support virtualization and kernel performance for AI cloud infrastructure. You will develop tools, optimize compute platforms, and collaborate to enhance performance for AI... ...has over 8 years of experience in Compute SRE and is proficient with Linux kernel internals,...Platform
$350k
...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is... ...access to the knowledge and tools to make AI work for their unique needs and goals.... ...novel use-cases. We're hiring to grow the platform alongside the Tinker community. About the...PlatformLocal areaVisa sponsorshipWork visaRelocation package- ...Sierra, we’re creating a platform to help businesses build... ...experiences with AI. We are primarily an in-... ...using Terraform and modern DevOps tooling. Improve the... ...effective operation. Lead improvements to deployment... ...Define the foundation of SRE practices at Sierra, influencing...PlatformFull timeFlexible hours
- ...mission to create the world's first AI-powered Personal &... ...a Site Reliability Engineer (SRE) at Air Apps, you will be responsible... ...closely with development and DevOps teams to improve system design... ...AWS, Azure, or Google Cloud Platform (GCP) . Participate in on-...PlatformTemporary workWorldwide
$180k - $250k
...engineers who can help scale the platform. This is a strong fit for... ...core systems that power secure AI agent execution. This person will... ...Site Reliability Engineer The SRE role is focused on keeping our... ..., and observability workflows Lead incident response, root cause analysis...PlatformFull timeImmediate start- ...scalability across Sierra's AI-driven infrastructure.... ...with product and platform engineers to design systems... ...Terraform and modern DevOps tooling. Improving the... ...operation. Leading improvements to deployment... ...Defining the foundation of SRE practices at Sierra, influencing...PlatformFull timeFlexible hours
- Slope in San Francisco is looking for a reliability engineer focused on managing call completion for its Voice AI platform. You will be key in establishing incident management processes and improving system stability through effective monitoring and capacity planning. Candidates...Platform
- A pioneering tech company in San Francisco is seeking a dynamic Marketing Lead to build and lead their marketing function. This role focuses on creating high-impact programs targeted at engineers and driving demand generation. The ideal candidate should have a strong technical...Platform
$150k - $220k
TrueML is looking for a Senior Manager, DevOps to lead infrastructure and platform engineering efforts in San Francisco. The role involves driving cloud... ...reliability. The ideal candidate will have 10+ years in DevOps/SRE, a Bachelor's degree in Computer Science, and expertise...Platform- ...new sales opportunities. The ideal candidate has over 3 years of sales experience, especially in software or SaaS, and is skilled in lead generation techniques. The position offers growth opportunities and a positive work environment focused on team success. Applicants...Platform
- Block in San Francisco is looking for a Site Reliability Engineer who will enhance platform reliability and support critical services. You will work in a dynamic environment utilizing AI-driven tooling to improve system observability and incident response, ensuring that...Platform
- A leading AI technology company in San Francisco is looking for a Site Reliability Engineer to build high-performance, scalable systems for model serving. You will collaborate with teams to deploy optimized NLP models and ensure high availability. Ideal candidates have...Platform
- CodeRabbit is looking for an experienced Site Reliability Engineer to join our Platform Engineering team in San Francisco. In this role, you'll ensure high availability and performance of our AI-powered code review platform. Ideal candidates will have over 7 years of...Platform
- A leading blockchain and AI platform in San Francisco is seeking a Senior Site Reliability Engineer to ensure the reliability and security of their production infrastructure. You will maintain blockchain infrastructure, build automation, and collaborate with teams to embed...PlatformRemote work
- Resolve AI in San Francisco seeks a skilled Backend Engineer to create AI-driven environments for site reliability tasks. The role... ...research scientists. Candidates should have experience in cloud platforms like AWS, Kubernetes, and building scalable services. The position...PlatformFlexible hours
$300k
...Join a stealth-mode startup building out their AI and cloud platform, powered by thousands of H100s, H200s, and B200s, ready for experimentation... .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale...PlatformPermanent employment$150k - $215k
Nscale is looking for a Principal Observability Platform Engineer in San Francisco, California. You will own the technical strategy for Nscale... ...at scale. The ideal candidate will have over 8 years in SRE or platform engineering with hands-on experience in tools like Prometheus...Platform- Nscale is seeking a Staff Observability Platform Engineer based in San Francisco, California, to design and implement observability solutions for GPUs and AI workloads. The role entails partnering with SRE and engineering teams to enhance system visibility and reliability...Platform
- ...seeking a Senior Product Manager for AI and Automation in San Francisco. The role focuses on leading product strategy and execution... ...for PagerDuty’s automation platform, collaborating closely with teams... ...management and a strong understanding of SRE, with a passion for automation...Platform
- Together AI is seeking a Staff Platform Engineer to join the Product Foundations engineering organization and... .... The ideal candidate brings strong SRE and platform engineering background, excels... ...‑functional collaboration, and can lead critical infrastructure initiatives...Platform
- ...seeking an experienced Senior / Staff Engineer for our SRE, InfraSec team in Seattle. The role involves leading the security of cloud-based infrastructure,... ...with a focus on security, along with strong cloud platform skills. This position offers a competitive salary and...PlatformRemote job
$190k - $240k
Tatari is looking for a Data Platform Engineer to ensure the reliability and stability of its data platform infrastructure in San Francisco... ...than data engineering tasks. The ideal candidate has strong SRE skills and experience with cloud infrastructure, and is comfortable...Platform$120k - $168.49k
A leading educational technology company in San Francisco seeks a Site Reliability Engineer to... ...service delivery. The role requires 2+ years in SRE or related fields, strong programming skills, and familiarity with cloud platforms. The total compensation is competitive,...Platform- Our Real-Time AI / ML Tool Optimizes Signal-To-Noise Ratios in Observability Tools (such as Datadog, Splunk, New Relic etc.... ...: Splunk, Datadog, New Relic, etc. Previous Experience as a Platform Engineer or SRE Engineer Excited to be the First Solutions Engineer, Help Establish...Platform
- ...ABOUT THE ROLE We are hiring a Platform Engineer to own the runtime... ...developer experience. You keep an AI-native, multi-tenant platform... ...developer experience. Lead on-call practices and incident... ...years in platform/infrastructure/DevOps/SRE with production ownership. Strong...Platform
- BTIG seeks a DevOps/SRE to join our technology team to improve developer velocity through production support, infrastructure evolution, and strong operational continuity. You will own critical platform systems and drive reliability, scalable services, and customer‑facing...Platform
- DoorDash is seeking a Software Engineer for its Reliability Platform team to enhance infrastructure services. This role offers the chance... ...in Go, and familiarity with AWS. Additionally, a background in SRE principles is highly valued. Join us in delivering innovative solutions...Platform
$180k - $240k
...play a key role in enhancing the overall platform security and compliance. The ideal candidate has over 5 years of experience in DevOps/SRE, particularly with Azure systems. Compensation... ...and location. Join us to shape the future of enterprise AI. #J-18808-Ljbffr rPotentialPlatform- ...is looking for a Manager for Site Reliability Engineering in San Francisco, California. This role involves managing SRE teams to support the IDaaS platform, driving the microservice journey, and ensuring workload reliability. Candidates should have a strong technical...Platform
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Platform DevOps & SRE Lead. Be the first to apply!
- devops team lead San Francisco, CA
- lead devops engineer San Francisco, CA
- devops director San Francisco, CA
- director of digital platform San Francisco, CA
- digital platform specialist San Francisco, CA
- power platform San Francisco, CA
- platform manager San Francisco, CA
- platform product manager San Francisco, CA
- devops intern San Francisco, CA
- devops San Francisco, CA


