Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE
$207k - $300kManage a team of Software/Systems Engineers on projects for users and remain directly responsible for uptime.Own the end-to-end availability and performance of key services, build automation to prevent problem recurrence, and automate responses to all non-exceptional service conditions.Mentor the team by example and establish credibility through quality technical execution.Coordinate on-call rotations across continents, using a follow-the-sun model.Design, write and deliver software to improve the availability, scalability, latency and efficiency of Google's services.Minimum qualifications:Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.8 years of experience with software development in one or more programming languages.3 years of experience designing, analyzing, and troubleshooting distributed systems.3 years of experience managing people or teams.3 years of experience leading projects.Preferred qualifications:Master's degree in Computer Science or Engineering.Experience effectively, efficiently, and responsibly applying AI tooling and workflows to engineering practices.Experience leading teams through organizational consolidations, team mergers, or significant transitions, unifying operational standards, aligning team cultures, and establishing sustainable multi-site on-call rotations across distributed global hubs.Deep practical expertise in Site Reliability Engineering practices, including SLO/SLI definition, capacity planning, disaster recovery, complex incident management, and postmortem culture for mission-critical services.Proven track record of collaborating closely with Software Engineering (SWE), product teams, and policy/compliance stakeholders.Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to Google Cloud, while using your expertise in coding, algorithms, complexity analysis and large-scale system design. SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow. You will have direct ownership of Google's foundational data pipeline and will manage the end-to-end web journey, from planetary-scale fetching via Harpoon and Trawler to high-throughput indexing through Raffia and Indexing Engine.You will work with systems that form the live data backbone powering Tier-0 Search surfaces and critical training pipelines for Google DeepMind and Gemini models globally; shapeing operational norms, build a sustainable cross-site rotation, and guide a mission-critical organization at the frontier of Google's AI future.Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next-generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefitsLearn more about benefits at Google.Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.8 years of experience with software development in one or more programming languages.3 years of experience designing, analyzing, and troubleshooting distributed systems.3 years of experience managing people or teams.3 years of experience leading projects.
$207k - $300k
...high-performing, SRE team to support... ...systems powering core AI infrastructure.... ..., maintain high reliability standards, and... ...experience with software development in... ...of experience managing and growing engineering teams, including... ...across multiple sites or timezones.Experience...Suggested- ...by the Illumio AI Security Graph... ...: 5 On-Site Days a Week in... ...Headquarters Our Engineering team is driven... ...Impact As an SRE Engineer II, you will be... ...responsible for managing our multi-... ...enhancing system reliability and... ...for automated software delivery and deployment...SuggestedWork experience placementImmediate start
- ...Role :- Site Reliability Engineer (SRE) Infrastructure & Agentic Automation Location :- Santa Clara, CA (Hybrid) Work Authorization... ...at the intersection of traditional infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry...Suggested
- ...leading technology company for AI and Bitcoin mining... ...and construction, equipment management, and daily operations. Bitdeer... ...and operates the fleet. The SRE Platform team builds the monitoring... ...build. As an entry-level Software Engineer on the SRE / Monitoring Platform...SuggestedFull timeContract workInternshipLocal area
- ...Job Description Job Description Site Reliability Engineer II Bay Area, offices in San Jose · Hybrid · 24/7 FedRAMP Operations · Rotational... ...genuinely can't afford downtime. If you're a couple of years into SRE or cloud support and want your next role to open doors into...SuggestedHourly payContract workFor contractorsShift workNight shiftWeekend work
- ...Experts to evaluate AI-powered workflows across software development,... ...infrastructure, DevOps, SRE, and platform engineering. You will test AI... ...for accuracy and reliability. Work with... ...Site Reliability Engineering... ...workflow and project management platforms such as...Remote jobFor contractors
$230k - $250k
...autonomous networking, giving engineers and AI agents the ability to know... ....Forward is looking for a Site Reliability EngineerAbout the Role This... ...not a "keep the lights on" SRE role. As our first or early... .... Experience with network management or observability platforms...Night shift- ...You Will Contribute:Software Engineer II (Full Stack)Are you excited... ...that power water management and conservation... ...and leverage modern AI-assisted development... ...while ensuring quality, reliability, and maintainabilityCollaborate... ...days per week on-site in Los Gatos,...Full timeTemporary workLocal areaFlexible hours3 days per week
- ...Superintelligence Cloud, is a leader in AI cloud infrastructure... ...is currently Tuesday.Engineering at Lambda is responsible... ...for system deployment, management and maintenance.What You... ...workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...Work at officeLocal areaWork from homeFlexible hours
- ...of a Technical Support Engineer within a SaaS (Software as a Service) environment... ...a growing focus on Site Reliability Engineering (SRE).The ideal candidate has... ...support, and scale an AI Security Public SaaS platform... ...with configuration management tools (e.g., Terraform)...Full timeLocal area
- ...testing, and debugging of platform and system software tasks of small to medium complexity.... ...varying scope, the Software Development Engineer II implements high-quality code and deploys... ...firmware interfaces.Experience using generative AI to optimize coding and software...Full timeWork at officeLocal areaRemote workHome office
$165.2k - $223.6k
As part of the AWS Applied AI Solutions organization, we... ...of companies worldwide to manage day-to-day operations. We will... ...their information.As a Software Development Engineer II, you'll take ownership of production... ...root causes to maintain reliable data deletion and opt-out...Permanent employmentInternshipLocal areaWorldwideFlexible hours$226.14k
...time to building systems that operate reliably on a global scale. When you work here,... ...we’d love to meet you. Job Title: Software Engineer II Location: 50 West San Fernando Street... ...improvement based on user feedback. Build AI-assisted tools and automation in Python...Full timeTemporary workWork at officeRemote workWorldwide$104.9k - $174.7k
...Site Reliability Engineer The Site Reliability Engineer role is responsible... ...gaps. Respond to system-management alerts and operational exceptions... ...Reliability Engineering (SRE), production operations,... ...troubleshoot, and support hardware, software, storage, network, cloud,...Temporary workLocal area$110k - $130k
...the World's leading AI-first Quality Engineering Company? Ready to advance... ...are looking for a Site Reliability Engineer to join our... ...our capacity management and performance management... ...Incidents. SRE Skillsets - Expectations... ...Engineer (SRE). ~ Software development "hands on...Casual workLocal areaFlexible hours$125.7k - $203.1k
Software Engineer Embedded Systems II Join a vibrant community of passionate professionals... ...and ensure product reliability. Apply systems... ...concurrency, memory management, and low-level... ...organizations in the AI era - and beyond. We... ...see the Cisco careers site to discover more...Full timeTemporary workApprenticeshipWork experience placementLocal areaFlexible hours$151k - $244.2k
...and Inclusion. We weave AI into the fabric of... ...and practices, having managed high cardinality metrics... ...collaborate closely with our engineering teams to develop... .... As a Senior Staff SRE with the Cortex Observability... ...product and ensure the reliability and availability of our...Full timeWork at office$192.4k - $275.8k
...enterprise customers, blending Site Reliability Engineering, Systems Engineering, and... ...Success, and Release Management turn to when navigating the... ...~7+ years of experience in SRE, cloud operations, and systems... ...protect organizations in the AI era – and beyond. We’ve...Full timeTemporary workLocal areaFlexible hours$248k - $396.75k
...Full time JR2023973 Site Reliability Engineering (SRE) at NVIDIA is an engineering... ...availability. It combines software and systems engineering... ...Kubernetes, databases, capacity management, continuous delivery, and... ...direction of NVIDIA’s AI Platform Runtime and lead...Full time$122.5k - $175k
...future of work is Human + AI and are building an... ...looking for a Staff Site Reliability Engineer to join our team.... ...department. You are an SRE with proven experience... ...infrastructure and managing platforms like Kubernetes... ...deploy systems and software in diverse...Full timeWork at officeLocal area3 days per week$122.5k - $175k
...future of work is Human + AI and are building an... ...looking for a Staff Site Reliability Engineer to join our team.... ...department. You are an SRE with proven experience... ...infrastructure and managing platforms like Kubernetes... ...deploy systems and software in diverse...Full timeWork at officeLocal area3 days per week- Software Engineer II TENEX is an AI-native, automation-first, built-for-scale Managed Detection and Response (MDR) provider. We are a force multiplier for defenders, helping... ...Design, develop, and deploy scalable and reliable backend services and APIs. Build and maintain...Remote workMonday to Friday
- ...• Solve complex reliability challenges at scale... ...Influence architecture and engineering culture at a company level... ..., implement, and manage highly available and scalable... ...as a Principal SRE with a strong focus on... ...artificial intelligence (AI) tools to support parts...
$207.4k - $259.2k
...artificial intelligence (“AI”) solutions, and other technologies... ...and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team.... ..., deployment, and release management.Champion cloud-first... ...reliability is built into the software development lifecycle from...Permanent employmentLocal areaWorldwideVisa sponsorship- ...technology company for AI and Bitcoin mining infrastructure... ..., equipment management, and daily operations.... ...contexts of the NeoCloud SRE platform — the multi-region... ...& SLO: alert-engine-framework, alert-correlation... .... Qualifications Software Engineering Experience:...Remote jobFull timeContract workLocal area
$168k - $270.25k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build... ...using the combination of software and systems engineering... ...coding, database, capacity management, continuous delivery and... ...NVIDIA’s next-generation AI-driven enterprise products...Full time$170k - $220k
...We are seeking a Senior Software Engineer in Test (SET II) to play a pivotal role in advancing our test engineering... ...device cameras in Android apps. Managing complex image manipulation and data... ...We may use artificial intelligence (AI) tools to support parts of the hiring...Full time$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (... ...of thousands of NVIDIA's software engineers worldwide. The cloud... ...you'll be doing:Serve as an SRE Architect part of GPU... ...speed and cost efficiency of AI development and testing systems...Full timeWork experience placementWorldwide$256k - $414k
...looking for a Senior Manager to lead the design,... ...gaming workloads, AI/ML training, and... ...throughput, and highly reliable interconnects... ...hardware vendors, and SRE groups to influence... ...Science or a related engineering field (or... ...Streaming, Network Site Reliability/. If you...Full timeLocal area$267k - $356k
...Cloud, is a leader in AI cloud... ....Lambda's Storage Engineering team is the backbone... ...industry, which means reliability and performance aren... ...behind Lambda's own software-defined data plane... ...new and existing sites using tools such as... ...with their management and data-plane APIs...Work experience placementWork at officeLocal areaWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE. Be the first to apply!
- site reliability engineer San Jose, CA
- site reliability engineer sre San Jose, CA
- software sales executive San Jose, CA
- software technology San Jose, CA
- software intern San Jose, CA
- entry level software sales San Jose, CA
- software sales representative San Jose, CA
- id software San Jose, CA
- software sales San Jose, CA
- internship software San Jose, CA



