Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE

$207k - $300k

Google

Manage a team of Software/Systems Engineers on projects for users and remain directly responsible for uptime.Own the end-to-end availability and performance of key services, build automation to prevent problem recurrence, and automate responses to all non-exceptional service conditions.Mentor the team by example and establish credibility through quality technical execution.Coordinate on-call rotations across continents, using a follow-the-sun model.Design, write and deliver software to improve the availability, scalability, latency and efficiency of Google's services.Minimum qualifications:Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.8 years of experience with software development in one or more programming languages.3 years of experience designing, analyzing, and troubleshooting distributed systems.3 years of experience managing people or teams.3 years of experience leading projects.Preferred qualifications:Master's degree in Computer Science or Engineering.Experience effectively, efficiently, and responsibly applying AI tooling and workflows to engineering practices.Experience leading teams through organizational consolidations, team mergers, or significant transitions, unifying operational standards, aligning team cultures, and establishing sustainable multi-site on-call rotations across distributed global hubs.Deep practical expertise in Site Reliability Engineering practices, including SLO/SLI definition, capacity planning, disaster recovery, complex incident management, and postmortem culture for mission-critical services.Proven track record of collaborating closely with Software Engineering (SWE), product teams, and policy/compliance stakeholders.Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to Google Cloud, while using your expertise in coding, algorithms, complexity analysis and large-scale system design. SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow. You will have direct ownership of Google's foundational data pipeline and will manage the end-to-end web journey, from planetary-scale fetching via Harpoon and Trawler to high-throughput indexing through Raffia and Indexing Engine.You will work with systems that form the live data backbone powering Tier-0 Search surfaces and critical training pipelines for Google DeepMind and Gemini models globally; shapeing operational norms, build a sustainable cross-site rotation, and guide a mission-critical organization at the frontier of Google's AI future.Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next-generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefitsLearn more about benefits at Google.Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.8 years of experience with software development in one or more programming languages.3 years of experience designing, analyzing, and troubleshooting distributed systems.3 years of experience managing people or teams.3 years of experience leading projects.

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE in San Jose, CA vacancy
  • $207k - $300k

     ...high-performing, SRE team to support...  ...systems powering core AI infrastructure....  ..., maintain high reliability standards, and...  ...experience with software development in...  ...of experience managing and growing engineering teams, including...  ...across multiple sites or timezones.Experience... 
    Suggested

    Google

    San Jose, CA
    1 day ago
  • $167.7k - $245.2k

     ...days per week on-site at Cisco offices in...  ...defining the future of AI resilience. Our...  ..., improving reliability and reducing risks...  ...confidently deploy and manage AI-powered...  ...Site Reliability Engineer (SRE), you will build,...  ...toolsCollaborate with software engineers and customers... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Jose, CA
    19 hours ago
  • $186.9k - $267.7k

     ...approximately 2 days per week on-site at Cisco offices in...  ...the future of AI resilience. Together our...  ...intended, improving reliability and reducing risks. This...  ...expertly deploy and manage AI-powered applications...  ...Site Reliability Engineer (SRE), you will provide technical... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Jose, CA
    19 hours ago
  •  ...by the Illumio AI Security Graph...  ...: 5 On-Site Days a Week in...  ...Headquarters Our Engineering team is driven...  ...Impact As an SRE Engineer II, you will be...  ...responsible for managing our multi-...  ...enhancing system reliability and...  ...for automated software delivery and deployment... 
    Suggested
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    4 days ago
  • $101k - $161k

     ...artificial intelligence, and software-defined networking to...  ...awards, such as Best Engineering Team, Best Company...  ...’re looking for Site Reliability Engineers to join our...  ...Service (CVaaS) global SRE team. SREs at Arista...  ...enterprise network management and streaming telemetry... 
    Suggested

    Arista Networks

    Santa Clara, CA
    1 day ago
  •  ...thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the...  ...ensuring enterprise-grade reliability by leveraging hardware acceleration...  ...and orchestration. Ensure software solutions are optimized for...  ...). Background in MLOps or SRE roles focused on high-... 
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    4 days ago
  •  ...Experts to evaluate AI-powered workflows across software development,...  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI...  ...for accuracy and reliability. Work with...  ...Site Reliability Engineering...  ...workflow and project management platforms such as... 
    Remote job
    For contractors

    YO AI Labs

    San Jose, CA
    2 days ago
  • $125.7k - $203.1k

     ...working together in Engineering, IT, Supply...  ...Impact As a Software Engineer Embedded Systems II, you will take...  ...ensure product reliability. Apply systems...  ...concurrency, memory management, and low-level...  ...organizations in the AI era - and...  ...Cisco careers site to discover more... 
    Full time
    Temporary work
    Apprenticeship
    Work experience placement
    Local area
    Flexible hours

    CISCO Systems

    Milpitas, CA
    3 days ago
  • $230k - $250k

     ...autonomous networking, giving engineers and AI agents the ability to know...  ....Forward is looking for a Site Reliability EngineerAbout the Role This...  ...not a "keep the lights on" SRE role. As our first or early...  .... Experience with network management or observability platforms... 
    Night shift

    Forward Networks

    Santa Clara, CA
    3 days ago
  • $148k - $235.75k

     ...unlimited potential of AI to define the next...  ...of innovative engineers who are building...  ...volume telemetry into reliable, job-centric...  ..., and safe change management. You’ll own SLOs/SLIs...  ...on. You’ll partner Software Engineering and...  ...distributed systems as SRE/DevOps/Platform... 
    Full time

    Nvidia

    Santa Clara, CA
    22 hours ago
  •  ...of a Technical Support Engineer within a SaaS (Software as a Service) environment...  ...a growing focus on Site Reliability Engineering (SRE).The ideal candidate has...  ...support, and scale an AI Security Public SaaS platform...  ...with configuration management tools (e.g., Terraform)... 
    Full time
    Local area

    F5 Networks

    San Jose, CA
    4 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure...  ...is currently Tuesday.Engineering at Lambda is responsible...  ...for system deployment, management and maintenance.What You...  ...workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...re looking for a Senior SRE to join our Compute...  ...ll harness the power of AI to deliver groundbreaking...  ...Infrastructure‑as‑Code) and config management to standardize and...  ...management, fleet reliability/auto-healing, E2E...  ...or Ruby.Mentored other engineers and influenced technical... 
    Full time

    Nvidia

    Santa Clara, CA
    22 hours ago
  • $201k - $220k

     ...seeks a SR Systems Engineer, Mission Platforms...  ...the Muon Mission Foundry to enable families...  ...hardware, flight software, ground segment interfaces...  ...and Product Management to define the end-...  ...experience in AI&T workflows and V&...  ...citizen or national, (ii) U.S. lawful,... 
    Permanent employment
    Full time
    Temporary work
    Remote work
    Flexible hours

    Muon Space

    San Jose, CA
    3 days ago
  • $212.5k - $250k

     ...connected platform and AI-native ecosystem for private...  ...brings together the software, services, and legal...  ...infrastructure that founders use to manage equity, fund managers...  ...lifecycle so Carta's engineers can do the best work of...  ...Software Engineer II on DevEx, you will:... 
    Full time

    Carta

    Santa Clara, CA
    4 days ago
  • $165.2k - $223.6k

     ...AWS Neuron SDK, the software development kit...  ...software boundary, our engineers build systematic...  ...what'spossible in AI acceleration.As...  ...Software Engineer II - AI/ML, you will...  ...engineers, and product managers to deliver state-...  ...(design patterns, reliability and scaling) of... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ...You Will Contribute:Software Engineer II (Full Stack)Are you excited...  ...that power water management and conservation...  ...and leverage modern AI-assisted development...  ...while ensuring quality, reliability, and maintainabilityCollaborate...  ...days per week on-site in Los Gatos,... 
    Full time
    Temporary work
    Local area
    Flexible hours
    3 days per week

    Badger Meter

    Los Gatos, CA
    19 hours ago
  • $212.5k - $250k

     ...connected platform and AI-native ecosystem for...  ...Carta brings together the software, services, and legal...  ...that founders use to manage equity, fund managers...  ...SolveAs a Senior Software Engineer II, KYC, you will lead...  ...standards for performance and reliability.Simplify Systems: Dig... 
    Full time

    Carta

    Santa Clara, CA
    2 days ago
  • $165.2k - $223.6k

    As part of the AWS Applied AI Solutions organization, we...  ...of companies worldwide to manage day-to-day operations. We will...  ...their information.As a Software Development Engineer II, you'll take ownership of production...  ...root causes to maintain reliable data deletion and opt-out... 
    Permanent employment
    Internship
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    Santa Clara, CA
    3 days ago
  •  ...testing, and debugging of platform and system software tasks of small to medium complexity....  ...varying scope, the Software Development Engineer II implements high-quality code and deploys...  ...firmware interfaces.Experience using generative AI to optimize coding and software... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    F5 Networks

    San Jose, CA
    22 hours ago
  • $192.4k - $275.8k

     ...enterprise customers, blending Site Reliability Engineering, Systems Engineering, and...  ...Success, and Release Management turn to when navigating the...  ....7+ years of experience in SRE, cloud operations, and systems...  ...protect organizations in the AI era - and beyond. We’ve... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    4 days ago
  • $122.5k - $175k

     ...resilient, and secure. As an AI-forward enterprise,...  ...looking for a Staff Site Reliability Engineer to join our team....  .... You are an SRE with proven experience...  ...building infrastructure and managing platforms like...  ...and deploy systems and software in diverse environments... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    4 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing...  ...availability. It combines software and systems engineering...  ..., databases, capacity management, continuous delivery, and...  ...technical direction of NVIDIA’s AI Platform Runtime and lead... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $207.4k - $259.2k

     ...artificial intelligence (“AI”) solutions, and other technologies...  ...and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team....  ..., deployment, and release management.Champion cloud-first...  ...reliability is built into the software development lifecycle from... 
    Permanent employment
    Local area
    Worldwide
    Visa sponsorship

    Archer Aviation

    San Jose, CA
    1 day ago
  • $110k - $130k

     ...the World's leading AI-first Quality Engineering Company? Ready to advance...  ...are looking for a Site Reliability Engineer to join our...  ...our capacity management and performance management...  ...Incidents. SRE Skillsets - Expectations...  ...Engineer (SRE). ~ Software development "hands on... 
    Casual work
    Local area
    Flexible hours

    QualiTest Group

    Santa Clara, CA
    3 days ago
  • $207k - $300k

     ...years of experience with software development in one or...  ...experience in a people management or team leadership role....  ...Master’s degree or PhD in Engineering, Computer Science, or a...  ...projects across multiple sites internationally.The Google Cloud AI Research team addresses... 

    Google

    Sunnyvale, CA
    1 day ago
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build...  ...using the combination of software and systems engineering...  ...coding, database, capacity management, continuous delivery and...  ...NVIDIA’s next-generation AI-driven enterprise products... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $122.5k - $175k

     ...future of work is Human + AI and are building an...  ...looking for a Staff Site Reliability Engineer to join our team....  ...department. You are an SRE with proven experience...  ...infrastructure and managing platforms like Kubernetes...  ...deploy systems and software in diverse... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    2 days ago
  • $119k - $170k

     ...future of work is Human + AI and are building an AI-native...  ...are looking for a Staff Site Reliability Engineer (Production Engineer) to...  ...) reporting to the Senior Manager, Site Reliability Engineering...  ...-region fleet. This is a software-first SRE role: you will write... 
    Full time
    Work at office
    Local area
    Remote work
    Shift work
    3 days per week

    Zscaler

    San Jose, CA
    2 days ago
  • $132.6k - $214.5k

     ...and Inclusion. We weave AI into the fabric of...  ...and practices, having managed high cardinality metrics...  ...collaborate closely with our engineering teams to develop...  .... As a Senior Staff SRE with the Cortex Observability...  ...product and ensure the reliability and availability of our... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE. Be the first to apply!