Staff Site Reliability Engineer, Semantic Understanding
$207k - $300kJobleads-US
Staff Site Reliability Engineer, Semantic Understanding
Share Staff Site Reliability Engineer, Semantic Understanding
corporate_fare Google place San Jose, CA, USA
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 8 years of experience with software development in one or more programming languages.
- 4 years of experience in applying Design for Reliability techniques.
- 3 years of experience as a Site Reliability Engineer.
- 3 years of experience leading projects.
- 3 years of experience designing, analyzing, and troubleshooting distributed systems.
Preferred qualifications:
- Master's degree in Computer Science or Engineering.
- Experience in Generative AI, Generative AI Agent, Google Infrastructure.
- Experience in large-scale and secure fleet management of servers and components.
- Experience in enterprise risk assessments.
- Experience with process improvement and automation.
About the job
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to Google Cloud, while using your expertise in coding, algorithms, complexity analysis and large-scale system design. SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow. Semantic Understanding Platform (SUP) is the standard platform for hosting image, text, video and data understanding signals at Google. SUP is a unified platform for understanding and generating content, and developing signals from raw data. Our team ensures the reliability, scalability, and operation of several Semantic Understanding Services, which we develop jointly with our partner SWE teams. This role is open to both Site Reliability Engineering (SRE) Systems Engineering and SRE Software Engineering applicants. Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google .
- Develop and drive the technical strategy and roadmap for critical areas within SU SRE, mentoring team members to enhance system reliability and efficiency.
- Initiate, own, and lead large-scale, complex projects and programs to significantly improve the reliability, scalability, and performance of SU services, often spanning multiple teams and systems.
- Partner with development teams to influence system design and architecture, embedding reliability principles throughout the development life-cycle. Drive alignment on technical direction across teams, navigating priorities.
- Identify, analyze, and mitigate risks in production. Design and implement architectural improvements to ensure long-term service health, scalability, and efficiency.
- Participate in the oncall rotation, respond to incidents, and drive postmortem actions to prevent recurrence.
Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy , Know your rights: workplace discrimination is illegal , Belonging at Google , and How we hire .
Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.
Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.
#J-18808-Ljbffr Jobleads-US- ...Google LLC is seeking a Staff Site Reliability Engineer in San Jose, CA to strengthen the Semantic Understanding Platform. You will drive reliability, scalability and performance for SU services, mentor teams, and lead cross-team initiatives in a blame-free environment...Suggested
$235k - $250k
...strategic, hands-on Global Manager of Site Reliability Engineering to lead the reliability, release... ...consumer lag, schema evolution, delivery semantics, replay, and failure recovery.... ...technical decisions.Business expertise: Understand how platform failures and delivery...SuggestedPermanent employmentFull timeLocal areaFlexible hours$182.8k - $247.3k
...learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform... ...operational excellence Support core infrastructure (i.e understand, diagnose, and debug these systems in production)...SuggestedWork experience placement- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base... ...and Go, the proactive AI assistant that understands context and delivers help automatically... ...for building software to ensure the reliability of our back-end systems, working with engineers...SuggestedWorldwideHome officeFlexible hours
$180k - $230k
...re looking for a Senior SRE to own the reliability, scalability, and observability of our... ...work closely with platform and data engineering to keep high-throughput, data-intensive... ...concepts and tooling (ArgoCD/Flux) ~ Solid understanding of networking, distributed systems,...SuggestedWork at officeLocal areaImmediate startRemote work3 days per week$90k - $100k
...Site Reliability Engineer The Opportunity We are looking for a highly capable engineer to join our Platform and Site Reliability engineering... ...building globally-distributed systems ~ Strong understanding of security best practices ~ Infrastructure as Code: Terraform...Flexible hours$115.5k - $164.8k
...where you matter. Your Impact As an engineer on the APX SRE CloudOps team, you will... ...required human intervention with reliable, tested automation. You will also participate... ...‑call rotations and incident response. Understanding production firsthand gives you the...Work experience placementWork at officeRemote work$114k - $148k
...related criteria. Summary As a Site Reliability Engineer, you will focus on ensuring the platform... .... You will interact with internal staff, managers, and customers to implement... ...others in several technical areas. Understanding practical use of SOC/FedRAMP controls...Work experience placement- ...to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance... ...used for automation (Python, Go, or similar) ~ Deep understanding of distributed systems ~ Experience with container...Full timeTemporary workWork at officeWorldwideFlexible hours
- ...Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute... ...Strong scripting skills in Python, Bash, or Go. Solid understanding of Linux and Windows administration. Database Knowledge...Local area
- ...create AI systems that can accurately understand the universe and aid humanity in its pursuit... ..., highly motivated, and focused on engineering excellence. This organization is for... ...teammates. ABOUT THE ROLE: As a Site Reliability Engineer focused on campus reliability...Night shift
- ...closely with other teams of world‑class engineers to tenaciously and creatively solve... ...Build and maintain a comprehensive understanding of the platform and custom application... ...loops. Evaluate emerging AI tooling for reliability and operations use cases, and advocate...Full timeCasual workRemote workFlexible hours
$110k - $145k
...liaise with product and engineering teams to ensure... ...platform and product reliability. The ideal candidate... ...product teams to best understand how to monitor their applications... ...to operations center staff on platform usage and... ...a platform engineer, site reliability engineer,...Work experience placement- ...As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly... ...automation and operational tooling. ~ Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure...Work experience placement
$148.5k - $223.9k
...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the... ...services maintain peak performance and reliability. Understanding of AI/ML concepts applied to operations (e.g.,...WorldwideWeekend work- ...next generation of enterprise-ready products. About the Site Reliability Engineering Team The Site Reliability Engineering (SRE) team... ...environments - Are curious and proactive, with a strong desire to understand systems end-to-end and uncover hidden failure modes -...Remote work
$150k - $220k
...Senior Site Reliability EngineerJob detailsDepartment / EngineeringRemoteFull-time$150,000 USD... ...The RoleAs a Senior Site Reliability Engineer, you own significant pieces of our infrastructure... ...problem-solving approach.* Solid understanding of Unix/Linux operating systems and...Full timeRemote work$135.2k - $181.2k
...mechanical, and sensor-based systems to ensure reliability and performance. Configure,... ...development environments. Applied understanding of observability principles and associated... ...an interest in emerging data engineering tools and methodologies. Preferred...Worldwide$110k - $130k
...Site Reliability Engineer Are you interested in working with the World’s leading AI-first Quality Engineering Company? Ready to advance... ...Expectations from Pricing & Settlements team: Good understanding of hybrid infrastructure. Expertise with AWS. Expertise...Casual workLocal areaFlexible hours$71.6k - $119.4k
## Site Reliability Engineer IIApplylocations: Home based-Georgia: Home based-California: Home Based - Pittsburgh, PAtime type: Full timeposted... ...with engineers and non-technical team members to understand requirements and translate them into technical tasks.* Learn...Temporary workInternshipLocal areaWork from home- ...As an SRE at Wordbricks, you will keep our systems fast, reliable, and boring. You'll own the infrastructure and operations behind... ...experience with cloud providers and infrastructure as code ~ Solid understanding of networking, databases, and distributed systems ~ Calm...Remote workFlexible hours
- ...around the world depend on to keep their engineering teams focused on what matters most -... ...culture (yep, even our GTM team understands and loves this technology)! We bring integrity... .... About the Role: As a Site Reliability Engineer, you will play a critical role...Remote workFlexible hours
$104k - $178k
## Sr. Site Reliability Engineer IApply: Hybrid: NYC Global HQ: Full time: Posted 12 Days Ago: JR00000779# ****Who We Are****DV is the leader... ...improvements.### ### ****Technical Knowledge***** Deep understanding of networking, DNS, load balancing, and CDN technologies....Full time- ...within a cloud platform (e.g., AWS, GCP, Azure), with a strong understanding of the underlying components making it easy to adapt to or... ...troubleshooting under pressure, and driving postmortems to improve system reliability over time Design systems with security in mind, applying...Work at officeRemote workHome office
$83.54k - $137.24k
...Select how often (in days) to receive an alert: Site Reliability Engineer Location: Bethpage, NY, US, 11714 Brand: Optimum Requisition... ...AHV and hyperconverged infrastructure concepts. Understanding of virtual machine lifecycle management and infrastructure...Local area$100.1k - $180.2k
...brands, offering comprehensive engineering, supply chain, and... ...a vast network of over 100 sites worldwide, Jabil combines global... ...Jabil is seeking a Lead Site Reliability Infrastructure and Security... ...industrial environments.Strong understanding of cybersecurity frameworks,...Temporary workWork at officeLocal areaRemote workWorldwide- ...casual atmosphere. FreedomPay is seeking an Associate Site Reliability Engineer to help maintain the availability, performance, and reliability... ...Responsibilities Learn and develop a strong understanding of FreedomPay's platform and application ecosystem. Work...Full timeCasual workInternshipFlexible hours
- ...Infrastructure Team as a technical leader driving reliability, automation, and scalability across the... ...practices across teams, mentor senior engineers, and be a primary escalation point for... ...Sentry, Signoz or equivalents) ~ Deep understanding of cloud performance, able to diagnose...Casual workWork at officeRemote workFlexible hours
$169.3k - $304.7k
...maintaining fast, efficient, scalable, and reliable routing software and infrastructure... ...global platform. As a Principal Site Reliability Engineer - Network, you will be responsible... ...distributed systems Have a deep understanding of TCP/IP, BGP, load balancing and...Work experience placementWork at officeRemote work$150k - $170k
...global investing! About The Role As the Manager of Site Reliability Engineering, you'll lead a team of SRE Automation Engineers while remaining... ...: Proficient in Linux administration with a deep understanding of the TCP/IP stack, OSI model, DNS, and network troubleshooting...Full timeWork at officeRemote workWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer, Semantic Understanding. Be the first to apply!
- software engineer staff Kentucky
- technology administrator Kentucky
- assistant engineer Kentucky
- staff engineer Kentucky
- senior staff systems engineer Kentucky
- senior staff engineer Kentucky
- engineering aide Kentucky
- site safety Kentucky
- on-site clinical research associate (traveling/remote) Kentucky
- site services specialist Kentucky


