Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$119.8k - $234.7k

Microsoft

Job ID: 200043923Posted: 2026-07-21Location: United States, Washington, RedmondSalary: USD $119,800 - $234,700 per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual ContributorTravel: Less than 25%Profession: Software EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence. The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is to build the data platform for the age of AI, powering a new class of data-first applications and driving a data culture.Within Azure Data, the databases team builds and maintains Microsoft's operational Database systems. We store and manage data in a structured way to enable multitude of applications across various industries. We are on a journey to enable developer friendly, mission-critical, AI enabled operational databases across relational, non-relational and OSS offerings.We believe in making the day in the life of the On-Call Engineer boring while living up to the expectations of a massive cloud service with stringent Service Level Objectives (SLO’s). We do this by thinking differently, stretching ourselves to go all the way to the root of the problem, keeping data in front and center for all our decisions and taking a systems approach for generating outcomes that far exceeds the expectations. Helping attain the aspirational Service Level Objectives (SLO’s) through pragmatic innovation is what sets the SRE’s in Cosmos DB apart. If you share the same purpose, cause and belief and have passion to follow this pursuit, please read the rest of the Job description on what we do, and we would love to have you join us!Azure Cosmos DB is Microsoft’s next generation of globally distributed, massively scalable, multi-model cloud database service. It is designed to enable developers to build planet-scale applications. Azure Cosmos DB is one of the fastest growing Azure services. Joining the Azure Cosmos DB team is a fantastic opportunity to work with incredibly talented engineers operating like a startup and be at the forefront of building and shaping the Livesite Automation and AI Ops stack in Cosmos DB and lead the path for broader adoption across Microsoft Azure.Cosmos DB is a database of choice for the spectrum spanning from the hobbyist developer to the largest of Fortune 500 companies. The database provides the data backbone of many critical systems in Health Care, Retail, Telecommunications, IoT and many more where the Service Availability and Latency is paramount. Cosmos DB provides financially backed SLA (service level agreements) around 99.99 Availability and < 10 MS Latency and we are responsible of upholding ourselves to even more stringent Service Level Objectives (SLO) that delight our customers. Other than a resilient and fault tolerant architecture, a key to attaining the SLO’s is automating the root cause analysis and mitigation of issues and a lot of times proactively addressing the issues even before any customer impact. This team supports on building systems where a vast majority of Livesite issues are automatically mitigated without the need for human intervention.We are looking for a self-driven Senior Site Reliability Engineer (SRE) who likes taking a data driven and systems-based approach to solve Service Reliability problems. You will be responsible for building and optimizing solutions that can analyze massive amounts of telemetry and other Service Health indicators in near real time and perform automated root cause analysis and necessary mitigations to restore SLO’s.Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.ResponsibilitiesCollaborating closely with engineering teams on building and enhancing tooling and automation solutions for faster resolution of issues impacting SLO’s and averting incidents altogether when possible.Collaborating with the customers to understand their pain points around supportability and SLO attainment and formulate strategies for addressing recurring issues in a sustainable way.Communicate on a deeply technical level and be the single point of contact for interfacing with enterprise customers for handling service escalations and driving the issues to resolution.Ability to design and implement any changes to service telemetry for the automation to consume if it is not already available.Enhancing customer facing experience by proactive alerting based on utilization, trends, resource health, etc.Analyze data and provide operational insights into customer experience to design and product teams, so that we can design features with supportability in mind.Embody our culture and values.QualificationsRequired/Minimum Qualifications:6+ years technical experience in software engineering, network engineering, or systems administrationOR Bachelor's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administrationOR Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration.Other Requirements:Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.Preferred/Additional Qualifications:4+ years of experience running large scale cloud services.2+ years of operational experience in improving Service Reliability, Availability and Performance.Understanding of Observability and MELT implementation patterns for large-scale services.Experience in Logic Apps and authoring Jupyter Notebooks.Experience in analyzing, troubleshooting, and automating root cause analysis and mitigation of incidents impacting large-scale distributed systems.Systematic problem-solving approach, coupled with effective communication skills and a sense of curiosity.Ability to deal with the ambiguity associated with working in a fast-paced environment.Influencing the product architecture and roadmap to make sure the customer-experienced supportability is always a key consideration when evolving the product. #azdat #azuredata #SRESite Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year. Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Redmond, WA vacancy
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual...  ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type:...  ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewAre...  ...than the Microsoft Defender engineering team. We are looking for a Senior Site Reliability Engineer who will... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago
  • $174k - $253k

     ...Minimum qualifications Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5...  ...s degree in Computer Science or Engineering. About The Job Site Reliability Engineering (SRE) is what you get when you treat operations... 
    Senior
    Temporary work

    Google

    Kirkland, WA
    3 days ago
  • $165k - $225.6k

     ...core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. THE SENIOR SITE RELIABILITY ENGINEER OPPORTUNITY Reporting to the Manager, Site Reliability Engineering, this role will help... 
    Senior
    Permanent employment
    Full time
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  • $165k - $270k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Senior
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Redmond, WA
    3 days ago
  • $165k - $230k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Senior
    Permanent employment
    Temporary work
    Work at office
    Immediate start
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    10 hours ago
  • $165k - $230k

     ...developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. KUBERNETES PLATFORM SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s... 
    Senior
    Permanent employment
    Temporary work
    Work at office
    Worldwide
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    3 days ago
  • $165k - $230k

     ...the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. HARDWARE / INFRASTRUCTURE SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s... 
    Senior
    Permanent employment
    Temporary work
    Work at office
    Worldwide
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    3 days ago
  • $102.1k - $202.2k

     ...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual...  ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...Sovereign team, you will collaborate with engineers across disciplines to deliver and maintain... 
    Ongoing contract
    Work experience placement
    Local area
    Remote work
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...a Principal Site Reliability Engineering Manager to lead a team responsible...  ...impact through others—developing senior and principal engineers,... 
    Ongoing contract
    Temporary work
    Fixed term contract
    Local area
    Immediate start
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago
  •  ...manage AI resources on Microsoft Azure, including AI Foundry and RAG solutions Monitor and ensure service uptime, availability, reliability, and latency Track and integrate SRE metrics with enterprise monitoring systems Support CI/CD and DevOps workflows using... 

    Tech M USAAvance Consulting

    Redmond, WA
    10 hours ago
  •  ...to align with the unique needs of each client.Job DescriptionHi,Hope you are doing great!!We have an urgent opening for senior reliability engineer and the job description is as follows :Location: Redmond, WADuration: FTEKeywords: Reliability, Polymer/plastics/adhesives... 
    Senior

    EROS Technologies

    Redmond, WA
    3 days ago
  • $184k - $287.5k

     ...their best work. Come join the team and see how you can make a lasting impact on the world.NVIDIA is seeking a Sr. Systems Software Engineer for the Apache Spark Acceleration group. Over the past five years GPU accelerated data processing has moved from proof of concept... 
    Senior
    Full time

    Nvidia

    Redmond, WA
    5 days ago
  • $184k - $287.5k

     .... Come join the team and see how you can make a lasting impact on the world!NVIDIA is searching for a highly motivated, technical engineer to join the Tegra system-on-chip (SoC) software organization. You will work on key aspects of our ARM SW ecosystem and system software... 
    Senior
    Full time
    Remote work

    Nvidia

    Redmond, WA
    3 days ago
  • $184k - $287.5k

     ...learning, supercomputing, gaming, and visualization. As a Senior System Software Engineer on the NvSci team, you will play an integral role in...  ...tools, and generative AI technologies to improve software reliability, maintainability, and scalability.What we need to see:BS... 
    Senior
    Full time
    Remote work

    Nvidia

    Redmond, WA
    1 day ago
  • $152k - $241.5k

     ...join the team and see how you can make a lasting impact on the world.Join the leading Tegra Tools team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding opportunity to work on breakthrough technology that drives everything from self-driving... 
    Senior
    Full time

    Nvidia

    Redmond, WA
    3 days ago
  • $119.8k - $234.7k

     ...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type:...  ...world is watching. We are looking for a Senior Software Engineer to help architect, build, and run it.This...  ...the technical direction for scale, reliability, performance, and security across the... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  • $168k - $270.25k

     ...their best work. Come join the team and see how you can make a lasting impact on the world.The NVIDIA Experience (NVEX) Solutions Engineering team is looking for an experienced software engineer focused on customer support of NVIDIA products including AI Enterprise,... 
    Senior
    Full time
    Weekend work

    Nvidia

    Redmond, WA
    3 days ago
  • $152k - $241.5k

     ...to assess the performance of future GPU hardware features.What we need to see:Masters or PhD degree in Computer Science, Computer Engineering, or related field (or equivalent experience).3+ years of relevant industry experience.Strong proficiency in C++ programming and... 
    Senior
    Full time

    Nvidia

    Redmond, WA
    1 day ago
  • $159.2k - $215.3k

     ...orbit satellite network. Our mission is to deliver fast, reliable internet connectivity to customers beyond the reach of existing...  ...asylum. We are seeking a highly skilled and motivated Senior Reliability Engineer to join the Hardware Reliability Engineering team within the... 
    Senior
    Permanent employment
    Flexible hours

    Amazon

    Redmond, WA
    2 days ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ..., AI strategy, full stack engineering, Security, Dataverse & D365?...  ...of AI-Led engineering.As a Senior/Principal Software Engineer,...  ...quality, performance, and service reliability. Reliability & Operational... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  • $168.1k - $227.4k

     ...Amazon's low Earth orbit satellite network delivering fast, reliable internet connectivity to customers beyond the reach of...  ...through rigorous automated testing.We are hiring a Senior Software Development Engineer to lead the design and development of automation that validates... 
    Senior
    Permanent employment
    Internship
    Local area
    Immediate start
    Flexible hours

    Amazon

    Redmond, WA
    3 days ago
  • $119.8k - $234.7k

     ...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual...  ...design and build services that empower engineers and scientists across the company to...  ...expertise in distributed systems, service reliability, and experimentation methodologies. You... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    3 days ago
  • $119.8k - $234.7k

     ...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual...  ...looking for a Streaming and Captioning Engineering Lead to develop and maintain it. You...  ...in real time, and render it all in fast, reliable, accessible players. You will set technical... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  • $119.8k - $234.7k

     ...9,800 - $234,700 per yearEmployment type: Full-TimeWork site: 4 days / week in-officeRole type: Individual ContributorTravel...  ...is focused on continuing to lead. The Edge Web Platform engineering team is looking for a Senior Software Engineer - Browser Platform to join its on-... 
    Senior
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    4 days ago
  •  ...Job Functions • Architect, design and implement firmware and low-level drivers for Intel SoC . • Interface with other members of engineering team while gathering requirements, use cases and validation scenarios. • Perform low-level debugging of various firmware and OS components... 
    Senior

    System Canada Technologies

    Redmond, WA
    4 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low-precision or sparsified... 
    Senior
    Full time

    Nvidia

    Redmond, WA
    3 days ago
  • $152k - $241.5k

     ...fast, functional, and timely kernel delivery to customers.What we need to see:Masters or PhD degree in Computer Science, Computer Engineering, or related field (or equivalent experience).3+ years of relevant industry experience.Strong proficiency in C++ programming and... 
    Senior
    Full time

    Nvidia

    Redmond, WA
    1 day ago
  • $137.2k - $205.38k

     ...including Urban Air Mobility (UTM).Echodyne is seeking a Senior Software Engineer, IoT Platform, to join our fast-growing team. Who You Are...  ...the components living on the radar are built with clarity, reliability, and real-world mission constraints in mind. Required Experience... 
    Senior
    Full time
    Temporary work
    Local area

    Echodyne

    Kirkland, WA
    3 days ago
  • $152k - $241.5k

     ...QA as the performance representative for the CUTLASS team.What we need to see:Masters or PhD degree in Computer Science, Computer Engineering, or related field (or equivalent experience).3+ years of relevant industry experience.Strong programming skills in Python and C++... 
    Senior
    Full time

    Nvidia

    Redmond, WA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!