Senior Site Reliability Engineer
$119.8k - $234.7kMicrosoft Corporation
Job ID: Posted: Location: United States, Washington, RedmondSalary: USD $119,800 - $234,700 per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual ContributorTravel: Less than 25%Profession: Software EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence. The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is tobuild the data platform for the age of AI, powering a new class of data-first applications and driving a data culture.Within Azure Data, the databases team builds and maintains Microsoft's operational Database systems. We store and manage data in a structured way to enable multitude of applications across various industries. We are on a journey to enable developer friendly, mission-critical, AI enabled operational databases across relational, non-relational and OSS offerings.We believe in making the day in the life of the On-Call Engineer boring while living up to the expectations of a massive cloud service with stringent Service Level Objectives (SLO’s). We do this by thinking differently, stretching ourselves to go all the way to the root of the problem, keeping data in front and center for all our decisions and taking a systems approach for generating outcomes that far exceeds the expectations. Helping attain the aspirational Service Level Objectives (SLO’s) through pragmatic innovation is what sets the SRE’s in Cosmos DB apart. If you share the same purpose, cause and belief and have passion to follow this pursuit, please read the rest of the Job description on what we do, and we would love to have you join us!Azure Cosmos DB is Microsoft’s next generation of globally distributed, massively scalable, multi-model cloud database service. It is designed to enable developers to build planet-scale applications. Azure Cosmos DB is one of the fastest growing Azure services. Joining the Azure Cosmos DB team is a fantastic opportunity to work with incredibly talented engineers operating like a startup and be at the forefront of building and shaping the Livesite Automation and AI Ops stack in Cosmos DB and lead the path for broader adoption across Microsoft Azure.Cosmos DB is a database of choice for the spectrum spanning from the hobbyist developer to the largest of Fortune 500 companies. The database provides the data backbone of many critical systems in Health Care, Retail, Telecommunications, IoT and many more where the Service Availability and Latency is paramount. Cosmos DB provides financially backed SLA (service level agreements) around 99.99 Availability and < 10 MS Latency and we are responsible of upholding ourselves to even more stringent Service Level Objectives (SLO) that delight our customers. Other than a resilient and fault tolerant architecture, a key to attaining the SLO’s is automating the root cause analysis and mitigation of issues and a lot of times proactively addressing the issues even before any customer impact. This team supports on building systems where a vast majority of Livesite issues are automatically mitigated without the need for human intervention.We are looking for a self-driven Senior Site Reliability Engineer (SRE) who likes taking a data driven and systems-based approach to solve Service Reliability problems. You will be responsible for building and optimizing solutions that can analyze massive amounts of telemetry and other Service Health indicators in near real time and perform automated root cause analysis and necessary mitigations to restore SLO’s.Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.ResponsibilitiesCollaborating closely with engineering teams on building and enhancing tooling and automation solutions for faster resolution of issues impacting SLO’s and averting incidents altogether when possible.Collaborating with the customers to understand their pain points around supportability and SLO attainment and formulate strategies for addressing recurring issues in a sustainable way.Communicate on a deeply technical level and be the single point of contact for interfacing with enterprise customers for handling service escalations and driving the issues to resolution.Ability to design and implement any changes to service telemetry for the automation to consume if it is not already available.Enhancing customer facing experience by proactive alerting based on utilization, trends, resource health, etc.Analyze data and provide operational insights into customer experience to design and product teams, so that we can design features with supportability in mind.Embody ourcultureandvalues.QualificationsRequired/Minimum Qualifications:6+ years technical experience in software engineering, network engineering, or systems administrationOR Bachelor's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administrationOR Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration.Other Requirements:Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.Preferred/Additional Qualifications:4+ years of experience running large scale cloud services.2+ years of operational experience in improving Service Reliability, Availability and Performance.Understanding of Observability and MELT implementation patterns for large-scale services.Experience in Logic Apps and authoring Jupyter Notebooks.Experience in analyzing, troubleshooting, and automating root cause analysis and mitigation of incidents impacting large-scale distributed systems.Systematic problem-solving approach, coupled with effective communication skills and a sense of curiosity.Ability to deal with the ambiguity associated with working in a fast-paced environment.Influencing the product architecture and roadmap to make sure the customer-experienced supportability is always a key consideration when evolving the product.#azdat #azuredata #SRESite Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year. Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
$119.8k - $234.7k
...thinking in a cloud-enabled world. Microsoft's Azure Data engineering team is leading the transformation of analytics in the... ...99.99 Availability and We are looking for a self-driven Senior Site Reliability Engineer (SRE) who likes taking a data driven and systems-...SeniorOngoing contractLocal area$160k - $210k
...change and achieving remarkable growth in a rapidly evolving industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale...SeniorWork at officeImmediate startRemote workWork from home$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability...SeniorOngoing contractLocal area3 days per week- ...Overview We are seeking a Senior Software Development Engineer to join a team focused on supporting hardware testing and validation environments... ...an existing codebase, troubleshoot issues, improve reliability, and work effectively at the intersection of software and...Senior
- ...to align with the unique needs of each client.Job DescriptionHi,Hope you are doing great!!We have an urgent opening for senior reliability engineer and the job description is as follows :Location: Redmond, WADuration: FTEKeywords: Reliability, Polymer/plastics/adhesives...Senior
- ...Implement quantized and sparse recipes within inference engines and manage model export pipelines to ensure correct serialization. Develop benchmarking harnesses, data analysis tools, and improve developer productivity through infrastructure and CI improvements. Requirements...Senior
- ...telemetry, and customer feedback. Requirements: Requires a Bachelor's degree in Computer Science and at least 4 years of technical engineering experience with C, C++, or C#. Preferred qualifications include experience with Windows client software and developing with large...Senior
$100k - $150k
...Senior .NET Solutions Developer - Remote Bright Vision Technologies is a technology... ...stacks. This role spans the full engineering lifecycle — requirements analysis, architectural... ..., observability, and operational reliability. Key Responsibilities Provide comprehensive...SeniorFull timeH1bLocal areaRemote workVisa sponsorship$140k - $160k
...addition to the responsibilities below, the Senior Project Integrator executes controls... ...standard platforms while collaborating with engineering to design robust industrial control... ...such as installation, commissioning, and site investigations. Perform systems integrations...SeniorTemporary workLocal areaRemote workDay shift- ...Participate in setting physics architecture: Define data models, simulation loops, solver strategies, and integration points across engine and simulation systems. Deliver core physics features: Collision detection (broad‑phase/narrow‑phase), entity controllers,...SeniorFlexible hours
$140k - $190k
...Senior Software Engineer Who We Are Kymeta revolutionizes satellite communications through Intelligent Communications Platforms (ICPs)... ...network, and system performance to diagnose issues and improve reliability, usability, and operational effectiveness ~...SeniorWorldwideFlexible hours$206.09k - $217k
...use AI and automation to amplify our impact. What We're Looking For Multiple position openings. We're looking for a Senior Software Engineer who will design and implement microservices-based architecture for Lightning AI's core platform services. Position allows...SeniorWork at officeRemote workWork from homeFlexible hours$140k - $200k
...Uncrewed Aircraft Systems (UAS), and Airspace Management including Urban Air Mobility (UTM). Echodyne is seeking a Sr. Software Engineer to help build our next-generation radar software platform, advancing mission-critical security for defense and automotive...SeniorFull timeContract workTemporary workFlexible hours$165k - $270k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts....SeniorPermanent employmentTemporary workWork at officeImmediate startMonday to FridayWeekend work$165k - $270k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...SeniorPermanent employmentTemporary workWorldwideWeekend work$165k - $230k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink...SeniorPermanent employmentTemporary workWorldwideWeekend work$104.1k - $162.9k
...Description Job Description The Senior Embedded Application Software Engineer is responsible for leading the... ...integrate, and maintain high-quality, reliable embedded software that enhances... ...Whether on highways, construction sites, urban centers, or forest roads, PACCAR...SeniorH1b- ...the ground and in the air. Kapta Space is seeking a Senior Embedded Firmware Engineer to lead the complete lifecycle development, including development... ..., hardware, FPGA and mechanical engineers to build reliable, embedded firmware for the payload Assist in Hardware...SeniorPermanent employmentFull time
- ...Senior Full Stack Java Developer Location: Issaquah, WA (100% Onsite) Job Type: Contract (Consulting) Duration:... ...Qualifications Bachelor's degree in Computer Science, Software Engineering, or a related field. Strong analytical, troubleshooting,...SeniorContract work
$165k - $230k
...the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. HARDWARE / INFRASTRUCTURE SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s...SeniorPermanent employmentTemporary workWork at officeWorldwideMonday to FridayWeekend work$130k - $195k
...Connectivity is a design house. We are a team of 600 talented engineers, and our main office is in Lund, Southern Sweden. Primarily, we... ...will never work alone. Job Overview: We are seeking a Senior Android Power and Thermal Engineer with a strong background in...SeniorTemporary workWork experience placementWork at officeFlexible hours$149.8k - $262.2k
...Company Description It all started when engineer Fred Luddy wrote code that automated a... ...About the team: We are seeking a Senior Staff Software Engineer (IC5 Level) to... ...them to resolution, with a strong focus on reliability, scalability, and operability — leading...SeniorPermanent employmentFull timeWork experience placementWork at officeImmediate startRemote workFlexible hours2 days per week- Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...
$232k - $319k
...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...with self-service Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and...SeniorPermanent employmentLocal areaWorldwideFlexible hours$190.9k - $334.1k
...It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful...SeniorWork at officeImmediate startRemote workFlexible hours$204k - $306k
...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from... ...week in our San Francisco Office.The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$102.1k - $202.2k
...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany:... ...Sovereign team, you will collaborate with engineers across disciplines to deliver and maintain...Ongoing contractWork experience placementLocal areaRemote work3 days per week$140k - $200k
...Senior Systems Engineer Kymeta revolutionizes satellite communications through Intelligent Communications Platforms (ICPs). Our electronically steered flat panel antennas enable seamless communications on-the-move. Kymeta solutions serve government, military, maritime...SeniorWorldwideFlexible hours$119.8k - $234.7k
...Microsoft Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft's expanding Cloud Infrastructure... ...architecture, and deliver against product goals for quality, reliability, and performance. Collaborate with internal, external, and...SeniorOngoing contractWork at officeLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Redmond, WA
- senior operations technician Redmond, WA
- senior cloud service delivery manager Redmond, WA
- senior it service manager Redmond, WA
- senior chief engineer Redmond, WA
- sr operations manager Redmond, WA
- senior physical design engineer Redmond, WA
- senior energy engineer Redmond, WA
- senior financial analyst remote Redmond, WA
- senior manager accounts payable Redmond, WA


