Principal Supercomputing Operations Software Engineer
$142.8k - $304.2kMicrosoft
Role Description
Microsoft Azure’s Artificial Intelligence and High Performance Computing (AI/HPC) organization powers some of the world’s largest cloud native supercomputers used for frontier AI training, scientific computing, and large scale distributed simulations. Our team builds and operates hyperscale GPU clusters that consistently place Azure among global leaders in the Top500, MLPerf, and Graph500 benchmarks. By joining us, you step into the engineering core responsible for ensuring these systems remain reliable, performant, and ready for the next wave of AI innovation.
As a Principal Supercomputing Operations Engineer, you serve as the technical authority and strategic owner for interconnect fabric operations across flagship AI supercomputing environments. You treat InfiniBand and GPU interconnect fabrics as a single end to end reliability domain, defining how they are operated, debugged, hardened, and scaled in production. This is a hands on, production first leadership role operating at the intersection of architecture, live operations, and reliability engineering.
You will lead the most complex and impactful fabric related incidents, making high stakes technical decisions under ambiguity while balancing availability, risk, long term correctness, and customer impact. Beyond resolving incidents, you define failure models, operational strategy, and systemic prevention mechanisms that reduce recurrence at fleet scale. Your impact multiplies through technical leadership: setting operational standards, influencing engineering direction across teams, mentoring senior engineers, and partnering deeply with platform, hardware, firmware, and service teams to drive durable reliability improvements.
You will architect and drive automation, diagnostics, and telemetry that materially improve operability and debuggability of interconnect fabrics, and author authoritative playbooks, TSGs, and escalation models relied on across the organization. Through your judgment, designs, and operational strategy, Azure’s largest AI platforms scale safely, predictably, and sustainably to meet the demands of next generation AI workloads.
Microsoft’s mission is to empower every person and organization on the planet to achieve more. We work with a growth mindset, innovate to empower others, and collaborate to realize shared goals. Our culture is rooted in respect, integrity, and accountability, and we strive to build an environment where every engineer can learn, grow, and have real impact. As part of this team, you’ll help shape the next generation of cloud scale AI infrastructure and contribute to an inclusive culture where your expertise makes a difference every day.
Responsibilities
- Serve as the technical authority and DRI for InfiniBand and GPU interconnect fabric operations across large scale AI supercomputing environments, ensuring sustained GPU availability, training stability, and SLA compliance.
- Lead and orchestrate complex, high severity fabric incidents end to end, including detection, triage, mitigation, recovery, and root cause analysis, making high impact decisions under ambiguity.
- Perform deep, multi layer systems debugging across InfiniBand, Subnet Manager, GPU interconnect, PCIe, GPUs, firmware, drivers, and OS layers to identify true root causes at fleet scale.
- Drive operational excellence and systemic prevention by identifying recurring failure patterns, defining reliability models and failure domains, and authoring authoritative TSGs, playbooks, and escalation frameworks adopted across teams.
- Architect and drive automation, telemetry, diagnostics, and tooling that materially improve detection, observability, debuggability, and mean time to mitigation, raising the operational bar for interconnect fabrics across the platform.
Qualifications
- Required Qualifications:
- Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
- Other Qualifications:
- Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
- Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
- Preferred Qualifications:
- Bachelor's Degree in Computer Science OR related technical field AND 10+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, OR Python OR Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
- 6+ years of experience operating large‑scale distributed systems, high‑performance computing (HPC), or artificial intelligence (AI) infrastructure in production environments.
- Demonstrated ownership of mission‑critical production infrastructure with direct impact on service availability, GPU workloads, and customer SLAs.
- Hands‑on experience operating and debugging interconnect fabrics supporting large‑scale compute workloads.
- Strong Linux systems knowledge with experience debugging low‑level infrastructure issues across operating systems, drivers, and services.
- Proven ability to reason across hardware, firmware, drivers, and software stacks to diagnose and resolve complex production issues.
Benefits
- Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year.
- There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.
- Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here .
- Job Posting TitleSenior Software Developer/Operations Engineer (CIS O&M)Job DescriptionJob Overview:Customer: Internal Revenue Services (IRS) Location: Remote, with occasional onsite support for meetings, deployments, production activities, or knowledge transfer as required...SuggestedContract workFor contractorsFor subcontractorRemote workWorldwideRelocation package
- ...Position Purpose: The Software Engineer Sr. Principal is responsible for guiding software design and development across multiple engineering teams and principal software engineers. Sr. Principals will help accelerate delivery speed and quality through technical expertise...PrincipalWork experience placementLocal areaRemote workNight shift
$160.3k - $240.5k
...à la recherche d’un(e) développeur(se) principal(e) en sécurité, à l’aise dans un rôle pratique... ...is seeking a senior, hands-on security engineer to serve as the team’s technical lead... ...We’re Looking ForExperience in Security Operations, Incident Response, Detection...PrincipalFull timeWork at officeRemote workWorldwide$143k - $160k
...areas of Cyber Security, Instructional Design and Training, Software Engineering and IT Support Services to improve the security and well-... ...to serve with MKS2. Software Developer - Journeyman / Operations Research Software Engineer Location: Naval Postgraduate...SuggestedFull timeFor contractorsWork at officeLocal areaRemote work$112k - $154k
...meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering... ...products we build improve patient outcomes globally. As a Principal Software Development Engineer in Test at Baxter, your role has a direct effect on...PrincipalFull timeTemporary workLocal areaFlexible hours$115k - $230k
...blockchain code. Come and join this ambitious mission as a research software engineer to work on automated analyses for provably secure and... ...needs, including analysis, design, automated testing, operations, CI/CD, measuring results, incorporating customer feedback,...PrincipalFull timeInternshipLocal areaRemote workFlexible hours$116.6k - $177.8k
...configuring, validating, and troubleshooting automation hardware and software to ensure reliable system performance and scalability.DUTIES... ...automation software development processes.Perform any other engineering function as assigned.EXPERIENCE AND QUALIFICATIONS:Bachelor’s...PrincipalFull timeTemporary workWork at officeRemote workFlexible hours$107.5k - $204.5k
...market leading businesses, world-class operations and investments in research and... ...Artificial Intelligence and Automation Engineer to drive engineering efficiency by leveraging... ...(AI) and solution automation. The Principal Software Engineer - Digital Engineering AI Enablement...PrincipalContract workTemporary workWork experience placementWork at officeRemote workFlexible hours- ...innovation and mentoring junior engineers are key.Job Description*This... ...innovating, building, and operating the best in class, most... ...About the Role:We’re seeking a Principal AI & Automation Engineer, to... ...automation expertise, modern software architecture, and AI/ML-...PrincipalFull timeWork at officeRemote workWorldwideShift work
$150k - $200k
Python, JAVA, C++, J avaScript, Spark analytics, Jupyter Notebooks, GHOSTMACHINE, MapReduce Due to federal contract requirements, United States citizenship and an active TS/SCI security clearance and polygraph are required for the position. Required: Must be ...PrincipalFull timeContract workTemporary workImmediate start$85.4k - $128k
...solutions for global security. Our Engineering and Sciences (E&S) organization pushes... ...be a part of our mission! As a Software Engineer/Principal Software Engineer at Northrop... ...Software Engineer designs, develops, operates, and maintains software and firmware...PrincipalFull timeContract workRelocation packageShift work- Software Engineer, Levels up to Principal and Lead, Golang, NATS, Microservices, Boston, MA Hybrid, 150k midpoint**Note: I also represent similar roles that are completely remote but must live in continental USCompensation Commensurate with experience, bonus, benefits,...PrincipalPermanent employmentLive inRemote work
$114k - $171k
...impossible. Our employees are not only part of history, they're making history.Northrop Grumman’s Space Sector is seeking a Principal Operations Systems Engineer - Level 3 to join our team in Aurora, CO. This position is 100% onsite and cannot accommodate telecommute work.This...PrincipalFull timeRemote workRelocation packageShift workNight shiftRotating shift$107.5k - $204.5k
...problems. With our three market leading businesses, world-class operations and investments in research and development, we offer... ...the strength of more than 100 years of experience and renowned engineering expertise to meet the needs of today’s mission and stay ahead...PrincipalTemporary workWork experience placementWork at officeRemote workFlexible hours- ...~Certificate/credential troubleshooting ~Consumer issue resolution ~Runbook updates ~24x7 operational readiness Qualifications ~Software engineering experience ~Enterprise API development ~Java ~Spring Boot ~REST and SOAP web services ~...Full time
$280k - $360k
...Full Stack Software Engineer (Staff, Principal, Distinguished) \ $280k – $360k • No equity \ \ We are hiring for a client in the social banking and digital wallet industry, driving innovation in consumer finance across emerging markets. The product focuses on...PrincipalFull timeWork at officeImmediate startRemote workRelocationFlexible hours- ..., and performance of enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across mission-critical systems. You... ...teams to embed SRE best practices throughout the software development lifecycle. Drive operational...PrincipalRemote workFlexible hours
$91.8k - $137.6k
...Aeronautics Systems sector team is seeking a highly motivated Software Engineer or Principal Software Engineer to join our team onsite in Redondo... ...that utilize Graphics Processor Units (GPU), supercomputers, and other advanced hardware for modeling & simulation,...PrincipalFull timeImmediate startRemote workRelocation packageFlexible hoursShift work$79.3k - $118.9k
...Northrop Grumman Aeronautics Systems Sector has an opening for a Software Engineer/Principal Engineer Softwareto join our Global Surveillance Division... ...office spaces to support the program and business needs. Operating on our 9/80 schedule means you will get every other Friday...PrincipalFull timeWork at officeRemote workRelocation packageShift work- ...around the world as an independent member of Nexia.We currently have an exciting career opportunity for an AI Solution Software Engineer (Staff/Principal) to join our Strategic AI team. CohnReznick is a hybrid firm and most of our professionals are located within a...PrincipalWork at officeLocal areaRemote workFlexible hours
$180k - $205k
...that even the most advanced supercomputers or AI systems will never reach... ...develops the algorithms and software needed to make these systems... ...Come join us to build the operating system for the world's first... ...PsiQuantum works closely alongside engineers and scientists in the...PrincipalFull timeShift work$75k
...We are seeking a visionary and highly technical AI Go-To-Market (GTM) Engineer to sit at the core of our revenue engine. This is not a traditional operations role; it is a hybrid of software engineering, data science, and revenue strategy. You will be responsible for...PrincipalFull timeLocal areaImmediate startRemote work$156k - $190k
Job Description:Software Engineer Principal I - RoboticsRedmond, WA (Hybrid - 3 days in-person)Join our team at Genie and embark on an exciting... ...providing innovative solutions, engaging our people, and operating in a sustainable way. We are committed to helping team members...PrincipalFull timeTemporary workLocal areaRemote workWorldwide2 days per week$91.8k - $137.6k
...Northrop Grumman Aeronautics Systems is looking to add a Software Engineer/Principal Software Engineer- Data Analyst to join our team on site... ...engineering teams.• Develop toolsets to exploit large sets of operational data supporting the maintenance of communication...PrincipalFull timeRemote workRelocation packageShift work$227.2k - $417k
...About the Role:As a Software Engineer on the ML Infrastructure team, you will collaborate closely... ...You will improve the way we deploy and operate our services and even contribute to... ...Software EngineerAdditional Details: As a Principal Engineer on the ML Infrastructure team,...PrincipalFull timeTemporary workLocal areaFlexible hours$150 per hour
Role Description GITI is seeking a Principal Software Engineer to support Cyber Operations Research and Development as the technical lead for production software development on a passive RF emitter identification and network analysis from real-time sensor data streams....PrincipalFull timeRemote work- ...website translation model. Our OneLink product is a collection of software services and tools that provide on-the-fly translated pages for the websites you use every day. Our team of software engineers & web developers create OneLink services and tools that provide the...PrincipalFull time
- ...develop security requirements, and continuous governance mechanisms. XFN collaboration Collaborate with the blue team, security operation, infrastructure, R & D, algorithm, data, and business teams to complete attack review, vulnerability repair, detection rule...PrincipalFull time
- ...shared experiences for everyone. The Engine Networking Team pulls the players... ...communication of the game state to all. As a Principal Engineer on this team you will help the... ...waste Understand what happens on the operating system level when certain code is completed...PrincipalFull timeWorldwide
- ...skills and performance. Job Description Position: Principal Software Engineer Location: Lowell, MA Duration: 12 Months Rate: Open... ...throughout DVT, detect board failures during normal operation where possible, and use in the manufacturing environment...PrincipalFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Supercomputing Operations Software Engineer. Be the first to apply!
- principal software engineer Remote
- senior principal software engineer Remote
- data center operations engineer Remote
- security operations center engineer Remote
- cloud operations engineer Remote
- production support engineer Remote
- remote operation drilling engineer Remote
- network operations center engineer Remote
- post production engineer Remote
- security operations engineer Remote





