Principal Systems Software Engineer - Observability and Telemetry Platform (Santa Clara)
$272k - $431.25kNvidia
Principal Systems Software Engineer at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized discipline which demands knowledge across different systems, networking, coding, database, capacity management, continuous delivery and deployment and open source cloud enabling technologies like Kubernetes and OpenStack. SRE at NVIDIA ensures that our internal and external facing GPU cloud services run maximum reliability and uptime as promised to the users and at the same time enabling developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency and performance. The Principal Systems Software Engineer is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on eliminating manual work through automation, performance tuning and growing efficiency of production systems.The lead Systems Software Engineer oversees how our systems connect and interact. We apply various tools and approaches to address a diverse set of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow.What you'll be doing:Design, implement and support operational and reliability aspects of large scale Observability & Telemetry collection platform with a focus on performance at scale, real time monitoring, logging and alertingEngage in and improve the whole lifecycle of services—from inception and design through deployment, operation and refinementSupport services before they go live through activities such as system design consulting, developing software tools, platforms and frameworks, capacity management and launch reviewsMaintain services once they are live by measuring and monitoring availability, latency and overall system healthScale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocityPractice sustainable incident response and blameless postmortemsBe part of an on call rotation to support production systemsWhat we need to see:BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience15+ years of experience with Infrastructure automation, distributed systems design, experience with design, develop tools for running large scale private or public cloud system in Production8+ years experience delivering foundational infrastructure and observability platforms.Experience in one or more of the following: Python, Go, Perl or RubyIn depth knowledge on Linux, Networking and ContainersWays to stand out from the crowd:Interest in crafting, analyzing and fixing large-scale distributed systemsSystematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive. Ability to debug and optimize code and automate routine tasksExperience in using or running large private and public cloud systems based on Kubernetes, OpenStack and Docker. Experience running Grafana, OpenTelemetry, Prometheus, and similar observability focused toolsYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 4, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time
$272k - $431.25k
...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU... ...geographiesFamiliarity with the open source data platform ecosystem (Apache Spark, Velox,... ...: US, CA, Santa Clara; US, IL, ChampaignType: Full time...PlatformFull timePart timeWork experience placement$272k - $431.25k
Principal System Software Engineer - Data Center MODS page is loaded## Principal System Software Engineer... ...Data Center MODSlocations: US, CA, Santa Clara: US, CA, Remotetime type: Full timeposted... ...), Redfish, HMC, BMC protocols and platform security.* Consistent track record...PlatformFull time$272k - $431.25k
...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team... ...deterministic compute architecture into a platform that compiler teams and data... ...RPC frameworks, coordination and telemetry patterns, MPI. Inference systems...PlatformFull timeShift work$272k - $431.25k
We are hiring senior engineers to work on the CUDA driver, a core component of our platform for accelerating general purpose computation... ...model across a range of system configurations and hardware capabilities... ...15+ years of relevant systems software development experience Strong...PlatformFull time$272k - $431.25k
...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation innovations in automotive platform software, system architecture, and performance engineering. In this highly visible technical leadership role, you will contribute directly...PlatformFull time$272k - $431.25k
...-on delivery across system software, drivers, and CUDA to... ...components, driver/platform layers, and performance... ...direction for an engineering team; mentor engineers... ...strict reliability, observability, and performance... ...SummaryLocation: US, CA, Santa Clara; US, TX, AustinType:...PlatformFull timePart time$272k - $431.25k
...looking for a Principal Software Engineer to join our CSP... ...firmware and GPU system software,... ...manageability, observability, security requirements... ...or accelerator platform engineering. BS... ...monitoring and telemetry: Xid errors,... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin...PlatformFull timePart timeRemote work- ...seeking a world‑class Principal Engineer (Sr Manager‑... ...and Platform Engineering (CIPE... ...deployment, and observability. Your work will... ...our standards for software quality, and unlock... ...at our dynamic Santa Clara California headquarters... ...of novel systems that leverage Large...PlatformFull timeWork at office3 days per week
$272k - $431.25k
...are looking for a Principal Software Engineer to join our DGX Cloud... ...the foundational systems that drive NVIDIA’s... ...directly craft the platform that fuels the future... ...) and innovative observability systems (e.g., Prometheus... ...: US, CA, Santa Clara; US, WA, SeattleType...PlatformFull timePart time- ...of technology. We are at the forefront of software and hardware innovation, pushing the... ...Location: Hybrid, working onsite at our Santa Clara, CA, headquarters 3 days per week. The role: Principal System Software Engineer, AI Inference Execution What you will...Full timeWork experience placement3 days per week
- ...NVIDIA is seeking a Sr. Principal Systems Software Engineer in Santa Clara to lead development of GPU-accelerated data processing for the Apache Spark ecosystem. You will design and implement Java, Scala, and CUDA/C++ libraries to speed up DataFrames, I/O, and interoperable...Full time
$272k - $431.25k
...NVIDIA is seeking a Principal Software Engineer for LPX System Software to build foundational software for a novel computing architecture in Santa Clara, California. The role involves shaping key system software components in Rust and driving architecture design while...Full time$272k - $431.25k
...We are looking for Principal Software Engineers to help shape the technical... ..., build critical systems, and turn ambiguous... ...repair, upgrades, observability, and readiness.... ...and influence platform, infrastructure, storage... ...: US, CA, Santa Clara; US, RemoteType: Full...PlatformFull timePart time$272k - $425.5k
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems page is loaded## Principal Software Engineer – Large-Scale... ...Storage Systemslocations: US, CA, Santa Clara: US, WA, Remote: US, MA,... ...budget of any single GPU, this platform enables efficient, resilient...PlatformFull timeLocal areaRemote work$2,972 per week
...Registered Nurse (RN) | Telemetry Location: Santa Clara, CA Agency: FlexCare Pay: $2,972 per week Shift Information: Days - 3 days... ...lasting relationships. · Fast-Track to Travel Platform: Our platform offers our clinicians an all-in-one hub for...PlatformFull timeContract workImmediate startShift work- ...We are seeking a Senior/Principal Software Engineer to join our dynamic Questa Sim (Simulation) R&D team at Siemens EDA in Santa Clara, USA . In this pivotal role, you will be instrumental... ...Concepts and Optimizations. Platform Proficiency: Experience working on UNIX...PlatformFull time
$272k - $431.25k
...We're looking for a Principal Engineer to join our CSP Engagements... ...targets on NVIDIA platforms. In this role, you... ...identify patterns and drive systemic improvements in... ...identify configuration, software, or workload... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,...PlatformFull timePart timeRemote work- ...Apply for Lead Principal Platform Software Engineer at Oracle in Santa Clara, CA, US. This Full time on site position offers great opportunities for career growth... ...and a passion for building large‑scale distributed systems to join our technical leadership team. Our team...PlatformFull time
$272k - $431.25k
...re looking for a Principal Software Engineer to join our CSP... ...to ensure NVIDIA platforms achieve target... ...incorporate their fleet telemetry and failure data... ...to distinguish systemic architectural... ...telemetry and observability systems: time-... ...SummaryLocation: US, CA, Santa ClaraType: Full...PlatformFull timePart time$272k - $431.25k
...assistants and engineering-productivity... ...we need a principal-level, hands... ...production systems and the architectural... ..., observability, release confidence... ...like mature software, not... ...frameworks, telemetry, and policy-... ...and related platform capabilities... ...SummaryLocation: US, CA, Santa ClaraType:...PlatformFull timePart timeLive in$208k - $260k
...offer a deep observability pipeline that... .... As a Principal Security Detections Engineer on the GigaSMART... ...and application telemetry into... ...high-performance systems software to identify suspicious... ...out of our Santa Clara, CA headquarters... ...security platforms. ~Author functional...PlatformLocal areaWorldwide3 days per week$221.2k - $387.1k
...all started when engineer Fred Luddy wrote... ...Our ServiceNow AI platform brings together any... ...Software EngineerThe engineering... ...this role: As a Principal Software Engineer... ...scale distributed systems, data ingestion,... ...scale, reliability, observability, fault tolerance,...PlatformPart timeWork at officeImmediate startRemote workFlexible hours$184k - $356.5k
...Senior System Software Engineer Platform - OpenBMC page is loaded Senior System Software Engineer Platform - OpenBMC Apply locations US, CA, Santa Clara US, Remote time type Full time posted on Posted 2 Days Ago job requisition id JR1999525 NVIDIA’s invention of the...PlatformFull timeSecond jobRemote work$248k - $391k
...We are looking for a Principal Software Engineer to drive the... ...learning loops.Own system design reviews, code... ...collaboration, and productivity platforms to create cohesive,... ...).Instrument robust observability, tracing, structured... ...: US, CA, Santa ClaraType: Full time...PlatformFull timePart time$248k - $391k
...seeking a highly skilled Principal Software Engineer to join our dynamic... ..., defining platform architecture, and optimizing... ...-class AI inference systems. Join us in this... ...hardware watchdogs, and telemetry for pre-release,... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType:...PlatformFull timePart time$272k - $431.25k
...working for us! The Cloud Engineering & Services team is... ...layer for agentic systems: signed policy, runtime... ...We are looking for a Principal Software Engineer, Agent Policy Fabric (APF) Core Platform, to join our Cloud... ...SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US,...PlatformFull timeContract workPart timeRemote work$99.6k - $234.6k
...Job Description The Principal AI Agent / ML Software Engineer is a Senior Staff-level,... ...operating next‑generation AI systems on Oracle Cloud... ...production‑grade agentic AI platforms, autonomous workflows, scalable... ..., evaluation, and observability. Design, architect, and...PlatformFull timeTemporary workFlexible hours$272k - $431.25k
...Security Data Engineer within our Infrastructure... ...fragmented telemetry from a 250,000... ...action our platform takes stands... ...the inventory systems, identity and... ...: A strong software engineering background... ...fleet-scale observability across... ...SummaryLocation: US, CA, Santa Clara; Canada,...PlatformFull timePart timeRemote work$272k - $431.25k
...We're looking for a Principal Software Engineer to join our CSP Engagements... ...for rack-scale system SW/FW, working with... ...recovery, health telemetry APIs, firmware update... ...in system software, platform firmware, or large-scale... ...: US, CA, Santa Clara; US, TX, Austin; US,...PlatformFull timePart timeRemote workShift work$272k - $431.25k
...seeking a highly skilled Principal Engineer to join our world-... ...scalability of the systems that power our global... ...projects across the platform, and be a major contributor... ...regulated enterprise software delivery.Your base... ...: US, CA, Santa Clara; US, CA, RemoteType:...PlatformFull timePart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Systems Software Engineer - Observability and Telemetry Platform (Santa Clara). Be the first to apply!
- principal software engineer Santa Clara, CA
- senior principal software engineer Santa Clara, CA
- IT system engineer Santa Clara, CA
- systems software developer Santa Clara, CA
- system programmer Santa Clara, CA
- platform engineering manager Santa Clara, CA
- platform developer Santa Clara, CA
- platform engineer Santa Clara, CA
- principal architect Santa Clara, CA
- senior principal cloud computing engineer Santa Clara, CA









