Principal Systems Software Engineer - Observability and Telemetry Platform
$272k - $431.25kNVIDIA
Principal Systems Software Engineer at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized discipline which demands knowledge across different systems, networking, coding, database, capacity management, continuous delivery and deployment and open source cloud enabling technologies like Kubernetes and OpenStack. SRE at NVIDIA ensures that our internal and external facing GPU cloud services run maximum reliability and uptime as promised to the users and at the same time enabling developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency and performance. The Principal Systems Software Engineer is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on eliminating manual work through automation, performance tuning and growing efficiency of production systems.The lead Systems Software Engineer oversees how our systems connect and interact. We apply various tools and approaches to address a diverse set of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow.What you'll be doing:Design, implement and support operational and reliability aspects of large scale Observability & Telemetry collection platform with a focus on performance at scale, real time monitoring, logging and alertingEngage in and improve the whole lifecycle of services—from inception and design through deployment, operation and refinementSupport services before they go live through activities such as system design consulting, developing software tools, platforms and frameworks, capacity management and launch reviewsMaintain services once they are live by measuring and monitoring availability, latency and overall system healthScale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocityPractice sustainable incident response and blameless postmortemsBe part of an on call rotation to support production systemsWhat we need to see:BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience15+ years of experience with Infrastructure automation, distributed systems design, experience with design, develop tools for running large scale private or public cloud system in Production8+ years experience delivering foundational infrastructure and observability platforms.Experience in one or more of the following: Python, Go, Perl or RubyIn depth knowledge on Linux, Networking and ContainersWays to stand out from the crowd:Interest in crafting, analyzing and fixing large-scale distributed systemsSystematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive. Ability to debug and optimize code and automate routine tasksExperience in using or running large private and public cloud systems based on Kubernetes, OpenStack and Docker. Experience running Grafana, OpenTelemetry, Prometheus, and similar observability focused toolsYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 4, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time
$195k - $290k
...most advanced AI-native platform. We work on large scale distributed systems, processing almost 3... ...endpoint sensor that observes system activity,... ...behavior, and streams telemetry to the Falcon cloud for... ...this possible.This is a Principal Software Engineer role and the senior...PlatformFull timeWork experience placementWork at officeLocal area$249k
...for travelers everywhere.Principal Software Engineer, Observability Introduction to the Team:... ...employees. A singular technology platform powered by data and... ...in cloud, distributed systems, and observability. You will... ...Architect and Build Core Telemetry Pipelines: Lead the...PlatformFull time$272k - $431.25k
...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team... ...deterministic compute architecture into a platform that compiler teams and data... ...RPC frameworks, coordination and telemetry patterns, MPI. Inference systems...PlatformFull timeShift work$272k - $431.25k
...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU accelerated data processing has moved from... ...boundaries and geographiesFamiliarity with the open source data platform ecosystem (Apache Spark, Velox, Presto, Apache Arrow,...PlatformFull timeWork experience placement- ...library. He/she will participate in the core system design and development. Our target... ...proven records on infrastructure level software development experience ~2+ years Clojure... ...communities ~ Linux (Debian in particular) platform programming experience is a big plus....PlatformFull time
$184k - $287.5k
...Infrastructure organization is seeking a Senior System Software Engineer to lead the evolution of our next-generation Data & Observability Platform. We serve and collaborate directly with... ...engineers rely on to visualize chip telemetry, debug distributed pipelines, and ensure...PlatformFull time$200k - $220k
...passionate, and committed engineers, technologists,... ...AI Network Software Solution... ...automation, network observability on a scale. You will... ...most demanding AI platforms.This role will be... ..., and monitoring.System Performance & ObservabilityImplement... ...telemetry pipelines and...PlatformWorldwide$272k - $431.25k
...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation innovations in automotive platform software, system architecture, and performance engineering. In this highly visible technical leadership role, you will contribute directly...PlatformFull time$272k - $431.25k
...Center MODS organization seeks a Principal Engineer to architect and scale next-... ...L10 and L11 diagnostic systems for Cloud Service Providers... ...distributed systems and hardware / software interfaces is essential for... ..., HMC, BMC protocols and platform security.Consistent track...PlatformFull time- ...right place. As a Principal Engineer, Developer Platform Engineering at JPMorganChase... ...extreme loadBuild observability-driven capacity... ...with production telemetry, real user monitoring... ...tools within the Software Development Life Cycle... ...and payment system performance, such as...PlatformEarly shift
$272k - $431.25k
...We're looking for a Principal Software Engineer to join our CSP... ...customers to ensure NVIDIA platforms achieve target MTBI... ...their fleet telemetry and failure data into... ...you to distinguish systemic architectural gaps... ...level telemetry and observability systems: time-series...PlatformFull time$272k - $431.25k
...and hands-on delivery across system software, drivers, and CUDA to make... ...-mode components, driver/platform layers, and performance counter... ...technical direction for an engineering team; mentor engineers,... ...drivers with strict reliability, observability, and performance...PlatformFull time$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements... ...firmware and GPU system software, working... ...manageability, observability, security requirements... ...firmware, or accelerator platform engineering. BS or... ...monitoring and telemetry: Xid errors, thermal...PlatformFull timeRemote work$165.6k - $296.4k
...than 25%Profession: Software... ...experimentation systems that scale across... ...designing foundational engineering systems (instrumentation... ...confidence. As a Principal Growth Engineer in... ...quality, better telemetry, safer rollouts,... ...CoreAI CoreAI - Platform and Tools is Microsoft...PlatformOngoing contractLocal area3 days per week$272k - $431.25k
...assistants and engineering-productivity... ...Now we need a principal-level, hands-on... ...harden production systems and the... ...performance, observability, release confidence... ...like mature software, not prototypes... ...evaluation frameworks, telemetry, and policy-... ..., and related platform capabilities...PlatformFull timeLive in$114.6k - $234.6k
...Infrastructure (OCI) is looking for a Principal Software Engineer to lead the development... ...secure infrastructure systems that underpin the core of OCI’s compute platform. This role sits within... ...firmware updates and telemetry-based observability. You’ll drive solutions...PlatformTemporary workFlexible hours$231.4k - $331.8k
...Team You will join Cisco’s Platform & Identity Engineering Group, a foundational... ...to provide the underlying systems that support Cisco’s innovation... ...direction of Cisco’s software and technology solutions,... ...services such as observability, data platforms, or AI/ML...PlatformFull timeTemporary workLocal areaFlexible hours- ...experience as a senior software engineer, including experience... ...complex, large-scale systems Expert-level... ...least one major cloud platform (AWS, GCP, or Azure);... ...-enabled impact.As a Principal Software Engineer, you... ...improving reliability, observability, and scalability of core...PlatformApprenticeship
$185.5k - $265k
...Zero Trust Exchange platform. This innovation... ...leverage intelligent systems to stay ahead of evolving... ...are looking for a Principal Software Development Engineer to join our team.... ...are resilient, observable, and easy to evolve... ...intelligent microservice telemetry to optimize high-...PlatformFull timeWork at officeLocal areaRemote work3 days per week$118.3k - $224.9k
...Technology (AST) is seeking a Principal Software Engineer with specialization in... ...patches, and configurations.System migrations and cross-... ...Terraform, Ansible, Puppet, Salt)Observability dashboard tools (Datadog,... ...software developmentCross-platform development, focusing on...PlatformTemporary workWork experience placementWork at officeRemote workFlexible hours$208k - $260k
...Place to Work, we offer a deep observability pipeline that efficiently... ...educational organizations. As a Principal Software Engineer, you will lead the design and... ...will work across distributed systems, backend services, and modern cloud platforms to deliver scalable, secure,...PlatformLocal areaWorldwide3 days per week$190k - $210k
...StatesProducts - Engineering /Fulltime /... ...detailsTitle of position: Principal Software EngineerPosition... ...Director of SW Systems... ...systems, cloud platforms, security, and real... ...secure, reliable, observable, adaptable, and... ...systems, network telemetry, or large-scale...PlatformFull timeH1bLocal areaWork from homeWork visaShift work$224k - $356.5k
...multi-GPU, high-bandwidth architecture of this platform.We are looking for a deeply technical systems software engineer who will own AI stack readiness on DGX Station... .... Ability to read GPU traces and translate observations into actionable optimizations.Strong understanding...PlatformFull timeLocal area$175k - $275k
Cerebras Systems builds the world's largest... ...of the Embedded Software team, you will help... ...Wafer Scale Engine (WSE)—the world’s... ...services, and the Linux platform and BSP layers... ...reliability, observability, and long-term maintainability... ...hardware level telemetry pipelines. The...Platform$83k - $166.1k
Implements features in system software modules and automation pipelines... ...to standards. Supports platform bring-up tasks and assists... ...support efforts with platform engineering, operations, firmware development... ...logging, monitoring, and observability for effective debugging,...PlatformTemporary workFlexible hours$248k - $391k
...world. We are looking for a Principal Software Engineer to drive the evolution of... ...learning loops.Own system design reviews, code reviews... ...collaboration, and productivity platforms to create cohesive,... ...CI/CD).Instrument robust observability, tracing, structured logging...PlatformFull time$184k - $287.5k
...production workflows? Our Engineering Workflow Platform team builds the user-facing... ...with hands-on systems engineering to solve challenges... ...migrations, and improve workflow observability.What we need to see:... ...least 12 years of relevant software engineering experience.Production...PlatformFull timeImmediate start$224k - $356.5k
...Do you look at a build system, a release pipeline, or... ...re looking for a senior engineer to invent, set, and... ...enables us to create the software and firmware that drive... ...signing services, and the observability that proves it all works. When our platforms are fast and...PlatformFull timeImmediate start$200k - $322k
...looking for a highly skilled Senior Software Engineer to design and develop AIOps & Observability platforms at NVIDIA. The platforms are... ..., reliable, and distributed systems that can handle high traffic... ...aggregating, and visualizing telemetry and operational metrics.Ways To...PlatformFull time$147k - $237.5k
...TeamEngineering - Our engineering team is at the... ...CareerAs a Principal Engineer on... ...of our platform. You’ll think... ...broadly about all system components, weigh... ...scalable software features and infrastructure... ...monitoring, observability, and alerting... ...over security telemetry, configuration...PlatformFull timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Systems Software Engineer - Observability and Telemetry Platform. Be the first to apply!
- principal software engineer Santa Clara, CA
- system programmer Santa Clara, CA
- IT system engineer Santa Clara, CA
- systems software developer Santa Clara, CA
- platform engineer Santa Clara, CA
- platform developer Santa Clara, CA
- principal Santa Clara, CA
- senior principal cloud computing engineer Santa Clara, CA
- principal architect Santa Clara, CA
- senior principal scientist Santa Clara, CA

