Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Systems Software Engineer - Observability and Telemetry Platform

$272k - $431.25k

NVIDIA

Principal Systems Software Engineer at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized discipline which demands knowledge across different systems, networking, coding, database, capacity management, continuous delivery and deployment and open source cloud enabling technologies like Kubernetes and OpenStack. SRE at NVIDIA ensures that our internal and external facing GPU cloud services run maximum reliability and uptime as promised to the users and at the same time enabling developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency and performance. The Principal Systems Software Engineer is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on eliminating manual work through automation, performance tuning and growing efficiency of production systems.The lead Systems Software Engineer oversees how our systems connect and interact. We apply various tools and approaches to address a diverse set of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow.What you'll be doing:Design, implement and support operational and reliability aspects of large scale Observability & Telemetry collection platform with a focus on performance at scale, real time monitoring, logging and alertingEngage in and improve the whole lifecycle of services—from inception and design through deployment, operation and refinementSupport services before they go live through activities such as system design consulting, developing software tools, platforms and frameworks, capacity management and launch reviewsMaintain services once they are live by measuring and monitoring availability, latency and overall system healthScale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocityPractice sustainable incident response and blameless postmortemsBe part of an on call rotation to support production systemsWhat we need to see:BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience15+ years of experience with Infrastructure automation, distributed systems design, experience with design, develop tools for running large scale private or public cloud system in Production8+ years experience delivering foundational infrastructure and observability platforms.Experience in one or more of the following: Python, Go, Perl or RubyIn depth knowledge on Linux, Networking and ContainersWays to stand out from the crowd:Interest in crafting, analyzing and fixing large-scale distributed systemsSystematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive. Ability to debug and optimize code and automate routine tasksExperience in using or running large private and public cloud systems based on Kubernetes, OpenStack and Docker. Experience running Grafana, OpenTelemetry, Prometheus, and similar observability focused toolsYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 4, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Systems Software Engineer - Observability and Telemetry Platform in Santa Clara, CA vacancy
  • $195k - $290k

     ...most advanced AI-native platform. We work on large scale distributed systems, processing almost 3...  ...endpoint sensor that observes system activity,...  ...behavior, and streams telemetry to the Falcon cloud for...  ...this possible.This is a Principal Software Engineer role and the senior... 
    Platform
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    1 day ago
  • $249k

     ...for travelers everywhere.Principal Software Engineer, Observability Introduction to the Team:...  ...employees. A singular technology platform powered by data and...  ...in cloud, distributed systems, and observability. You will...  ...Architect and Build Core Telemetry Pipelines: Lead the... 
    Platform
    Full time

    Expedia

    San Jose, CA
    7 hours ago
  • $272k - $431.25k

     ...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team...  ...deterministic compute architecture into a platform that compiler teams and data...  ...RPC frameworks, coordination and telemetry patterns, MPI. Inference systems... 
    Platform
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU accelerated data processing has moved from...  ...boundaries and geographiesFamiliarity with the open source data platform ecosystem (Apache Spark, Velox, Presto, Apache Arrow,... 
    Platform
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    7 hours ago
  •  ...library. He/she will participate in the core system design and development. Our target...  ...proven records on infrastructure level software development experience ~2+ years Clojure...  ...communities ~ Linux (Debian in particular) platform programming experience is a big plus.... 
    Platform
    Full time

    Integrated Resources Inc.

    Santa Clara, CA
    11 hours ago
  • $184k - $287.5k

     ...Infrastructure organization is seeking a Senior System Software Engineer to lead the evolution of our next-generation Data & Observability Platform. We serve and collaborate directly with...  ...engineers rely on to visualize chip telemetry, debug distributed pipelines, and ensure... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $200k - $220k

     ...passionate, and committed engineers, technologists,...  ...AI Network Software Solution...  ...automation, network observability on a scale. You will...  ...most demanding AI platforms.This role will be...  ..., and monitoring.System Performance & ObservabilityImplement...  ...telemetry pipelines and... 
    Platform
    Worldwide

    Supermicro

    San Jose, CA
    2 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation innovations in automotive platform software, system architecture, and performance engineering. In this highly visible technical leadership role, you will contribute directly... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $272k - $431.25k

     ...Center MODS organization seeks a Principal Engineer to architect and scale next-...  ...L10 and L11 diagnostic systems for Cloud Service Providers...  ...distributed systems and hardware / software interfaces is essential for...  ..., HMC, BMC protocols and platform security.Consistent track... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...right place. As a Principal Engineer, Developer Platform Engineering at JPMorganChase...  ...extreme loadBuild observability-driven capacity...  ...with production telemetry, real user monitoring...  ...tools within the Software Development Life Cycle...  ...and payment system performance, such as... 
    Platform
    Early shift

    JP Morgan Chase

    Palo Alto, CA
    7 hours ago
  • $272k - $431.25k

     ...We're looking for a Principal Software Engineer to join our CSP...  ...customers to ensure NVIDIA platforms achieve target MTBI...  ...their fleet telemetry and failure data into...  ...you to distinguish systemic architectural gaps...  ...level telemetry and observability systems: time-series... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...and hands-on delivery across system software, drivers, and CUDA to make...  ...-mode components, driver/platform layers, and performance counter...  ...technical direction for an engineering team; mentor engineers,...  ...drivers with strict reliability, observability, and performance... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements...  ...firmware and GPU system software, working...  ...manageability, observability, security requirements...  ...firmware, or accelerator platform engineering. BS or...  ...monitoring and telemetry: Xid errors, thermal... 
    Platform
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $165.6k - $296.4k

     ...than 25%Profession: Software...  ...experimentation systems that scale across...  ...designing foundational engineering systems (instrumentation...  ...confidence. As a Principal Growth Engineer in...  ...quality, better telemetry, safer rollouts,...  ...CoreAI CoreAI - Platform and Tools is Microsoft... 
    Platform
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Mountain View, CA
    3 days ago
  • $272k - $431.25k

     ...assistants and engineering-productivity...  ...Now we need a principal-level, hands-on...  ...harden production systems and the...  ...performance, observability, release confidence...  ...like mature software, not prototypes...  ...evaluation frameworks, telemetry, and policy-...  ..., and related platform capabilities... 
    Platform
    Full time
    Live in

    Nvidia

    Santa Clara, CA
    1 day ago
  • $114.6k - $234.6k

     ...Infrastructure (OCI) is looking for a Principal Software Engineer to lead the development...  ...secure infrastructure systems that underpin the core of OCI’s compute platform. This role sits within...  ...firmware updates and telemetry-based observability. You’ll drive solutions... 
    Platform
    Temporary work
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    3 days ago
  • $231.4k - $331.8k

     ...Team You will join Cisco’s Platform & Identity Engineering Group, a foundational...  ...to provide the underlying systems that support Cisco’s innovation...  ...direction of Cisco’s software and technology solutions,...  ...services such as observability, data platforms, or AI/ML... 
    Platform
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    1 day ago
  •  ...experience as a senior software engineer, including experience...  ...complex, large-scale systems Expert-level...  ...least one major cloud platform (AWS, GCP, or Azure);...  ...-enabled impact.As a Principal Software Engineer, you...  ...improving reliability, observability, and scalability of core... 
    Platform
    Apprenticeship

    McKinsey & Company

    San Jose, CA
    7 hours ago
  • $185.5k - $265k

     ...Zero Trust Exchange platform. This innovation...  ...leverage intelligent systems to stay ahead of evolving...  ...are looking for a Principal Software Development Engineer to join our team....  ...are resilient, observable, and easy to evolve...  ...intelligent microservice telemetry to optimize high-... 
    Platform
    Full time
    Work at office
    Local area
    Remote work
    3 days per week

    Zscaler

    San Jose, CA
    1 day ago
  • $118.3k - $224.9k

     ...Technology (AST) is seeking a Principal Software Engineer with specialization in...  ...patches, and configurations.System migrations and cross-...  ...Terraform, Ansible, Puppet, Salt)Observability dashboard tools (Datadog,...  ...software developmentCross-platform development, focusing on... 
    Platform
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Raytheon

    San Jose, CA
    7 hours ago
  • $208k - $260k

     ...Place to Work, we offer a deep observability pipeline that efficiently...  ...educational organizations. As a Principal Software Engineer, you will lead the design and...  ...will work across distributed systems, backend services, and modern cloud platforms to deliver scalable, secure,... 
    Platform
    Local area
    Worldwide
    3 days per week

    Gigamon

    Santa Clara, CA
    1 day ago
  • $190k - $210k

     ...StatesProducts - Engineering /Fulltime /...  ...detailsTitle of position: Principal Software EngineerPosition...  ...Director of SW Systems...  ...systems, cloud platforms, security, and real...  ...secure, reliable, observable, adaptable, and...  ...systems, network telemetry, or large-scale... 
    Platform
    Full time
    H1b
    Local area
    Work from home
    Work visa
    Shift work

    Extreme Networks, Inc.

    San Jose, CA
    1 day ago
  • $224k - $356.5k

     ...multi-GPU, high-bandwidth architecture of this platform.We are looking for a deeply technical systems software engineer who will own AI stack readiness on DGX Station...  .... Ability to read GPU traces and translate observations into actionable optimizations.Strong understanding... 
    Platform
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $175k - $275k

    Cerebras Systems builds the world's largest...  ...of the Embedded Software team, you will help...  ...Wafer Scale Engine (WSE)—the world’s...  ...services, and the Linux platform and BSP layers...  ...reliability, observability, and long-term maintainability...  ...hardware level telemetry pipelines. The... 
    Platform

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $83k - $166.1k

    Implements features in system software modules and automation pipelines...  ...to standards. Supports platform bring-up tasks and assists...  ...support efforts with platform engineering, operations, firmware development...  ...logging, monitoring, and observability for effective debugging,... 
    Platform
    Temporary work
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    7 hours ago
  • $248k - $391k

     ...world. We are looking for a Principal Software Engineer to drive the evolution of...  ...learning loops.Own system design reviews, code reviews...  ...collaboration, and productivity platforms to create cohesive,...  ...CI/CD).Instrument robust observability, tracing, structured logging... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...production workflows? Our Engineering Workflow Platform team builds the user-facing...  ...with hands-on systems engineering to solve challenges...  ...migrations, and improve workflow observability.What we need to see:...  ...least 12 years of relevant software engineering experience.Production... 
    Platform
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $224k - $356.5k

     ...Do you look at a build system, a release pipeline, or...  ...re looking for a senior engineer to invent, set, and...  ...enables us to create the software and firmware that drive...  ...signing services, and the observability that proves it all works. When our platforms are fast and... 
    Platform
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    1 day ago
  • $200k - $322k

     ...looking for a highly skilled Senior Software Engineer to design and develop AIOps & Observability platforms at NVIDIA. The platforms are...  ..., reliable, and distributed systems that can handle high traffic...  ...aggregating, and visualizing telemetry and operational metrics.Ways To... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $147k - $237.5k

     ...TeamEngineering - Our engineering team is at the...  ...CareerAs a Principal Engineer on...  ...of our platform. You’ll think...  ...broadly about all system components, weigh...  ...scalable software features and infrastructure...  ...monitoring, observability, and alerting...  ...over security telemetry, configuration... 
    Platform
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Systems Software Engineer - Observability and Telemetry Platform. Be the first to apply!