Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Systems Software Engineer

Oracle

Principal Systems Software EngineerOracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems software and platform-management capabilities that power next-generation GPU infrastructure.This is a hands-on systems engineering role at the intersection of software, firmware, hardware, and large-scale cloud infrastructure. You will work on complex GPU and server platforms, owning critical capabilities spanning BMC/service-processor software, platform management, firmware lifecycle, reliability and serviceability, telemetry, power management, hardware bring-up, and fleet operations.You will work closely with silicon, firmware, hardware, compute, and fleet engineering teams, as well as technology and manufacturing partners, to bring new platforms from initial hardware enablement through production deployment and ongoing operation at cloud scale.This role is well suited for an engineer with deep experience in computer systems, embedded or platform firmware, and hardware/software integration who enjoys solving difficult problems that cross traditional engineering boundaries.ResponsibilitiesKey Responsibilities:Design, develop, and maintain complex BMC, service-processor, and platform-management software for GPU and server infrastructure.Develop capabilities for host management, baseboard management, telemetry, platform monitoring, power control, and system diagnostics.Design and implement secure, reliable firmware-management and update workflows across complex, multi-vendor platforms.Define robust interfaces and integration contracts between platform firmware, hardware components, operating systems, drivers, and higher-level infrastructure services.Systems & Firmware DevelopmentDesign and implement low-level systems software and firmware using technologies such as C, C++, Python, and Bash.Develop maintainable software for managing, monitoring, diagnosing, and provisioning server and GPU systems.Build automation and tooling that improves platform provisioning, onboarding, validation, diagnostics, and fleet operations.Develop and debug advanced platform capabilities involving areas such as RAS, telemetry, power management, high-speed I/O, and chipset/SoC services.Conduct design and code reviews and help establish strong engineering practices for maintainability, testing, observability, security, and reliability.Hardware Bring-Up & Cross-Layer DebuggingPlay a leading role in initial board, device, and platform bring-up for new GPU and server systems.Diagnose difficult failures spanning hardware, firmware, bootloaders, operating systems, drivers, and platform services.Use hardware and software diagnostic techniques—including logs, schematics, JTAG, logic analyzers, emulators, and platform instrumentation—to isolate root causes.Partner with silicon, board, firmware, and manufacturing teams to validate end-to-end platform behavior and resolve integration issues.Turn complex or recurring failures into durable engineering fixes, improved diagnostics, automation, and preventive controls.GPU Reliability, Serviceability & OperationsDevelop reliability, availability, and serviceability (RAS) capabilities for large-scale GPU and server environments.Improve telemetry, fault detection, logging, observability, and diagnostics used to identify and resolve platform issues.Develop and improve power-control and power-capping capabilities and related platform instrumentation.Design systems with fleet-scale reliability, fault tolerance, secure firmware lifecycle, and operational serviceability in mind.Support difficult platform incidents and escalations and help translate field findings into long-term product and engineering improvements.Technical LeadershipOwn technically complex and sometimes ambiguous areas from architecture and design through implementation, validation, and deployment.Drive technical decisions and establish clear interfaces across teams responsible for different layers of the platform.Lead deep technical investigations and help teams reach evidence-based root causes for difficult system failures.Raise engineering standards through architecture and design reviews, code reviews, testing practices, automation, and diagnostic tooling.Mentor and provide technical guidance to engineers while remaining actively involved in design, coding, bring-up, and debugging.Collaborate effectively across software, firmware, hardware, silicon, compute, fleet, support, and external partner organizations.In addition, you have:6+ years of programming and/or scripting experience, with relevant languages such as C, C++, Python, or Bash.Strong systems-software, embedded-software, or firmware development experience.Experience integrating software or firmware with complex hardware systems.Strong understanding of computer hardware fundamentals and the interaction between hardware, firmware, operating systems, and software.Demonstrated ability to diagnose complex problems that cross hardware and software boundaries.Experience with software development practices including design, implementation, debugging, code review, testing, automation, and quality assurance.Experience working in Linux/Unix-based development or systems environments.Ability to independently own technically complex projects and collaborate across multiple engineering organizations.Preferred Qualifications:BMC, OpenBMC, service processors, or server platform-management technologies.Server, GPU, accelerator, or other complex compute-platform firmware.GPU or server RAS, telemetry, fault management, observability, or serviceability.Board, device, or system bring-up.Hardware debugging using JTAG, logic analyzers, emulators, schematics, or related diagnostic tools.Firmware lifecycle management and secure firmware-update mechanisms.Power management, power control/capping, thermal management, or platform telemetry.Hardware interfaces and low-level communication protocols.CPU, GPU, SoC, ASIC, or FPGA-based systems.ARM, AMD, Intel, NVIDIA, or similarly complex compute platforms.Automation and diagnostics for server provisioning, validation, or fleet operations.Large-scale cloud or data-center infrastructure.Technical leadership, mentoring, architecture/design ownership, and cross-functional engineering coordination.Location: On-Site | Santa Clara, CA

Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the Principal Systems Software Engineer in Santa Clara, CA vacancy
  •  ...testing framework and library. He/she will participate in the core system design and development. Our target system is based on jepsen...  ...~5+ years proven records on infrastructure level software development experience ~2+ years Clojure development experience... 
    Suggested
    Full time

    Integrated Resources Inc.

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team builds the foundational software that turns a novel deterministic compute architecture into a platform that compiler teams and data center operators can rely... 
    Suggested
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    7 hours ago
  •  ...transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what...  ...Santa Clara, CA, headquarters 3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires... 
    Suggested
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU accelerated data processing has moved from proof of concept to production deployments. Enterprises now recognize the need for accelerated computing to handle large... 
    Suggested
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $146.3k - $306.4k

    Defines architecture for large-scale systems software, firmware integration, and fleet automation that impacts multiple teams or products...  ...reliability, performance, and cost at hyperscale. Champions engineering excellence: coding standards, threat modeling, resource management... 
    Suggested
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation innovations in automotive platform software, system architecture, and performance engineering. In this highly visible technical leadership role, you will contribute directly... 
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • NVIDIA seeks a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group to drive GPU-accelerated data processing in production deployments. You'll develop Java, Scala, and CUDA/C++ libraries to accelerate Spark DataFrames and I/O on common formats... 

    NVIDIA

    Santa Clara, CA
    3 days ago
  •  ...Defines architecture for large-scale systems software, firmware integration, and fleet automation that impacts multiple teams or products...  ...reliability, performance, and cost at hyperscale. Champions engineering excellence: coding standards, threat modeling, resource management... 
    Full time
    Flexible hours

    Oracle

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

    NVIDIA is seeking a Principal Software Engineer for LPX System Software to build foundational software for a novel computing architecture in Santa Clara, California. The role involves shaping key system software components in Rust and driving architecture design while... 

    NVIDIA

    Santa Clara, CA
    5 days ago
  •  ...Oracle Cloud Infrastructure seeks a Principal Systems Software Engineer to build and evolve low-level systems software for next‑gen GPU infrastructure. This is a hands-on role spanning software, firmware, hardware and cloud-scale fleet operations. You’ll work across... 

    Socket.dev

    Santa Clara, CA
    11 hours ago
  • $272k - $431.25k

     ...!The Data Center MODS organization seeks a Principal Engineer to architect and scale next-generation L10 and L11 diagnostic systems for Cloud Service Providers (CSPs). In this...  ...Proficiency in distributed systems and hardware / software interfaces is essential for success.What... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems software and platform-management capabilities that power next-generation GPU infrastructure. This is a hands-on systems engineering role... 
    Flexible hours

    Socket.dev

    Santa Clara, CA
    11 hours ago
  • $200k - $220k

     ...seek talented, passionate, and committed engineers, technologists, and business leaders to...  ...Supermicro is seeking an experienced AI Network Software Solution Architect to lead the design...  ..., configuration, and monitoring.System Performance & ObservabilityImplement telemetry... 
    Worldwide

    Supermicro

    San Jose, CA
    2 days ago
  • $140k - $190k

    WiFi team is looking for a Principal Embedded Software Engineer with C programming and networking knowledge to join our team. This is a great opportunity...  ...skills.Strong knowledge or experience with network and system designPassion and talent on Linux Kernel and application... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    2 days ago
  • Bachelor's or Master's degree in Computer science, Software Engineering, or related field. 10+ years of experience in Android mobile app development, with at least 3+ years in medical device application. Experience of architecting, designing and developing of mobile applications... 
    Full time

    Boston Scientific

    San Jose, CA
    1 day ago
  •  ...motivated and technically exceptional PMTS AI/ML Compiler Engineer to join AMD's AI Software organization. In this role, you will drive compiler and...  ...profiling and optimizing performance-critical software systems.Proficiency with Linux development environments, source... 

    AMD

    San Jose, CA
    4 days ago
  • $272k - $431.25k

    We're looking for a Principal Engineer to join our CSP Engagements team as the technical focal...  ...enables you to identify patterns and drive systemic improvements in documentation,...  ...pattern analysis — identify configuration, software, or workload differences that explain... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $110k - $220k

     ...generation content. What you’ll do:Guide and mentor, a team of engineers, conducting code reviews and leading design discussions to...  ...goals and scalability requirements.  Architect complex software systems, ensuring performance, security, and scalability needs are met... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    7 hours ago
  • $143k - $286k

     ...What you'll do...Job DescriptionPrincipal, Software EngineerWe are seeking a talented and passionate Principal, Software Engineer to join our International Technology Organization...  ....   Architect complex software systems, ensuring performance, security, and scalability... 
    Full time
    Temporary work
    Part time
    Work at office
    Flexible hours

    Walmart

    Sunnyvale, CA
    2 days ago
  • $272k - $431.25k

     ...heterogeneous clusters so that many accelerators feel like a single system at datacenter scale. As large language models rapidly...  ...deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap for memory management of... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $231.4k - $331.8k

     ...You will join Cisco’s Platform & Identity Engineering Group, a foundational group responsible...  ...groups to provide the underlying systems that support Cisco’s innovation. The work...  ...for the technical direction of Cisco’s software and technology solutions, influencing industry... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    1 day ago
  • $185.5k - $265k

     ...thrive in an environment where we leverage intelligent systems to stay ahead of evolving threats. We believe in transparency...  ...the future of cybersecurity.RoleWe are looking for a Principal Software Development Engineer to join our team. This is a hybrid (3 days/week) or... 
    Full time
    Work at office
    Local area
    Remote work
    3 days per week

    Zscaler

    San Jose, CA
    1 day ago
  • $272k - $431.25k

     ...You will lead the architecture and hands-on delivery across system software, drivers, and CUDA to make profiling continuously available...  ...signals into actionable insights.Set technical direction for an engineering team; mentor engineers, drive technical planning to mitigate... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration,...  ...future of AI and beyond. Together, we advance your career. PRINCIPAL SOFTWARE DEVELOPMENT ENGINEER - PYTORCH TRAINING FRAMEWORKS THE ROLE: AMD is looking... 

    AMD

    San Jose, CA
    2 days ago
  • $272k - $431.25k

     ...smart personal assistants and engineering-productivity tools to data-...  ...the company. Now we need a principal-level, hands-on engineering...  ...depth to harden production systems and the architectural vision...  ...applications behave like mature software, not prototypes.Build... 
    Full time
    Live in

    Nvidia

    Santa Clara, CA
    1 day ago
  • $240.8k - $321.5k

     ...communities to create economic opportunity for all.Role SummaryThe Principal Engineer in the Identity Domain provides senior technical leadership...  ...for core identity capabilitiesReview and influence system designs and code to ensure security, scalability, and correctnessAuthentication... 
    Immediate start

    eBay

    San Jose, CA
    1 day ago
  • $221.2k - $387.1k

     ...DescriptionIt all started when engineer Fred Luddy wrote code that...  ...people.Job DescriptionPrincipal Software EngineerThe engineering...  ...get to do in this role: As a Principal Software Engineer, you will architect...  ...span large-scale distributed systems, data ingestion, context... 
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    7 hours ago
  • $155.8k - $224.2k

     ...pressing challenges of the 21st century.We are looking for a Principal Software Engineer, Edge Compute to join our team in one of today’s most...  ...Role and Responsibilities:Design, implement, and optimize system‑level and backend services in Rust or Go.Debug complex issues... 
    Full time
    Work at office
    Worldwide

    Bloom Energy

    San Jose, CA
    4 days ago
  •  ...years of professional experience as a senior software engineer, including experience with complex, large-scale systems Expert-level proficiency in Python, Golang, React...  ...distinctive, tech-enabled impact.As a Principal Software Engineer, you will lead the design and... 
    Apprenticeship

    McKinsey & Company

    San Jose, CA
    7 hours ago
  • $143k - $286k

     ...Engagement Services Technology Engineering TeamThe Customer Engagement...  ...world-class experiences and software solutions for Contact Centers and our customers. As a Principal, Software Engineer, you will...  ...focus on Java & Javascript based systems including React, React Native... 
    Full time
    Temporary work
    Part time
    Worldwide

    Walmart

    Sunnyvale, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Systems Software Engineer. Be the first to apply!