Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC Architect, Automation and At-Scale Deployment

$184k - $287.5k

NVIDIA

NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can tackle, and that matter to the world. This is our life’s work, to amplify human imagination and intelligence. Make the choice, join our diverse team today!We are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. Be a key player to enable the most exciting computing hardware and software and contribute to the latest breakthroughs in artificial intelligence and GPU computing. Provide insights on and implement at-scale system administration and tuning mechanisms for large-scale compute runs. You will work with the latest accelerated computing and Deep Learning software and hardware platforms, and with many scientific researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.What you’ll be doing:Provide engineering solutions to operationalize the latest GPU Computing products and software stacks, ensure technical relationships with internal and external engineering teams, and assisting systems, machine learning/deep learning engineers in building creative solutions based on NVIDIA technology.Be an internal reference for system administration, at-scale system analysis, and other datacenter and large-scale GPU-accelerated system solutions among the NVIDIA technical community.What we need to see:8+ years of experience using in accelerated computing for datacenter/HPC-based Enterprise computing solutions.Solid understanding of accelerated computing scheduling and I/O stacks.C/C++/Python/Bash programming/scripting experience.Experience working with engineering or academic research community supporting high performance computing or deep learning.Experience with parallel filesystems.Strong teamwork and communication skills, both verbal and written.Ability to multitask effectively in a dynamic environment.Action driven with strong analytical skills.Desire to be involved in multiple diverse and innovative projects.BS (or equivalent experience) in Engineering, Mathematics, Physics, or Computer Science. MS or PhD desirable.Ways to stand out from the crowd:Deep Learning framework skills.Exposure to using and deploying telemetry and visualization pipelinesExposure to container technology and Linux performance tools.Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 3, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Remote; US, CO, Remote; US, WA, Remote; US, CA, Remote; US, Remote; US, AR, Remote; US, NC, Remote; US, IL, Remote; US, OR, Remote; US, UT, Remote; US, NM, Remote; US, MA, RemoteType: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior HPC Architect, Automation and At-Scale Deployment in Santa Clara, CA vacancy
  • $224k - $356.5k

     ...high-performance computing (HPC), cloud service providers (CSP...  ...Analyze and debug performance scaling bottlenecks on multi-core and...  ...with CPU and interconnect architects to improve future CPU and system...  ..., and scalable cloud deployments.Your base salary will be determined... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...world.As an AI Storage Platform Architect at NVIDIA, this position will...  ...platforms and real-world AI deployments - translating the...  ...aligned with NVIDIA Dynamo), large-scale foundation model training, and...  ...architecting datacenter-scale AI, HPC, or storage infrastructure as... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    10 hours ago
  • $208k - $327.75k

     ...connects, and accelerates workloads at scale, and we’re seeking a visionary Product Architect with strong expertise in systems...  ...interoperability, and datacenter deployment requirements Lead...  ...experience architecting datacenter-scale HPC or AI infrastructure as a... 
    Senior

    NVIDIA

    Santa Clara, CA
    4 days ago
  •  ...advance your career. THE ROLE: We are seeking a Robotics AI Architect to define and scale next-generation Physical AI systems, with a focus on...  ...authority, you will synthesize learnings from real-world deployments and translate them into platform-defining capabilities, shaping... 
    Senior

    AMD

    San Jose, CA
    3 days ago
  • $152k - $241.5k

    We are now looking for a Senior Performance Architect for Nemotron! At NVIDIA, we are redefining the future...  ...choices translate into real-world deployment efficiency. You will ensure that...  ...architecture and system efficiency at scale. This role sits at the center of Generative... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $208k - $327.75k

     ...defined architectures.We are looking for a Senior AI Architect to help define the next generation of...  ...validate scalability, efficiency, and deployment feasibility.Partner closely with AI...  ...in modern AI architectures and large-scale model systemsExperience mapping AI workloads... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    10 hours ago
  • $184k - $287.5k

     ...reality. We are seeking visionary computer architects to design and develop models for the...  ...of cost and power modeling of datacenter-scale accelerated computing platforms Evangelizing...  ...server architecture design and deployment, especially up to hyperscale. Background... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $120k - $243k

    Senior AI Driven Test Automation ArchitectThis role has been designed as 'Hybrid' with a requirement that...  ...architecture and evolution of enterprise-scale automation platforms spanning UI,...  ...across the SDLC.Responsibilities:Architect and scale enterprise-grade automation... 
    Senior
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start
    Shift work
    2 days per week

    Hewlett Packard Enterprise

    San Jose, CA
    2 days ago
  •  ...organization is transforming the AI and HPC landscape. Our mission is to...  ...Principal Modeling Architect to join the Product Architecture...  ...networks), datatypes, and scaling methodologies to anticipate future...  ...customers on workload deployment and optimization.This role is... 
    Remote work

    AMD

    San Jose, CA
    4 days ago
  • $184k - $287.5k

     ...every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers...  ...Design and implement computationally performant features for large scale, CUDA-backed ML training frameworks, using low level acceleration... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...of a curious and motivated Senior Solutions Architect to join our NVIDIA Infrastructure...  ...apply observability and automation to improve our benchmarking...  ...throughput, latency, and scaling efficiency.Work closely...  ...managing Linux-based systems in HPC, distributed systems, or AI... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $255k - $340k

     ...responsible for building and scaling the physical infrastructure powering...  ...validation for new HPC AI/ML, general purpose compute...  ...throughout hardware NPI and deployment readiness.Development and execution...  ...new hardware systems. Enable automated tools for doing scale testing... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    10 hours ago
  • $152k - $241.5k

     ...seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU Compute Clusters for EDA (Electronic Design Automation) and high-performance computing...  ...strategic guidance for managing large-scale HPC systems, including the deployment... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $255k - $340k

     ...responsible for building and scaling the physical...  ...integrating OEM and white-label HPC AI/ML, general purpose...  ....Partner with HPC architects to translate platform...  ...fleet engineering, deployment, operation and datacenter...  ...configuration and automation.Background in performance... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    9 hours ago
  • $152k - $241.5k

     ...highly motivated Senior Software Engineers...  ...focus on NVLink Rack-Scale Systems Stability...  ...closely with architects and developers building...  ...datacenter deployments. What you will be...  ...tools, diagnostics, automation, and infrastructure...  ...and large-scale AI/HPC clusters such as NVIDIA... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    10 hours ago
  • $200k - $275k

     ...world, and help drive advancements in automation and robotics, mobility, healthcare, energy...  ...and X.Principal AI Accelerator Architect Lead Boston, MA; San Jose, CA Team: MicroAI...  ...between academic research and real-world deployment, ensuring ideas withstand contact with... 
    Permanent employment
    Full time
    Work at office
    Day shift

    Analog Devices

    San Jose, CA
    1 day ago
  • $272k - $431.25k

     ...technical focal point for fleet-scale reliability, working directly...  ...work (health monitoring deployment, telemetry pipeline, alerting...  ...failure models for accelerator or HPC infrastructureBackground in...  ...fleet-level dashboards and automated remediationNVIDIA is leading... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...network fabrics for OCI's largest AI and HPC customers. These fabrics are the...  ...on our team supports the design, deployment, and operations of a large-scale global Oracle cloud computing...  ...of a deep network understanding and automation skills to operate a production environment... 

    Oracle

    Santa Clara, CA
    3 days ago
  •  ...Headquartered in Singapore, Bitdeer has deployed data centers across multiple...  ...and hands-on Cloud SRE Architect to lead the design,...  ...platform solutions that power large-scale AI and enterprise workloads....  ...partitioning, optical health), HPC storage (Lustre, NetApp/Pure/DDN... 
    Senior
    Full time
    Contract work
    Local area
    Shift work

    Bitdeer Technologies Group

    San Jose, CA
    5 days ago
  •  ...driven data economy. As AI scales, so does data. Every...  ...and OEM partners. As a Storage Architect on this team, you own the understanding...  ...executive team. This is a senior individual contributor role...  ...Translate customer pain points and deployment challenges into actionable... 
    Senior
    Temporary work
    Immediate start
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Western Digital

    San Jose, CA
    18 days ago
  • $272k - $431.25k

     ...agentic workloads, deep learning (DL), high-performance computing (HPC), cloud service providers (CSP), gaming, virtual reality, and...  ...use-cases, efficient data processing, and scalable cloud deployments.Your base salary will be determined based on your location, experience... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...enable massive model training at scale. Your expertise will drive 2-...  ...single-GPU to thousand-GPU deployments.THE PERSON:You have a deep...  ...REQUIRED EXPERIENCE:Extensive and Senior experience optimizing large-...  ...track record architecting distributed training systems... 
    Remote work

    AMD

    Santa Clara, CA
    1 day ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica -...  ...excellence of Sumo’s planet-scale observability and security products...  ...and design, through deployment, operation, and refinement.Participate...  ...managing SLOsWrite code and automation to reduce operational... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    10 hours ago
  • $168k - $258.75k

     ...We are seeking an innovative Senior Customer Success Operations Manager – Technology to help scale NVIDIA's Customer Success operating...  ...workflows, prototypes, automations, and other digital assets where...  ...and other teams to design and deploy Customer Success workflows and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $262k - $365k

     ...of services—from inception and design, through to deployment, operation and refinement.Support services before...  ...availability, latency and overall system health.Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve... 
    Senior

    Google

    San Jose, CA
    10 hours ago
  • $248k - $396.75k

     ...building a solid foundation with automation. We are looking for a...  ...interfaces, addressing, and deployment intent.Generate intent-based...  ...to design and deliver large-scale network automation using IaC...  ...GPU cluster networking, and HPC environments.Cloud and hybrid... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...obstacles, catching issues early, and turning deployments into a strength rather than a liability...  ...CD rollouts, StackStorm event-driven automation, HashiCorp Vault for credential storage...  ...operators in Go for scheduling, auto-scaling, and compliance across data centers... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • $136k - $218.5k

    NVIDIA is looking for a Senior CPU Tooling and Design Automation Engineer! Do you want to help...  ...-performance computing (HPC), cloud service providers...  ..., working closely with architects, chip leads and designersPush...  ...-team methodologies for deployment of this toolingWhat We... 
    Senior
    Full time
    Work experience placement
    Night shift

    Nvidia

    Santa Clara, CA
    1 day ago
  • $174k - $253k

     ...of services—from inception and design, through to deployment, operation and refinement.Support services before...  ...availability, latency and overall system health.Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve... 
    Senior

    Google

    Sunnyvale, CA
    10 hours ago
  •  ...network infrastructure challenges, develop automation to improve operational efficiency, and...  ...complex network incidents across large-scale cloud environments; designing and developing...  ...issues; and participating in network deployment, expansion, and upgrade projects.... 
    Senior

    Oracle

    Santa Clara, CA
    1 hour ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC Architect, Automation and At-Scale Deployment. Be the first to apply!