Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Solutions Architect

$170k - $190k

ePlus Technology, inc.

We are seeking an elite Solutions Architect to lead the end-to-end design, sizing, and deployment of NVIDIA AI Factory-aligned infrastructure. In this highly technical, customer-facing role you will translate complex AI and machine learning workload requirements into fully engineered infrastructure solutions spanning colocation facilities, GPU compute, high-performance networking, parallel storage, and the complete NVIDIA AI software stack.

You will serve as a trusted technical advisor to enterprise and hyperscale customers, partnering with sales, product, and engineering teams to win and deliver transformational AI infrastructure programs. Your expertise will directly shape how organizations build and operate production AI Factories capable of training frontier models, running large-scale inference fleets, and accelerating data science pipelines at scale.

Your Impact

  • Lead discovery workshops to capture AI/ML workload requirements, including model training scale, inference SLAs, data pipeline throughput, and multi-tenancy needs.
  • Architect full-stack AI Factory solutions aligned to NVIDIA reference architectures, integrating colocation, GPU compute, networking, storage, and software layers.
  • Develop detailed Bills of Materials (BOMs), rack elevation diagrams, network topology drawings, and power/cooling budgets for customer proposals.
  • Define GPU cluster architectures using NVIDIA DGX, HGX, and MGX systems with B200, B300, and GB300 Blackwell SXM and NVLink-Switch configurations.
  • Design RTX PRO 6000 Blackwell Server Edition deployments for inference-optimized and enterprise AI workloads.
  • Conduct workload sizing and TCO/ROI modeling to validate infrastructure dimensioning for training, finetuning, and inference at scale.

Colocation & Facility Planning

  • Specify colocation requirements including critical power load (MW-scale), UPS and generator configurations, and PUE targets.
  • Design high-density GPU deployments utilizing air-cooled, direct liquid cooling (DLC), and rear-door heat exchanger configurations.
  • Define meet-me room (MMR) and cross-connect requirements; specify carrier-neutral telecom diversity strategies.
  • Engage colocation providers and data center operators to validate capacity availability and negotiate technical SLAs.
  • Coordinate with facilities and MEP engineers to validate power infrastructure from utility feed through PDU to rack level.

GPU Compute Infrastructure

  • Architect multi-node GPU clusters optimized for large language model (LLM) pre-training, fine-tuning, and reinforcement learning from human feedback (RLHF).
  • Size and configure DGX SuperPOD, HGX H/B-series, and MGX modular systems based on model parameter count, dataset size, and iteration timelines.
  • Define server firmware, BIOS, BMC, and DGXOS baselines for production GPU infrastructure.
  • Establish GPU health monitoring, RAS (Reliability, Availability, Serviceability) policies, and lifecycle management procedures.

High-Performance Networking

  • Design backend GPU fabric networks using NVIDIA Quantum InfiniBand (NDR 400Gb/s and HDR 200Gb/s) for distributed training traffic.
  • Architect Spectrum-X Ethernet-based AI networking solutions for inference clusters requiring highbandwidth, low-latency connectivity.
  • Specify ConnectX-8/7 HCA deployments and configure RDMA over Converged Ethernet (RoCEv2) or InfiniBand transport for NCCL collective operations.
  • Integrate BlueField-3 DPUs for GPU-accelerated network functions, storage offload, zero-trust security isolation, and bare-metal provisioning.
  • Design leaf-spine and fat-tree topologies for non-blocking bisectional bandwidth in GPU training clusters.
  • Define Quality of Service (QoS) policies separating storage, compute fabric, and management plane traffic.
  • Design high-performance parallel file system solutions using VAST Data, Hammerspace, and Pure Storage FlashBlade//E for AI training and checkpoint storage.
  • Size storage capacity, IOPS, and throughput based on dataset characteristics, checkpoint frequency, and concurrent reader/writer counts.
  • Architect multi-tier storage hierarchies: hot NVMe flash (VAST/FlashBlade) for active datasets, warm object storage for model archives, and cold tape/cloud for long-term retention.
  • Configure VAST Data Universal Storage for disaggregated storage with NFS, S3, and POSIX access; tune for large sequential read performance.
  • Deploy Hammerspace Global Data Environment for distributed data management and NFS-over-RDMA acceleration across geographically dispersed GPU clusters.
  • Define data pipeline architectures ingesting from cloud object stores (S3, GCS, ABS) to local flash for GPU local data loading without I/O bottlenecks.

AI Software Stack & Orchestration

  • Deploy and configure NVIDIA AI Enterprise (NVAIE) software stack including NVIDIA GPU Operator, NIM microservices, and RAPIDS accelerated data science libraries.
  • Architect inference serving infrastructure using NVIDIA NIM (NVIDIA Inference Microservices) for optimized LLM and vision model deployment with autoscaling.
  • Implement NVIDIA Dynamo for distributed inference and disaggregated serving of large-scale generative AI models.
  • Configure and optimize CUDA toolkit, cuDNN, NCCL communication libraries, and custom kernel environments for training workloads.
  • Deploy Base Command Manager and DGXOS for cluster lifecycle management, node provisioning, health dashboards, and job scheduling integration.
  • Integrate NVIDIA Mission Control for AI Factory operations, observability, and multi-cluster fleet management.
  • Design and deploy Kubernetes-based AI platforms using NVIDIA GPU Operator, integrating with Run:ai for dynamic GPU resource scheduling and multi-tenant workload isolation.
  • Configure SLURM workload manager for traditional HPC-style job scheduling on bare-metal GPU clusters, including preemption policies, fair-share scheduling, and burst-to-cloud integration.
  • Establish MLOps toolchain integrations with popular frameworks (PyTorch, JAX, TensorFlow) and experiment tracking platforms (MLflow, Weights & Biases).
  • Serve as primary technical point of contact throughout the pre-sales and delivery lifecycle, from initial discovery through post-deployment optimization.
  • Produce and present architecture design documents, technical proposals, and executive-level briefings to CTO/CIO and VP-level stakeholders.
  • Lead proof-of-concept (POC) and pilot deployments, including benchmark design, execution, and results analysis.
  • Collaborate with procurement, logistics, and deployment teams to ensure on-time delivery of complex infrastructure programs.
  • Provide post-deployment hypercare support, performance tuning, and capacity planning advisory services.
  • Contribute to internal knowledge bases, solution playbooks, and reference architectures for repeatable AI Factory deployments.

Candidates must demonstrate deep, hands-on expertise across the following technology domains:

GPU Compute

DGX B200 / B300, DGX H100 / H200, HGX B200 / B300, HGX H100 / H200, MGX platforms, GB300 NVL72 / GB200 NVL72, RTX PRO 6000 Blackwell Server Edition, NVLink Switch System, NVLink-C2C

Networking

NVIDIA Quantum InfiniBand (NDR 400G, HDR 200G), Spectrum‑X Ethernet, ConnectX‑8 / ConnectX‑7 HCAs, BlueField‑3 DPU, SHARP in‑network computing, UFM Fabric Manager, RDMA / RoCEv2 / InfiniBand

Storage

VAST Data Universal Storage (NFS/S3/POSIX), Hammerspace Global Data Environment, Pure Storage FlashBlade//E (Evergreen//One), NFS-over-RDMA, parallel file systems (Lustre, GPFS/WEKA), S3‑compatible object storage

AI Software

NVIDIA AI Enterprise (NVAIE), NIM Microservices, RAPIDS (cuDF, cuML, cuGraph), NVIDIA Dynamo, CUDA Toolkit, cuDNN, NCCL, TensorRT, Triton Inference Server. Base Command Manager, DGXOS, NVIDIA Mission Control, DGX Cloud, UFM, IPMI / Redfish BMC management.

Orchestration

Run:ai, SLURM, etc.

Colocation

Design and capacity planning for data center environments.

Frameworks

PyTorch, JAX, TensorFlow, Hugging Face Transformers, DeepSpeed, Megatron‑LM, vLLM, LMDeploy

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related technical discipline; Master's degree preferred.
  • 8+ years of solutions architecture, systems engineering, or technical pre-sales experience, with at least 4 years focused on GPU infrastructure or HPC environments.
  • Proven track record designing and deploying NVIDIA DGX or HGX-based GPU clusters in production AI/ML environments.
  • Deep understanding of distributed deep learning concepts: tensor parallelism, pipeline parallelism, data parallelism, gradient checkpointing, and mixed‑precision training.
  • Hands‑on experience with InfiniBand or high-speed Ethernet fabric design, RDMA configuration, and collective communication tuning (NCCL, MPI).
  • Direct experience sizing and deploying parallel storage systems (VAST, Hammerspace, or Lustre/WEKA/GPFS) for AI training workloads.
  • Strong working knowledge of Kubernetes, GPU Operator, and at least one GPU workload scheduler (Run:ai or SLURM).
  • Experience with Linux system administration, CUDA development environment configuration, and GPU driver/firmware management.
  • Demonstrated ability to create compelling technical proposals, architecture diagrams (Visio/Lucidchart/draw.io), and BOM-level documentation.
  • Exceptional communication skills with proven ability to present to both deep technical audiences and C‑level executives.

Preferred Qualifications:

  • NVIDIA-certified professional credentials (DCA-Core, NCP-DS, or equivalent).
  • Experience with NVIDIA Base Command Platform or Mission Control for multi‑cluster AI Factory operations.
  • Familiarity with sovereign AI, government cloud, or regulated industry AI infrastructure requirements.
  • Experience integrating AI Factory infrastructure with public cloud (AWS, Azure, GCP) for hybrid and burst‑to‑cloud architectures.
  • Background in MLOps, LLMOps, or platform engineering for production AI model lifecycle management.
  • Prior experience with colocation data center procurement, RFP development, and SLA negotiation.
  • Contributions to open-source AI infrastructure projects or published technical content (blogs, whitepapers, conference presentations).
  • Active participation in the NVIDIA Partner Network (NPN) ecosystem or prior experience at an NVIDIA Elite Solution Provider.

Core Competencies

Technical Depth

End-to-end AI infrastructure expertise from silicon to software; ability to go deep on any layer of the stack.

Systems Thinking

Ability to reason holistically about performance, reliability, power, cost, and operability trade-offs across complex integrated systems.

Customer Obsession

Relentless focus on understanding customer AI objectives and delivering solutions that accelerate time-to-value.

Executive Presence

Confidence and clarity when presenting complex technical architectures to senior business and technology leaders.

Analytical Rigor

Data-driven approach to workload sizing, performance modeling, and TCO analysis with attention to detail.

Ability to lead cross-functional pursuit teams, align internal stakeholders, and orchestrate complex delivery programs.

Position Specifics

The initial base salary range for this position is expected to be between $170,000 and $190,000 annually. The final base salary offered will be determined by multiple factors, including, but not limited to, job-related knowledge, depth of experience, skills, certifications, and geographic location. In addition to the base salary, our compensation structure may include other components such as commissions and discretionary bonuses.

ePlus offers a full range of medical, financial, and/or other benefits (including 401(k) eligibility, employee stock purchase program and various paid time off benefits, such as vacation, sick time, and personal leave), dependent on the position offered. Details of participation in these benefit plans will be provided if an offer of employment is extended.

If hired, employee will be in an “at‑will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Who We Are

At ePlus, we believe technology is a people business. Our team is passionate, skilled, and driven to deliver solutions that make a real difference. Join us and be part of a culture that values collaboration, innovation, and extraordinary results.

Corporate Values

  • Respectful communication and cooperation: We prioritize respectful communication, fostering an environment where everyone is treated with dignity and respect.
  • Teamwork and employee participation: Collaboration and teamwork thrive through diverse perspectives, both within our teams and in our interactions with our customers.
  • Work/life balance that supports our employees’ varying needs: We value the well‑being of our employees, recognizing that a healthy work‑life balance is pivotal to our collective success.
  • Embracing communities: We embrace and support the communities that nurture us. Our employees' dedication to fostering positive change is a source of immense pride for us.

Commitment to Diversity, Inclusion and Belonging

  • We are an equal opportunity employer that does not discriminate or allow discrimination based on race, color, religion, sex, sexual orientation, gender identity, age, national origin, citizenship, disability, veteran status, or any other classification protected by federal, state, or local law.
  • ePlus is dedicated to fostering, cultivating, and preserving a culture that represents diversity, enables inclusion, and makes our employees feel comfortable bringing their full, unique selves to work.

Physical Requirements

  • While performing this role, you will engage in both seated and occasional standing or walking activities. We provide reasonable accommodations, in accordance with relevant laws, to support success in this position.

ePlus maintains a California Consumer Privacy Act (CCPA) Privacy Notice on our Trust Center, available here: CCPA Privacy Notice.

#J-18808-Ljbffr
Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the Principal Solutions Architect in Irvine, CA vacancy
  • $170k - $190k

    OverviewWe are seeking an elite Solutions Architect to lead the end-to-end design, sizing, and deployment of NVIDIA AI Factory-aligned infrastructure. In this highly technical, customer-facing role you will translate complex AI and machine learning workload requirements... 
    Principal
    Contract work
    Local area

    ePlus

    Irvine, CA
    3 days ago
  • $124.5k - $239k

     ....What you’ll be doing...You’ll be using your technical expertise and sales skills to help develop complex wireless and networking solutions for our customers. You’ll provide customer training and advisement, conducting needs analyses, developing appropriate solutions, and... 
    Principal
    Full time
    Temporary work
    Part time
    Work experience placement
    Shift work

    Verizon

    Irvine, CA
    2 days ago
  • $182.8k - $247.3k

    Application deadline: Aug 3, 2026Amazon Web Services (AWS) is looking for a highly motivated Principal Solutions Architect to help accelerate our growing Global Defense Partners business.As a Principal Solutions Architect within AWS, you will have the opportunity to help... 
    Principal
    Local area
    Flexible hours

    AmazonWebServices

    Irvine, CA
    1 day ago
  • $153.6k - $207.8k

     ...understand each customer's unique challenges, then craft innovative solutions that accelerate their success. This customer-first approach is...  ...(AWS), we're hiring Application Modernization Solution Architects to work with and enable our enterprise account teams and customers... 
    Suggested
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    Irvine, CA
    2 days ago
  • $153.6k - $207.8k

     ...understand each customer's unique challenges, then craft innovative solutions that accelerate their success. This customer-first approach is...  ...us and help us grow.The Amazon Web Services (AWS) Solutions Architect team partners with customers to design and build some of the... 
    Suggested
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    Irvine, CA
    1 day ago
  •  ...Finance, we deliver innovative financing, leasing, and insurance solutions to more than 3 million customers and businesses nationwide.We’...  ...of movement. Apply today.WHAT YOU WILL DOThe Sr. Solution Architect, Customer Service Platform, leads the CRM architecture strategy... 
    Full time
    Work at office
    Local area
    Immediate start

    Hyundai Capital America

    Irvine, CA
    2 days ago
  • $77k - $202k

    Industry/SectorNot ApplicableSpecialismFinanceManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Zuora Solution Architect - Senior Associate, you will play a pivotal role in helping clients optimize operational efficiency through the... 
    Full time
    H1b

    PwC

    Irvine, CA
    2 days ago
  • $77k - $202k

     ...business applications, helping clients optimise operational efficiency. These individuals analyse client needs, implement software solutions, and provide training and support for seamless integration and utilisation of business applications, enabling clients to achieve... 
    Full time
    H1b

    PwC

    Irvine, CA
    3 days ago
  • $77k - $202k

     ...situations. This role offers the chance to collaborate with senior architects and technical teams to deliver scalable, exceptional, and...  ...enterprise environments- Analyze complex issues and develop strategic solutions- Mentor and guide junior associates in their professional... 
    Full time
    H1b

    PwC

    Irvine, CA
    1 day ago
  • $141.2k - $278.3k

    Position Summary Manufacturing Executions Systems Solution Architect Manager (AVEVA) Position Summary We are a team of strategic advisors, architects, and implementers who drive business transformations. Our diverse talent energizes clients' business functions and... 
    Local area
    Visa sponsorship

    Deloitte

    Costa Mesa, CA
    1 day ago
  • $30.47 - $56.35 per hour

    Team Name:LocalizationJob Title:Localization Technical Solution Architect | Irvine, CARequisition ID:R027631Job Description:We invite you to join us! Our team is looking for a Localization Technical Solution Architect to help scale and optimize our global localization... 
    Hourly pay
    Full time
    Temporary work
    Part time
    Local area
    Work from home
    Worldwide
    Relocation package

    Blizzard Entertainment

    Irvine, CA
    4 days ago
  • $146k - $194k

     ...technology built and maintained by us. We call this system-of-systems “ArsenalOS”.ABOUT THE ROLEWe are seeking a Oracle Fusion Solution Architect with deep Oracle Fusion experience to join the team. You will be responsible for orchestrating the many cross-functional,... 
    Full time
    Work experience placement
    Work at office
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    1 day ago
  • $110k - $150k

     ...geographically located in or near Irvine, California, with a willingness to work on-site in our Irvine office two days per week.As a Solutions Architect (Enterprise Networking), you will function as the technical lead for ePlus for targeted, key solution sets. You will be a key... 
    Contract work
    Work experience placement
    Work at office
    Local area
    2 days per week

    ePlus

    Irvine, CA
    2 days ago
  • $99k - $232k

    Industry/SectorNot ApplicableSpecialismFinanceManagement LevelManagerJob Description & SummaryThe OpportunityAs a Zuora Solution Architect - Manager, you will play a pivotal role in helping clients optimize their operational efficiency through the implementation and support... 
    Full time
    H1b

    PwC

    Irvine, CA
    2 days ago
  • $143k - $180k

    Why you’ll love Softchoice:We are a software-focused IT solutions and services provider that equips organizations to be agile and innovative...  ...and proofs of concept. You’ll have the unique opportunity to architect and deliver AI strategies that become the backbone of our... 
    Full time
    Remote work
    Flexible hours

    Softchoice

    Irvine, CA
    8 hours ago
  • $129k - $171k

     ...parallel at any given time, enabling operations across the U.S. and internationally.ABOUT THE JOBWe are seeking an experienced Solutions Architect to join the M&A Business Systems team and own hands-on integration work across our full acquisition portfolio. This is not... 
    Full time
    Contract work
    Work experience placement
    Work at office
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    4 days ago
  • $124k - $280k

     ...efficiency. These individuals analyze client needs, implement software solutions, and provide training and support for seamless integration and...  ...ability to working with Microsoft Dynamics 365 Technical Architects, clients, engineering teams including developers, testers and... 

    PwC (US)

    Irvine, CA
    3 days ago
  •  ...Verizon in Irvine is seeking a Senior Solutions Architect to design and deliver complex wireless and networking solutions for enterprise and public sector clients. You will translate business needs into architectures, present to executives, and lead cross‑functional teams... 

    Verizon

    Irvine, CA
    13 hours ago
  •  ...Edwards Lifesciences seeks a Distinguished Anaplan Solutions Architect to lead enterprise planning initiatives, partnering with Finance, Supply Chain, HR, Commercial, and IT to define direction and guide delivery teams. On-site in Irvine, CA. The role emphasizes strategic... 

    Edwards Lifesciences

    Irvine, CA
    13 hours ago
  •  ...The Trade Desk seeks an experienced Anaplan Solution Architect to lead the design, development, and deployment of Anaplan models across various business functions. This role involves collaborating with stakeholders, ensuring seamless system integration, and enhancing... 

    The Trade Desk

    Irvine, CA
    13 hours ago
  •  ...and integration between the application. Defining end-to-end solutions for applications that includes areas like, but not limited to,...  ...systems architecture. JD: Liaison with Customer & Cognizant Chief Architects and other Lead Architects, present architectural solutions in... 
    Temporary work

    TechDigital Group

    Irvine, CA
    4 days ago
  •  ...Solution ArchitectLocation: Irvine, CAJob DescriptionSolution Architect Team Lead for Observability and Anomaly Detection Platform Build leveraging AIProject is to build a new observability platform for middle office, which will serve the need for anomaly detection, data... 
    Work at office

    AceStack LLC

    Irvine, CA
    3 days ago
  • Job ID: 42578Location: Los Angeles, CA, United States | Irvine, CA, United States | San Diego, CA, United StatesDepartment: Water EngineeringWork Type: HybridDate Posted: 2026-07-16
    Principal

    Arcadis

    Irvine, CA
    8 hours ago
  • $160.8k - $241.2k

     ...for data center applications. You will work hands‑on with customers early in the design cycle to understand challenges and shape solutions that support the build‑out and operation of their facilities. You will also identify potential offer enhancements and collaborate... 
    Ongoing contract
    Full time
    Temporary work
    Immediate start
    Flexible hours

    Schneider Electric

    Costa Mesa, CA
    2 days ago
  • $99k - $232k

     ...PwC tax and audit guidance), the Firm's code of conduct, and independence requirements.The OpportunityAs part of the P&C Technical Solutions team, you lead the design and implementation of innovative solutions that modernize insurance operations. As a Manager, you... 
    Full time
    H1b

    PwC

    Irvine, CA
    2 days ago
  • $151k - $204.3k

     ...unmatched technology, and unwavering support. We dive deep to understand each customer's unique challenges, then craft innovative solutions that accelerate their success. This customer-first approach is how we built the world's most adopted cloud. Join us and help us grow... 
    Local area
    Worldwide
    Flexible hours

    Amazon Web Services (AWS)

    Irvine, CA
    13 hours ago
  •  ...Restaurant365, a leading SaaS platform for restaurant operations, seeks a Customer Architect to guide Enterprise clients from design to deployment. You’ll shape scalable solutions across product configuration, integrations, data flows, and workflows, partnering with Success... 

    Restaurant365

    Irvine, CA
    4 days ago
  •  ...as a “Best Place to Work” in Southern California and one of INC.’s 5000 fastest-growing private companies in the U.S. As a Solution Architect, you bridge the gap between visionary ideas and rock-solid enterprise systems, designing end-to-end architectures that... 
    Local area
    Immediate start

    CompassX Group

    Irvine, CA
    a month ago
  • A global technology firm is seeking an experienced Data Scientist / Senior Consultant to drive the strategy and execution of their search and recommendation systems. The ideal candidate will possess a Master’s or Ph.D. in a quantitative field and over a decade of relevant...
    Principal

    Ingram Micro, Inc.

    Irvine, CA
    13 hours ago
  • $131.3k - $177.6k

     ...behalf of our customers, whether that is how we build products and solutions, how we sell, how we deliver, or how we partner. AWS WWPS...  ...how we partner. As an Amazon Web Services (AWS) Solutions Architect in AWS WWPS HCLS segment, you are responsible for partnering... 
    Local area
    Flexible hours

    Amazon Web Services (AWS)

    Irvine, CA
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Solutions Architect. Be the first to apply!