Principal Solutions Architect
$170k - $190kePlus Technology, inc.
Overview We are seeking an elite Solutions Architect to lead the end-to-end design, sizing, and deployment of NVIDIA AI Factory-aligned infrastructure. In this highly technical, customer-facing role you will translate complex AI and machine learning workload requirements into fully engineered infrastructure solutions spanning colocation facilities, GPU compute, high-performance networking, parallel storage, and the complete NVIDIA AI software stack. You will serve as a trusted technical advisor to enterprise and hyperscale customers, partnering with sales, product, and engineering teams to win and deliver transformational AI infrastructure programs. Your expertise will directly shape how organizations build and operate production AI Factories capable of training frontier models, running large-scale inference fleets, and accelerating data science pipelines at scale. Your Impact
Solution Design & Architecture
AI Software
NVIDIA AI Enterprise (NVAIE), NIM Microservices, RAPIDS (cuDF, cuML, cuGraph), NVIDIA Dynamo, CUDA Toolkit, cuDNN, NCCL, TensorRT, Triton Inference Server Cluster Mgmt Base Command Manager, DGXOS, NVIDIA Mission Control, DGX Cloud, UFM, IPMI / Redfish BMC management Orchestration Kubernetes (K8s), NVIDIA GPU Operator, Run:ai GPU scheduling, SLURM, OpenMPI, Helm, Argo Workflows, Kubeflow, KServe Colocation Critical power design (kW - MW), UPS / generator, CRAC / CRAH / DLC / immersion cooling, hot-aisle containment, PUE optimization, carrier-neutral telecom, cross-connects, MMR design Frameworks PyTorch, JAX, TensorFlow, Hugging Face Transformers, DeepSpeed, Megatron-LM, vLLM, LMDeploy Qualifications
If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors. #LI-DY1 #IND1 Who We Are At ePlus, we believe technology is a people business. Our team is passionate, skilled, and driven to deliver solutions that make a real difference. Join us and be part of a culture that values collaboration, innovation, and extraordinary results. Corporate Values
ePlus maintains a California Consumer Privacy Act (CCPA) Privacy Notice on our Trust Center, available here: CCPA Privacy Notice. Notice to Recruiting Agencies: ePlus only accepts unsolicited resumes when presented directly by a candidate. Unsolicited resumes submitted to ePlus from any other source will be considered ePlus property and will not qualify for any placement or referral fees. ePlus will only pay such fees in connection with a valid written agreement between ePlus and the referring agency, and then only after providing advance written approval to the referring agency to submit resumes in connection with a particular opportunity.
Solution Design & Architecture
- Lead discovery workshops to capture AI/ML workload requirements, including model training scale, inference SLAs, data pipeline throughput, and multi-tenancy needs.
- Architect full-stack AI Factory solutions aligned to NVIDIA reference architectures, integrating colocation, GPU compute, networking, storage, and software layers.
- Develop detailed Bills of Materials (BOMs), rack elevation diagrams, network topology drawings, and power/cooling budgets for customer proposals.
- Define GPU cluster architectures using NVIDIA DGX, HGX, and MGX systems with B200, B300, and GB300 Blackwell SXM and NVLink-Switch configurations.
- Design RTX PRO 6000 Blackwell Server Edition deployments for inference-optimized and enterprise AI workloads.
- Conduct workload sizing and TCO/ROI modeling to validate infrastructure dimensioning for training, finetuning, and inference at scale.
- Specify colocation requirements including critical power load (MW-scale), UPS and generator configurations, and PUE targets.
- Design high-density GPU deployments utilizing air-cooled, direct liquid cooling (DLC), and rear-door heat exchanger configurations.
- Define meet-me room (MMR) and cross-connect requirements; specify carrier-neutral telecom diversity strategies.
- Engage colocation providers and data center operators to validate capacity availability and negotiate technical SLAs.
- Coordinate with facilities and MEP engineers to validate power infrastructure from utility feed through PDU to rack level.
- Architect multi-node GPU clusters optimized for large language model (LLM) pre-training, fine-tuning, and reinforcement learning from human feedback (RLHF).
- Size and configure DGX SuperPOD, HGX H/B-series, and MGX modular systems based on model parameter count, dataset size, and iteration timelines.
- Define server firmware, BIOS, BMC, and DGXOS baselines for production GPU infrastructure.
- Establish GPU health monitoring, RAS (Reliability, Availability, Serviceability) policies, and lifecycle management procedures.
- Design backend GPU fabric networks using NVIDIA Quantum InfiniBand (NDR 400Gb/s and HDR 200Gb/s) for distributed training traffic.
- Architect Spectrum-X Ethernet-based AI networking solutions for inference clusters requiring highbandwidth, low-latency connectivity.
- Specify ConnectX-8/7 HCA deployments and configure RDMA over Converged Ethernet (RoCEv2) or InfiniBand transport for NCCL collective operations.
- Integrate BlueField-3 DPUs for GPU-accelerated network functions, storage offload, zero-trust security isolation, and bare-metal provisioning.
- Design leaf-spine and fat-tree topologies for non-blocking bisectional bandwidth in GPU training clusters.
- Define Quality of Service (QoS) policies separating storage, compute fabric, and management plane traffic.
- Design high-performance parallel file system solutions using VAST Data, Hammerspace, and Pure Storage FlashBlade//E for AI training and checkpoint storage.
- Size storage capacity, IOPS, and throughput based on dataset characteristics, checkpoint frequency, and concurrent reader/writer counts.
- Architect multi-tier storage hierarchies: hot NVMe flash (VAST/FlashBlade) for active datasets, warm object storage for model archives, and cold tape/cloud for long-term retention.
- Configure VAST Data Universal Storage for disaggregated storage with NFS, S3, and POSIX access; tune for large sequential read performance.
- Deploy Hammerspace Global Data Environment for distributed data management and NFS-over-RDMA acceleration across geographically dispersed GPU clusters.
- Define data pipeline architectures ingesting from cloud object stores (S3, GCS, ABS) to local flash for GPUlocal data loading without I/O bottlenecks.
- Deploy and configure NVIDIA AI Enterprise (NVAIE) software stack including NVIDIA GPU Operator, NIM microservices, and RAPIDS accelerated data science libraries.
- Architect inference serving infrastructure using NVIDIA NIM (NVIDIA Inference Microservices) for optimized LLM and vision model deployment with autoscaling.
- Implement NVIDIA Dynamo for distributed inference and disaggregated serving of large-scale generative AI models.
- Configure and optimize CUDA toolkit, cuDNN, NCCL communication libraries, and custom kernel environments for training workloads.
- Deploy Base Command Manager and DGXOS for cluster lifecycle management, node provisioning, health dashboards, and job scheduling integration.
- Integrate NVIDIA Mission Control for AI Factory operations, observability, and multi-cluster fleet management.
- Design and deploy Kubernetes-based AI platforms using NVIDIA GPU Operator, integrating with Run:ai for dynamic GPU resource scheduling and multi-tenant workload isolation.
- Configure SLURM workload manager for traditional HPC-style job scheduling on bare-metal GPU clusters, including preemption policies, fair-share scheduling, and burst-to-cloud integration.
- Establish MLOps toolchain integrations with popular frameworks (PyTorch, JAX, TensorFlow) and experiment tracking platforms (MLflow, Weights & Biases).
- Serve as primary technical point of contact throughout the pre-sales and delivery lifecycle, from initial discovery through post-deployment optimization.
- Produce and present architecture design documents, technical proposals, and executive-level briefings to CTO/CIO and VP-level stakeholders.
- Lead proof-of-concept (POC) and pilot deployments, including benchmark design, execution, and results analysis.
- Collaborate with procurement, logistics, and deployment teams to ensure on-time delivery of complex infrastructure programs.
- Provide post-deployment hypercare support, performance tuning, and capacity planning advisory services.
- Contribute to internal knowledge bases, solution playbooks, and reference architectures for repeatable AI Factory deployments.
AI Software
NVIDIA AI Enterprise (NVAIE), NIM Microservices, RAPIDS (cuDF, cuML, cuGraph), NVIDIA Dynamo, CUDA Toolkit, cuDNN, NCCL, TensorRT, Triton Inference Server Cluster Mgmt Base Command Manager, DGXOS, NVIDIA Mission Control, DGX Cloud, UFM, IPMI / Redfish BMC management Orchestration Kubernetes (K8s), NVIDIA GPU Operator, Run:ai GPU scheduling, SLURM, OpenMPI, Helm, Argo Workflows, Kubeflow, KServe Colocation Critical power design (kW - MW), UPS / generator, CRAC / CRAH / DLC / immersion cooling, hot-aisle containment, PUE optimization, carrier-neutral telecom, cross-connects, MMR design Frameworks PyTorch, JAX, TensorFlow, Hugging Face Transformers, DeepSpeed, Megatron-LM, vLLM, LMDeploy Qualifications
- Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related technical discipline; Master's degree preferred.
- 8+ years of solutions architecture, systems engineering, or technical pre-sales experience, with at least 4 years focused on GPU infrastructure or HPC environments.
- Proven track record designing and deploying NVIDIA DGX or HGX-based GPU clusters in production AI/ML environments.
- Deep understanding of distributed deep learning concepts: tensor parallelism, pipeline parallelism, data parallelism, gradient checkpointing, and mixed-precision training.
- Hands-on experience with InfiniBand or high-speed Ethernet fabric design, RDMA configuration, and collective communication tuning (NCCL, MPI).
- Direct experience sizing and deploying parallel storage systems (VAST, Hammerspace, or Lustre/WEKA/GPFS) for AI training workloads.
- Strong working knowledge of Kubernetes, GPU Operator, and at least one GPU workload scheduler (Run:ai or SLURM).
- Experience with Linux system administration, CUDA development environment configuration, and GPU driver/firmware management.
- Demonstrated ability to create compelling technical proposals, architecture diagrams (Visio/Lucidchart/draw.io), and BOM-level documentation.
- Exceptional communication skills with proven ability to present to both deep technical audiences and Clevel executives.
- NVIDIA-certified professional credentials (DCA-Core, NCP-DS, or equivalent).
- Experience with NVIDIA Base Command Platform or Mission Control for multi-cluster AI Factory operations.
- Familiarity with sovereign AI, government cloud, or regulated industry AI infrastructure requirements.
- Experience integrating AI Factory infrastructure with public cloud (AWS, Azure, GCP) for hybrid and burstto-cloud architectures.
- Background in MLOps, LLMOps, or platform engineering for production AI model lifecycle management.
- Prior experience with colocation data center procurement, RFP development, and SLA negotiation.
- Contributions to open-source AI infrastructure projects or published technical content (blogs, whitepapers, conference presentations).
- Active participation in the NVIDIA Partner Network (NPN) ecosystem or prior experience at an NVIDIA Elite Solution Provider.
If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors. #LI-DY1 #IND1 Who We Are At ePlus, we believe technology is a people business. Our team is passionate, skilled, and driven to deliver solutions that make a real difference. Join us and be part of a culture that values collaboration, innovation, and extraordinary results. Corporate Values
- Respectful communication and cooperation: We prioritize respectful communication, fostering an environment where everyone is treated with dignity and respect.
- Teamwork and employee participation: Collaboration and teamwork thrive through diverse perspectives, both within our teams and in our interactions with our customers.
- Work/life balance that supports our employees' varying needs: We value the well-being of our employees, recognizing that a healthy work-life balance is pivotal to our collective success.
- Embracing communities: We embrace and support the communities that nurture us. Our employees' dedication to fostering positive change is a source of immense pride for us.
- We are an equal opportunity employer that does not discriminate or allow discrimination based on race, color, religion, sex, sexual orientation, gender identity, age, national origin, citizenship, disability, veteran status, or any other classification protected by federal, state, or local law.
- ePlus is dedicated to fostering, cultivating, and preserving a culture that represents diversity, enables inclusion, and makes our employees feel comfortable bringing their full, unique selves to work.
- While performing this role, you will engage in both seated and occasional standing or walking activities. We provide reasonable accommodations, in accordance with relevant laws, to support success in this position.
- By embracing our values, you will contribute to our collective mission of making a positive impact within our organization and the broader community. We understand that this job description serves as a guide and is not an employment contract.
ePlus maintains a California Consumer Privacy Act (CCPA) Privacy Notice on our Trust Center, available here: CCPA Privacy Notice. Notice to Recruiting Agencies: ePlus only accepts unsolicited resumes when presented directly by a candidate. Unsolicited resumes submitted to ePlus from any other source will be considered ePlus property and will not qualify for any placement or referral fees. ePlus will only pay such fees in connection with a valid written agreement between ePlus and the referring agency, and then only after providing advance written approval to the referring agency to submit resumes in connection with a particular opportunity.
Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the Principal Solutions Architect in San Ramon, CA vacancy
- Description Responsibilities:Solutions Architect (level 4: 12-15 Yrs)Experience requirements 12-15 years of experience in leading Product development, Consulting, Solutioning and Architecture Architecture and TechnologyMust Haves Strong experience building cloudnative microservicesbased...SuggestedShift work
$129.7k - $176.7k
...supportive people, willing to listen to your ideas. Job ResponsibilitiesUnderstanding the client’s needs, demonstrating the product and solution capabilities, scoping services, designing, and proposing delivery options, addressing any objections, and writing contracts and...SuggestedFull timeContract workWork experience placementWork at officeLocal areaFlexible hours- ...n and implementation of the overall solution architecture comprising of conceptual (functional and non functional), technical and physical architecture. Demonstrate Thought Leadership towards white space solutions. Provide system & application level solutions framework...Suggested
- ...AI Solution Architect The AI Solution Architect will be responsible for translating complex business challenges into viable, scalable, and secure AI/ML solutions. This role requires a deep understanding of AI/ML methodologies, data architectures, cloud platforms, and...Suggested
$133.3k - $181.7k
...willing to listen to your ideas. We are seeking a client-focused Data Analytics, AI & Microsoft Fabric Pre-Sales Architect to join our Artificial Intelligence Solutions practice. This role blends deep expertise in modern data platforms with the ability to shape and sell AI-...SuggestedFull timeContract workLocal areaFlexible hours$158k - $175k
Title: Principal Project Manager - Transmission & SubstationLocation: Northern California (San Ramon), Reno, NV, or Las Vegas, NVHire Type: Direct HireSalary: $158,000 - $175,000 (based on education and experience)Benefits: Medical, dental, visionSterling Engineering is...PrincipalWork at office- ...We’re seeking a seasoned Principal Transmission & Substation Project Manager to lead transformative infrastructure projects that shape... ...a proven leader, and passionate about delivering high-impact solutions, we want to hear from you. About the Role Lead major capital projects...Principal
- ...Relanto is seeking a highly skilled Anaplan Solution Architect to lead the design, development, and implementation of scalable Anaplan models that drive planning, forecasting, and analysis across departments. You will collaborate with finance, IT, and business stakeholders...
- PG&E is seeking an IT Infrastructure/Engineering leader to provide technical leadership for Common/Critical Facilities infrastructure. The role focuses on IT operations and project delivery, collaborating with internal and external teams to design, build, and support essential...
$170.36k - $230k
...kinds can tap into the world’s largest network of branded payment solutions. BHN helps businesses grow revenue, increase loyalty, motivate... ...reshaping how we operate. We are looking for a hands-on AI Architect who can partner directly with business leaders to identify,...Full timeWork experience placementWork at officeLocal areaRemote workFlexible hours- Hyve Company Overview Hyve Solutions transforms complex engineering challenges into production reality for technology innovators building... ...’s architecture and engineering organizations. The Solutions Architect will work directly with hyperscalers, cloud providers, AI-...Full timeFlexible hours
$98.04k - $154.8k
Position OverviewWe are seeking a high-energy, passionate, and visionary Solution Architect to join our Enterprise Strategy & Architecture group. Reporting to the Lead Enterprise Architect, you will act as a Master Planner with a big-picture mindset, responsible for turning...Full timeTemporary workWork at office2 days per week1 day per week- ...above and beyond to aid stores and customers and deliver timely solutions to benefit all members of Grocery Outlet. Our team consists of... ...We are seeking a highly experienced Senior Data & AI Solution Architect to lead the design and implementation of scalable data...Full time
$155k - $165k
...company. We provide integrated credit score and personal finance solutions to 1,600 + bank and credit union partners nationally. The... ...to the office for in-person meetings. The Senior Solution Architect is a highly skilled individual contributor at the heart of SavvyMoney...Full timeLocal areaRemote workWork from homeFlexible hours$251.84k - $314.79k
...Labs throughout our recruiting process as we integrate our teams, systems, and career sites.About the RoleFivetran needs an engineer-architect who can make the AI analyst experience real, dependable, and extensible. This role owns the architecture behind the analyst AI...PrincipalFull timeWork at officeRemote workFlexible hours- ...Solution Architect Location: Hybrid in Dublin, CA; 1 day a week on-site Qualifications and Special Skills Required: • Bachelor's Degree in Computer Science, Information Technology, or related field and 10+ years’ experience in information technology (IT), technology...1 day per week
$141k - $307k
...quality of customers installed base performance and deliver service and lifecycle solutions for their most critical equipment and processes. The impact you’ll make As a Solution Architect at Lam, you will be at the forefront of innovation by designing, developing, and troubleshooting...Full timeLocal areaRemote workFlexible hours2 days per week3 days per week1 day per week- ...But Are Not Limited To: Collaborates with the Enterprise Architect (EA), business and the project team to understand business... ...reusable service components and patterns. Ensures that the solution architecture and design align with the Target Architecture for...Work experience placement
- ...Team Please dont Submit Salesforce Developer or lead. Look for Solutions Architect, with data Modelling . Lower rate is the best , Solution Architect Location: Pleasanton, California Onsite Required Skills - MUST Salesforce Lightning Salesforce Security...
$98.04k - $154.8k
...Solution Architect We are seeking a high-energy, passionate, and visionary Solution Architect to join our Enterprise Strategy & Architecture group. Reporting to the Lead Enterprise Architect, you will act as a Master Planner with a big-picture mindset, responsible...Full timeTemporary workWork at officeFlexible hours2 days per week1 day per week- ...Solution Architect Oakland, CA 12+ months Required Experience: Minimum of five (5) years of solution architecture (SA) and played a crucial role in shaping an organization's digital transformation, ensuring that technology investments align with business goals...
- ...Job Description Position Overview: We are seeking an experienced Solution Architect with over 10 years of expertise in Field Service Management tools and a strong background in the utilities and energy industries. The successful candidate will lead the design...
- ...Prellis Bio in Berkeley, CA seeks a Senior/Principal Scientist — Antibody Protein Engineering to drive late-stage antibody design and optimization. You will mature hits into development-ready leads while building an integrated wet-lab and computational engineering platform...Principal
$128.56k - $160.7k
...communications, all from the comfort of our homes. We deliver innovative solutions tohundreds of thousands of businessesand empower millions of... ...at Twilio Join the team as Twilio's next Staff Solutions Architect, Contract LifeCycle Management. About the job We are...Contract workLocal areaRemote workWorldwide- ...Senior Talent Acquisition Specialist – Tata Technologies As a 3DEXPERIENCE Solution Architect, you will lead the design and implementation of PLM (Product Lifecycle Management) solutions using Dassault Systèmes’ 3DEXPERIENCE platform. You’ll work closely with clients...
- ...delivering the efficiencies, automation, and developer experience of the cloud in a form that businesses can own. We’re seeking Solutions Architects to serve as a senior technical voice in customer engagements, with prospects, and with partners. As a Solutions Architect...Shift work
- ...Ascentt Business Systems, Inc. is seeking a qualified professional to fill the position of Solution Architect based in Fremont, CA. Responsibilities will include: Develop software applications for data integration requirements on advanced analytics, data science and machine...
- ...Posting Seeking an experienced Associate Principal to lead application architecture and... ...initiatives using AWS Well Architected Tools and cloud architecture best practices... ...architecture best practices into existing and new solutions Ensure architectural compliance with...
- ...above and beyond to aid stores and customers and deliver timely solutions to benefit all members of Grocery Outlet. Our team consists of... ...solving important problems.About the Role: The Data Solution Architect will help define strategy both within and outside of SAP,...Full time
$137k - $287k
The group you’ll be a part ofGlobal Information SystemsThe impact you’ll makeAs a Data architect supporting digital transformation in GBIS (Global Business Intelligence Solutions) group at Lam Research, you'll be responsible for designing and developing data engineering...Local areaRemote workFlexible hours2 days per week3 days per week1 day per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Solutions Architect. Be the first to apply!

