Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI HPC Infrastructure Engineer

$150k - $170k

Analysis Group

OverviewAnalysis Group is one of the largest international economics consulting firms, with more than 1,500 professionals across 15 offices in North America, Europe, and Asia. Since 1981, we have provided expertise in economics, finance, health care analytics, and strategy to top law firms, Fortune Global 500 companies, and government agencies worldwide. Our internal experts, together with our network of affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure environment. The engineer maintains the Linux-based clustered computing platform that supports both traditional HPC/analytical workloads and large-scale AI/ML training and inference, ensuring systems run efficiently, GPUs and other accelerators are current and well-utilized, and operations are monitored, documented, and reported — including change management and performance statistics — across both domains.Essential Job Functions and ResponsibilitiesMaintain, tune, and manage the analytical and AI computing environment for researchers and data scientists, including Posit Workbench (RStudio Server Pro) environments.Optimize systems and infrastructure performance using parallelization technologies (MPI, OpenMP) and distributed/multi-GPU training strategies (e.g., PyTorch Distributed, Horovod, DeepSpeed).Design, deploy, and maintain GPU-accelerated compute infrastructure for large-scale model training and inference.Manage GPU scheduling, multi-tenancy, and utilization across SLURM and/or Kubernetes-based environments.Administer the NVIDIA software stack — drivers, CUDA, cuDNN, NCCL — and coordinate firmware and health monitoring across GPU fleets.Tune and optimize LLM training and inference performance — including batching, quantization, KV-cache utilization, parallelism strategies, and throughput/latency across GPU clusters.Build and maintain MLOps pipelines for model training, versioning, deployment, and monitoring (e.g., MLflow, Kubeflow).Manage container orchestration and runtimes (Docker, Kubernetes, Singularity/Apptainer) supporting both HPC jobs and ML workloads.Manage access authentication including PAM, LDAP integration, and single sign-on.Design and develop scripts for system administration, automating tasks, monitoring, and usage reporting across HPC and AI resources.Manage high-performance storage and data pipelines for AI training datasets and HPC workloads, primarily on GPFS (IBM Spectrum Scale).Troubleshoot, isolate, and resolve application, systems, and other technical problems (hardware, software, network, and GPU-specific issues).Develop and implement backup and recovery programs.Research, deploy, and manage general infrastructure, including development of policies and procedures for both HPC and AI/ML environments.Migrate data from heterogeneous environments to Linux, on-prem clusters, or cloud.Collaborate with data scientists and ML engineers to support the model development lifecycle and translate research needs into infrastructure requirements.Evaluate emerging AI hardware, accelerators, and cloud AI services, and recommend adoption where beneficial.Monitor performance, troubleshoot problem areas, and provide statistics and reports across compute, storage, and network.Create and maintain documentation related to system configuration, processes, change management, inventory, and service records.Ensure continuous network connectivity of all equipment.Conduct research and report on products, services, protocols, and standards to remain abreast of developments in HPC and AI infrastructure.Participate in a 24x7 on-call rotation; troubleshoot and resolve issues remotely or onsite as necessary.QualificationsBachelor's degree required; degree in computer science, electrical engineering, or a related field preferred.A minimum of 5 years of experience as a hands-on Linux Systems Administrator in a research, HPC, or production setting.An ideal candidate will have 5 to 10 years of substantive relevant experience. Experience managing Posit Workbench (RStudio Server Pro), Python, and R environments; strong Posit Workbench administration experience is a significant plus.Experience with SLURM, Platform LSF, or other job schedulers required; experience scheduling GPU resources strongly preferred.Hands-on experience with NVIDIA GPU infrastructure and software stack (CUDA, cuDNN, NCCL, NVIDIA GPU Operator) strongly preferred.Experience with Kubernetes and container orchestration for AI/ML workloads highly desired.Familiarity with ML/AI frameworks (PyTorch, TensorFlow) and distributed training patterns highly desired.Experience with MLOps tooling (MLflow, Kubeflow, Weights & Biases, or similar) is a plus.Experience with Bright Cluster Manager is highly desired.Experience with Ansible is highly desired.Experience with containerization (Docker, Singularity/Apptainer) is highly desired.Proficiency with remote access technologies and tools such as RDP, SSH, and emulation softwareHands-on experience with GPFS (IBM Spectrum Scale) required.Demonstrated experience tuning LLM training and/or inference performance (e.g., batching, quantization, KV-cache management, parallelism strategies) required.Experience with AI Gateways (e.g., LiteLLM, Kong AI Gateway, Portkey, or similar) is a very nice to have.Excellent hardware troubleshooting experience, including GPU-specific diagnostics.Knowledge of applicable data privacy practices and laws.Strong interpersonal, written, and oral communication skills.Highly self-motivated and directed, with keen attention to detail.Proven analytical and problem-solving abilities.Strong customer service orientation.Experience working in a collaborative environment.An inclusive and growth-oriented mindset, strong interpersonal skills, and an ability to work across functions.To the extent permitted by applicable law, eligible candidates must be authorized to work in the United States, without sponsorship or restriction, now and in the future.Analysis Group embraces equal opportunity. We are committed to building teams that bring a variety of backgrounds, perspectives, and skills, as we believe that a strong and inclusive workforce directly supports our goal of providing the highest-quality work. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other class protected under applicable federal, state, or local law, and we encourage candidates of all backgrounds to apply.Analysis Group offers competitive compensation and a comprehensive benefits package. The estimated salary range for this position is $150,000–$170,000. Compensation offered will be based on a number of factors including work experience, education, and skill level. This role is eligible for a discretionary annual bonus that is determined in large part by individual performance. To learn more about our benefit offerings, click here.#LI-HybridPrivacy NoticeFor information about Analysis Group’s privacy practices, please refer to the applicable Analysis Group privacy policy.­Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities.Please view the EEOC’s “Know Your Rights” poster here.Job SummaryCategory: Information Technology

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI HPC Infrastructure Engineer in Boston, MA vacancy
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for...  ..., and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks...  ...of 1,000+ GPU clusters for AI, HPC, and agentic AI workloads with infrastructure... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Boston, MA
    1 day ago
  •  ...approach. Teradata delivers real business value with AI. What You'll Do We are seeking a Staff Software Engineer to lead the design, development, and evolution of...  ...Technical Skills: Strong background in HPC or large-scale distributed systems development.... 
    Suggested
    Permanent employment
    Flexible hours

    Teradata

    Boston, MA
    4 days ago
  • $240k - $293k

     ...efficient, high-quality, and consistent service, incorporating AI tools and outcome-focused strategies across the product lifecycle...  ...drives sales, delivery, and client outcomes across product engineering and cloud modernization, supporting clients as they build, scale... 
    Suggested
    Temporary work
    Local area

    Slalom

    Boston, MA
    1 day ago
  •  ...which every person has a personalized, AI-enabled doctor always in their pocket. To...  ...the line between "application code" and "infrastructure" keeps blurring and we need someone who...  ...re looking for a Backend Infrastructure Engineer who's equal parts software engineer and... 
    Suggested
    Local area
    Immediate start
    Remote work

    Counsel Health

    Boston, MA
    2 days ago
  • $120k - $175k

    Watertown, MAR&D - Software /Full-time /On-siteRobotics Infrastructure Engineer: Systems, Infrastructure & ReliabilityThe CompanyWe believe general...  ..., and researchers who want to be part of that loop.As an AI robotics company that deploys its inventions directly into the... 
    Suggested
    Full time
    Work at office
    Shift work
    Night shift

    Tutor Intelligence

    Watertown, MA
    4 days ago
  • $87.5k - $131.3k

     ...Network Engineer Boston, Massachusetts, United States We take play seriously. We're...  ...AWS cloud. Our network is evolving: an AI-driven Juniper/Mist campus fabric, a Palo...  ...operate our Juniper/Mist campus and branch infrastructure. Engineer and maintain the Palo Alto... 
    Work experience placement
    Local area

    Hasbro

    Boston, MA
    2 days ago
  • Senior IT Recruitment Consultant - Greater Boston AI & Systems Engineer - Legal Tech | Boston, MA (Hybrid - 3 days)| Law Firm A prestigious...  ...: 5-7 years of experience in IT systems engineering or infrastructure, ideally with recent AI project exposure Deep familiarity... 
    Full time

    Franklin Fitch

    Boston, MA
    2 days ago
  • $135k - $205k

     ...Network Engineer SimSpace serves as an AI Proving Ground where organizations can confidently train, test, and outmaneuver adversaries in any...  ...monitoring, maintenance, and security of enterprise network infrastructure across SimSpace corporate offices and data centers... 
    For subcontractor
    Local area
    Remote work
    Worldwide
    Flexible hours

    SimSpace Corporation

    Boston, MA
    3 days ago
  • $245k - $275k

     .... By weaving together advances in cloud infrastructure, automation and analytics, and software...  ...AHEAD. The AHEAD Senior Specialty Solutions Engineer-Network, will be focused on the core...  ...benefits for additional details. Use of AI:We may use artificial intelligence (AI)... 
    Full time
    Work at office

    AHEAD

    Boston, MA
    3 days ago
  • $180k - $240k

    About WorkatoWorkato delivers enterprise infrastructure for the agentic era, redefining iPaaS...  ...unify data, applications, processes, and AI into a single, governed platform. A leader...  ...are hiring a Senior Infrastructure Engineer to join our Global Core Infrastructure team... 
    Full time
    Remote work
    Flexible hours

    Workato

    Boston, MA
    4 days ago
  •  ...CGS is seeking an experienced Network Engineer to join a team focused on the evaluation...  ...includes both wired and wireless network infrastructure and related hardware & software. The project...  ...We may use artificial intelligence (AI) tools to support parts of the hiring process... 
    Full time
    Local area
    Monday to Friday
    Flexible hours

    CGS Federal (Contact Government Services)

    Boston, MA
    4 days ago
  •  ...ID, User-ID, Content-ID, Decryption, and AI-Ops. Proven ability to design,...  ...Functional Skills Level 2/3 Network Security Engineer Deep and strong understanding of firewall...  ...troubleshoot complex enterprise network infrastructure. Good understanding of Remote Access products... 
    Local area
    Remote work

    HCLTech

    Quincy, MA
    1 day ago
  • $90 - $100 per hour

    Get AI-powered advice on this job and more exclusive features. Direct message the job...  ...Technology Partners Sr. Network Engineer - 4 Days / Week onsite in Boston - Contract...  ...implement, and support enterprise networking infrastructure Architect, deploy, and operate SASE... 
    Full time
    Contract work
    Remote work

    Maverick Technology Partners

    Boston, MA
    2 days ago
  • $244k - $366k

     ...independently drive their growth, and the Engineering Department's contribution to this mission is crucial. As Director of Production Infrastructure, you'll helm the creation and management...  ...accountability. Optimize the use of AI to enhance infrastructure management and... 

    Klaviyo

    Boston, MA
    4 days ago
  • $153.6k - $207.8k

     ...business objectives and translate them into innovative cloud and AI solutions that drive meaningful impact in pharmaceutical...  ...implementation experience- Bachelor's degree in Computer Science, Engineering, a related field, or equivalent experience- Experience positioning... 
    Flexible hours

    AmazonWebServices

    Boston, MA
    1 day ago
  •  ...capabilities across data, edge, integrated infrastructure and applications, deep ecosystem skills,...  ...—including Gemini Enterprise, Data & AI, CES, Security, and Gen AI—that empower...  ...Build a clear career pathway toward senior engineering, architecture, or leadership roles... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Boston, MA
    4 days ago
  • $180.5k - $236.91k

    Hi, we're Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team.Oscar is the first health insurance...  ...wellness time and reimbursements.Artificial Intelligence (AI): Our AI Guidelines outline the acceptable use of artificial... 
    Full time
    Work at office
    Remote work

    Oscar Health Insurance

    Boston, MA
    3 days ago
  • $107.45k - $199.55k

    Join Our Cloud Engineering TeamAre you passionate about building scalable, secure, and modern...  ...implement cloud-based applications and infrastructure solutions.Build and maintain cloud environments...  ..., such as artificial intelligence (AI), and automated processing tools, to... 
    Full time
    Temporary work
    Local area
    Flexible hours

    John Hancock

    Boston, MA
    2 days ago
  • $90k - $210k

     ...specially selected team of scientists, engineers, and software developers to deliver best...  ...support projects that provide critical AI, machine learning, simulation, situational...  ...using Python.Develop and maintain cloud infrastructure using infrastructure-as-code practices.Oversee... 
    Full time

    MORSE Corp

    Cambridge, MA
    18 hours ago
  • $123.1k - $186.3k

     ...within 12 months to ensure you are not duplicating efforts.Job CategoryCustomer SuccessJob DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a... 
    Full time
    Work experience placement
    Remote work

    Salesforce

    Boston, MA
    18 hours ago
  • $71.25k - $143.75k

     ...we are now searching for a Senior Cloud Engineer to join our team in Boston.Your mission:...  ...Reporting directly to the Head of Technology Infrastructure, the Senior Cloud Engineer will fill a...  ...portalThe Process We use the power of AI to help our partners make decisions. If... 
    Full time
    Temporary work
    Local area

    Sophia Genetics

    Boston, MA
    4 days ago
  •  ...seeking a highly skilled and motivated Senior Microsoft 365 Cloud Engineer to join our Workplace Engineering team. This role is responsible...  ...vendors to deliver innovative collaboration, productivity, and AI-powered solutions that support Sallie Mae's digital workplace... 
    Full time
    Temporary work
    Local area
    Flexible hours

    SallieMae

    Newton, MA
    3 days ago
  • $110k - $140k

     ...We are seeking a highly skilled Network Engineer to join our Managed Service Provider (MSP...  ..., and managing complex network infrastructures with a particular focus on Palo Alto firewalls...  ...including implementation and management of AI‑driven wireless networks Experience... 
    Remote work
    Flexible hours
    Afternoon shift

    SHI GmbH

    Boston, MA
    4 days ago
  •  ...like silicon production, IoT, and critical infrastructure with full device ownership, control, and...  .... Role Overview As a Software Engineer at zeroRISC, you will develop our suite...  ...We may use artificial intelligence (AI) tools to support parts of the hiring process... 
    Full time

    Zerorisc

    Boston, MA
    22 hours ago
  • $140k - $180k

     ...insurance subsidiaries. TEAM OVERVIEW  The ADAPT (AI, Data, and Platform Technologies) Engineering team is integral to KKR's technological strategy,...  ...environments, with familiarity in orchestration, infrastructure automation, and platform observability tooling.... 
    Full time
    Local area

    Careers at KKR

    Boston, MA
    4 days ago
  • $152k - $221k

     ...facing or support role.Experience with cloud engineering, on-premise engineering, virtualization,...  ...in one or more of the following: infrastructure modernization, application modernization...  ...data management, data analytics, cloud AI, networking, migrations, security.Experience... 

    Google

    Cambridge, MA
    18 hours ago
  •  ...investment firm based in Boston is seeking a Senior Cloud Security Engineer to shape and deliver cloud security strategy in a collaborative...  ...improvements Work closely with firmware teams to embed and test AI features on hardware platforms Set up and oversee tools for... 
    Full time

    Motion Recruitment

    Boston, MA
    3 days ago
  • $110k - $207.5k

    Corporate Functions Senior AWS Cloud Engineer - VP IIIWho We Are Looking ForWe are seeking...  ...complex AWS-based applications and AI-enabled platforms across Corporate Functions...  ...application, architecture, cybersecurity, infrastructure, and business teams to deliver... 
    Full time
    Temporary work
    Flexible hours

    State Street Bank

    Boston, MA
    2 days ago
  •  ...aviation, defense, energy, and other critical infrastructure domains. Backed by top-tier investors...  ...-stake efforts to deliver the Flyways AI Platform to our government customers....  ...software solutions alongside our product engineers, from adding new features to integrating... 

    Air space Intelligence

    Boston, MA
    4 days ago
  • $148k - $222k

     ...creators to own their own destiny.The mission of the Platform Engineering team is to provide infrastructure primitives, platforms, tooling, and guidance to Klaviyo...  ...processing tooling.You’ve already experimented with AI in work or personal projects, and you’re excited to... 
    Work at office

    Klaviyo

    Boston, MA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI HPC Infrastructure Engineer. Be the first to apply!