Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI HPC Infrastructure Engineer

$150k - $170k

Analysis Group

Overview Analysis Group is one of the largest international economics consulting firms, with more than 1,500 professionals across 15 offices in North America, Europe, and Asia. Since 1981, we have provided expertise in economics, finance, health care analytics, and strategy to top law firms, Fortune Global 500 companies, and government agencies worldwide. Our internal experts, together with our network of affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure environment. The engineer maintains the Linux-based clustered computing platform that supports both traditional HPC/analytical workloads and large-scale AI/ML training and inference, ensuring systems run efficiently, GPUs and other accelerators are current and well-utilized, and operations are monitored, documented, and reported — including change management and performance statistics — across both domains.Essential Job Functions and ResponsibilitiesMaintain, tune, and manage the analytical and AI computing environment for researchers and data scientists, including Posit Workbench (RStudio Server Pro) environments.Optimize systems and infrastructure performance using parallelization technologies (MPI, OpenMP) and distributed/multi-GPU training strategies (e.g., PyTorch Distributed, Horovod, DeepSpeed).Design, deploy, and maintain GPU-accelerated compute infrastructure for large-scale model training and inference.Manage GPU scheduling, multi-tenancy, and utilization across SLURM and/or Kubernetes-based environments.Administer the NVIDIA software stack — drivers, CUDA, cuDNN, NCCL — and coordinate firmware and health monitoring across GPU fleets.Tune and optimize LLM training and inference performance — including batching, quantization, KV-cache utilization, parallelism strategies, and throughput/latency across GPU clusters.Build and maintain MLOps pipelines for model training, versioning, deployment, and monitoring (e.g., MLflow, Kubeflow).Manage container orchestration and runtimes (Docker, Kubernetes, Singularity/Apptainer) supporting both HPC jobs and ML workloads.Manage access authentication including PAM, LDAP integration, and single sign-on.Design and develop scripts for system administration, automating tasks, monitoring, and usage reporting across HPC and AI resources.Manage high-performance storage and data pipelines for AI training datasets and HPC workloads, primarily on GPFS (IBM Spectrum Scale).Troubleshoot, isolate, and resolve application, systems, and other technical problems (hardware, software, network, and GPU-specific issues).Develop and implement backup and recovery programs.Research, deploy, and manage general infrastructure, including development of policies and procedures for both HPC and AI/ML environments.Migrate data from heterogeneous environments to Linux, on-prem clusters, or cloud.Collaborate with data scientists and ML engineers to support the model development lifecycle and translate research needs into infrastructure requirements.Evaluate emerging AI hardware, accelerators, and cloud AI services, and recommend adoption where beneficial.Monitor performance, troubleshoot problem areas, and provide statistics and reports across compute, storage, and network.Create and maintain documentation related to system configuration, processes, change management, inventory, and service records.Ensure continuous network connectivity of all equipment.Conduct research and report on products, services, protocols, and standards to remain abreast of developments in HPC and AI infrastructure.Participate in a 24x7 on-call rotation; troubleshoot and resolve issues remotely or onsite as necessary.QualificationsBachelor's degree required; degree in computer science, electrical engineering, or a related field preferred.A minimum of 5 years of experience as a hands-on Linux Systems Administrator in a research, HPC, or production setting.An ideal candidate will have 5 to 10 years of substantive relevant experience. Experience managing Posit Workbench (RStudio Server Pro), Python, and R environments; strong Posit Workbench administration experience is a significant plus.Experience with SLURM, Platform LSF, or other job schedulers required; experience scheduling GPU resources strongly preferred.Hands-on experience with NVIDIA GPU infrastructure and software stack (CUDA, cuDNN, NCCL, NVIDIA GPU Operator) strongly preferred.Experience with Kubernetes and container orchestration for AI/ML workloads highly desired.Familiarity with ML/AI frameworks (PyTorch, TensorFlow) and distributed training patterns highly desired.Experience with MLOps tooling (MLflow, Kubeflow, Weights & Biases, or similar) is a plus.Experience with Bright Cluster Manager is highly desired.Experience with Ansible is highly desired.Experience with containerization (Docker, Singularity/Apptainer) is highly desired.Proficiency with remote access technologies and tools such as RDP, SSH, and emulation softwareHands-on experience with GPFS (IBM Spectrum Scale) required.Demonstrated experience tuning LLM training and/or inference performance (e.g., batching, quantization, KV-cache management, parallelism strategies) required.Experience with AI Gateways (e.g., LiteLLM, Kong AI Gateway, Portkey, or similar) is a very nice to have.Excellent hardware troubleshooting experience, including GPU-specific diagnostics.Knowledge of applicable data privacy practices and laws.Strong interpersonal, written, and oral communication skills.Highly self-motivated and directed, with keen attention to detail.Proven analytical and problem-solving abilities.Strong customer service orientation.Experience working in a collaborative environment.An inclusive and growth-oriented mindset, strong interpersonal skills, and an ability to work across functions.To the extent permitted by applicable law, eligible candidates must be authorized to work in the United States, without sponsorship or restriction, now and in the future.Analysis Group embraces equal opportunity. We are committed to building teams that bring a variety of backgrounds, perspectives, and skills, as we believe that a strong and inclusive workforce directly supports our goal of providing the highest-quality work. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other class protected under applicable federal, state, or local law, and we encourage candidates of all backgrounds to apply.Analysis Group offers competitive compensation and a comprehensive benefits package. The estimated salary range for this position is $150,000–$170,000. Compensation offered will be based on a number of factors including work experience, education, and skill level. This role is eligible for a discretionary annual bonus that is determined in large part by individual performance. To learn more about our benefit offerings, click here.#LI-Hybrid Privacy Notice For information about Analysis Group’s privacy practices, please refer to the applicable Analysis Group privacy policy.

  • Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities.• Please view the EEOC's "Know Your Rights" poster here.Job SummaryCategory: Information Technology

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI HPC Infrastructure Engineer in Boston, MA vacancy
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for...  ..., and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks...  ...of 1,000+ GPU clusters for AI, HPC, and agentic AI workloads with infrastructure... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Boston, MA
    1 day ago
  • $114.6k - $234.6k

     ...accelerated High-Performance Compute (HPC), Artificial Intelligence and...  ...tailored specifically for AI, ML, HPC workloads. We...  ...-scale global Oracle Cloud Infrastructure (OCI). Primarily focused on the...  ...Collaborate with engineers from L1 optical engineering team... 
    Suggested
    Temporary work
    Flexible hours

    Oracle

    Boston, MA
    4 days ago
  • $122.2k - $240.5k

    Position Summary Senior Network Engineer Our Enterprise Networks practice, part of Hybrid Cloud Infrastructure within AI & Engineering, helps clients architect, modernize, and operate the network foundation that underpins their digital and AI-driven ambitions. We... 
    Suggested
    Local area

    Deloitte

    Boston, MA
    2 days ago
  • $147.25k - $176.75k

    General Information Job Title Staff I Infrastructure Engineer, Tech Solutions Group Job ID 109696 Work Areas Technology & Engineering...  ...Engineer I will collaborate mostly with members of the AI Tech team, within the Platform Infrastructure group, using... 
    Suggested
    Permanent employment
    Full time
    Local area

    Bain & Company

    Boston, MA
    3 days ago
  • $240k - $293k

     ...efficient, high-quality, and consistent service, incorporating AI tools and outcome-focused strategies across the product lifecycle...  ...drives sales, delivery, and client outcomes across product engineering and cloud modernization, supporting clients as they build, scale... 
    Suggested
    Temporary work
    Local area

    Slalom

    Boston, MA
    2 days ago
  • $112.5k - $202.5k

     ...quickly to solve problems? Join our Network Engineering team The Network Engineering team...  ..., configuring and maintaining network infrastructure in support of Akamai's global platform...  ...'s biggest moments without a glitch. AI : Enabling our customers to build, secure... 
    Work experience placement
    Work at office

    Akamai

    Cambridge, MA
    1 day ago
  • $65.7k - $118.3k

     ...improve our operational efficiency? Join our critical Network Infrastructure Engineering team Our Network Infrastructure Engineering Team takes...  ...: Scaling the world's biggest moments without a glitch. AI : Enabling our customers to build, secure, and scale AI apps... 
    Work experience placement
    Work at office

    Akamai

    Boston, MA
    5 days ago
  • Senior IT Recruitment Consultant - Greater Boston AI & Systems Engineer - Legal Tech | Boston, MA (Hybrid - 3 days)| Law Firm A prestigious...  ...: 5-7 years of experience in IT systems engineering or infrastructure, ideally with recent AI project exposure Deep familiarity... 
    Full time

    Franklin Fitch

    Boston, MA
    2 days ago
  •  ...which every person has a personalized, AI-enabled doctor always in their pocket. To...  ...the line between "application code" and "infrastructure" keeps blurring and we need someone who...  ...re looking for a Backend Infrastructure Engineer who's equal parts software engineer and... 
    Local area
    Immediate start
    Remote work

    Counsel Health

    Boston, MA
    2 days ago
  • $100k - $300k

    Watertown, MAR&D - Software /Full-time /On-siteRobotics Infrastructure Engineer: Systems, Infrastructure & ReliabilityThe CompanyWe believe general...  ..., and researchers who want to be part of that loop.As an AI robotics company that deploys its inventions directly into the... 
    Full time
    Work at office
    Shift work
    Night shift

    Tutor Intelligence

    Watertown, MA
    4 days ago
  • $127k - $160.55k

     ...measurable results for clients.At Zelis, AI is woven into the fabric of how we work....  ....Position OverviewThe Principal Network Engineer will be responsible for the design,...  ...of the organization's enterprise network infrastructure across on-premises, cloud, and hybrid environments... 
    Full time
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Zelis

    Boston, MA
    3 days ago
  • $95k - $145k

    Infrastructure / Data Center Operations Engineer SimSpace serves as an AI Proving Ground where organizations can confidently train, test, and outmaneuver adversaries in any environment. Trusted by allied governments, militaries, enterprises, and research institutions worldwide... 
    For subcontractor
    Local area
    Remote work
    Worldwide
    Flexible hours

    SimSpace Corporation

    Boston, MA
    2 days ago
  •  ...ready to innovate workflow. SUMMARY The Senior Solutions Engineer, Google Cloud & Maps Platform supports the Spatial Intelligence...  ...with emphasis on BigQuery, data pipelines, and Gemini or Vertex AI where the use case calls for it. Model consumption and unit economics... 
    Contract work
    Live in
    Local area
    Remote work

    Sanborn Map Company

    Boston, MA
    a month ago
  • Sr. Network Engineer - Juniper Operating Systems (Houston, TX or Westford, MA)This role has been designed as ‘’Onsite’ with an expectation...  ...-managed networking, SD-WAN, security-adjacent workflows, and AI-driven support experiences.The ideal candidate brings deep hands... 
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hewlett Packard Enterprise

    Boston, MA
    2 days ago
  • $245k - $275k

     .... By weaving together advances in cloud infrastructure, automation and analytics, and software...  ...AHEAD. The AHEAD Senior Specialty Solutions Engineer-Network, will be focused on the core...  ...benefits for additional details. Use of AI:We may use artificial intelligence (AI)... 
    Full time
    Work at office

    AHEAD

    Boston, MA
    3 days ago
  •  ...Somerville, MA   Our client is seeking a highly skilled Senior Infrastructure Engineer Consultant to support a large-scale infrastructure...  ...administrative tasks and improve operational efficiency. Leverage AI-assisted scripting tools to accelerate automation efforts... 
    Hourly pay
    Local area
    3 days per week

    Eliassen Group

    Somerville, MA
    14 days ago
  •  ...Description Do you have a passion for optical network infrastructure and cutting edge DWDM technologies? Join our DWDM Network Engineering Team! Our DWDM network engineering team takes...  ...for production network issues. Leveraging AI-driven tools and scripting to automate... 
    Work at office
    Worldwide

    Akamai Technologies

    Cambridge, MA
    6 days ago
  •  ...position: FEDITC is seeking a Network Engineer II to join our team in the Texarkana, TX...  ...Mobile Radio System hardware, software, and infrastructure components, including application...  ...applicants for employment. We do not employ AI tools in our decision-making processes.... 
    Contract work
    For contractors
    For subcontractor
    Local area
    Remote work
    Worldwide

    Federal IT Consulting (FEDITC)

    Boston, MA
    6 days ago
  •  ...InBev, backed by tier-1 investors in Physical AI.  We’re a polymathic team of 50 in Somerville: hardware and software engineers, chemists, AI researchers, factory...  ...grows and scales, we are excited for a ML Infrastructure Engineer to join the team! We are looking... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Laminar

    Somerville, MA
    3 days ago
  •  ...Job Summary for Network Engineer III Senior (Remote, 3 Months): - Minimum 5 years of experience in Network Security, especially...  ...duration. - Nice-to-have: Experience with automation, scripting, or AI-assisted security operations. - Nice-to-have: Experience with... 
    Work at office
    Remote work

    Expert Technology Services

    Quincy, MA
    2 days ago
  •  ...of network security best practices and Palo Alto Networks platforms.Experience with backbone infrastructure (routing, switching, SD-WAN) and CI/CD automation.Familiarity with AI/ML applications for network optimization is a plus.Excellent problem-solving and leadership... 

    Diverse Lynx

    Boston, MA
    5 days ago
  • $90k - $210k

     ...specially selected team of scientists, engineers, and software developers to deliver best...  ...various projects that provide critical AI, machine learning, simulation, situational...  ...using Python  Develop and maintain cloud infrastructure  oversee the deployment of cloud... 
    Full time

    Morse Corp

    Cambridge, MA
    more than 2 months ago
  •  ...never stand still. Our recent innovations include generative AI capabilities, machine learning insights, and advanced...  ...The Opportunity Northern Light is seeking a Senior Linux Infrastructure Engineer to take hands-on ownership of the Linux infrastructure behind... 
    Full time
    Casual work
    Work at office
    Local area
    Remote work
    Work from home
    Visa sponsorship

    Northern Light

    Somerville, MA
    4 days ago
  • $114.6k - $234.6k

     ...standards and architectures. Mentors junior engineers, provides input on team decisions, and...  ...attack surface. Automation, CI/CD, and Infrastructure as Code: • Develop reusable Terraform...  ...innovations to life-saving care. And with AI embedded across our products and... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Boston, MA
    5 days ago
  • $140k - $200k

     ...around the globe work on Speechify in a 100% distributed setting – Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups like... 
    Remote job
    Full time
    Work at office

    Speechify

    Boston, MA
    more than 2 months ago
  • $86k - $146k

    As a Software Engineer reporting to the Senior Director of Software Engineering, you'll play a key role in building and improving the Nasdaq...  ...with confidence. Explore emerging technologies — including AI-assisted development and intelligent automation — and bring ideas... 
    Full time
    Temporary work
    H1b
    Work at office
    Worldwide
    Home office
    Flexible hours
    3 days per week

    NASDAQ OMX

    Boston, MA
    1 day ago
  • $150k - $215k

     ...platform provide personalized insights that help individuals optimize their health, fitness, and recovery. As a Senior Software Engineer on the AI team, you will play a key role in building and scaling the systems that power WHOOP AI’s internal and member-facing... 
    Full time
    Work at office
    Relocation

    Whoop

    Boston, MA
    more than 2 months ago
  • $125k - $175k

     ...optimize their health, fitness, and recovery. As a Software Engineer II (Frontend), you will play a key role in shaping and delivering...  ...concept through production, leveraging modern web technologies and AI-assisted development workflows to rapidly prototype, iterate,... 
    Full time
    Work at office
    Relocation

    Whoop

    Boston, MA
    more than 2 months ago
  • $118.7k - $197.9k

     ...possess these skills: ·Work with customer engineers to evaluate technical requirements and...  ...·Support application integration, cloud infrastructure, networking, security, and migration activities...  ...clear guidance to others The Team AI & Engineering leverages cutting-edge... 
    Local area

    Deloitte

    Boston, MA
    2 days ago
  • $154.39k - $247.02k

     ...ownership and drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter.AI Infrastructure Engineer, Corporate AI TeamTeam & Role OverviewAxon’s Corporate AI Team sits within Business Technology and builds internal-facing... 
    Work experience placement

    Axon

    Boston, MA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI HPC Infrastructure Engineer. Be the first to apply!