Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Inference Optimization Intern - Performance Modeling

Institute of Foundation Models

Job Description

Job Description

About the Institute of Foundation Models

The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting-edge foundation models while pushing the limits of high-performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real-world impact.

As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures. This internship provides hands-on experience in low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.

Key Responsibilities

This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.

Responsibilities include:

  • Develop analytical performance models for GPU kernels and inference workloads.

  • Build and validate a simulator to estimate theoretical hardware performance limits.

  • Compare measured kernel performance against architectural peak throughput.

  • Identify performance bottlenecks in compute, memory, communication, and scheduling.

  • Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.

  • Investigate PTX and SASS code generation to understand low-level execution behavior.

  • Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.

  • Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.

  • Design profiling methodologies for Hopper and Blackwell architectures.

  • Document findings and provide actionable recommendations for performance improvements.

Academic Qualifications

Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.

Preferred Qualifications

  • Experience with CUDA programming and GPU kernel development.

  • Understanding of NVIDIA GPU architecture and memory hierarchy.

  • Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.

  • Knowledge of PTX, SASS, and low-level GPU execution.

  • Experience optimizing CUDA kernels for throughput and latency.

  • Understanding of roofline analysis, performance modeling, and hardware utilization metrics.

  • Experience with deep learning frameworks such as PyTorch or TensorFlow.

  • Strong programming skills in C++, CUDA, and Python.

Desired Skills

  • Performance engineering mindset.

  • Strong analytical and debugging abilities.

  • Interest in AI systems, inference optimization, and hardware-software co-design.

  • Ability to work independently on research and engineering challenges.

  • Excellent written and verbal communication skills.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Inference Optimization Intern - Performance Modeling in Sunnyvale, CA vacancy
  • $195.2k - $361.2k

     ...by design. Small, efficient models run directly on the user's...  ...hardware people actually own. You optimize inference engines (llama.cpp, vLLM)...  ...tiers and publish honest performance comparisonsUpstream fixes...  .... You will develop:The internals of modern inference engines... 
    Internship
    Performance
    Full time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    1 day ago
  • $193.3k - $261.5k

     ...machine learning accelerators. Join us to optimize the latest models to run really fast on the Trainium...  ...Development Engineer on the Inference Model Enablement team, you will onboard...  ...job responsibilities* Deliver high-performance models using distributed inference libraries... 
    Internship
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $165.2k - $223.6k

     ...enabling unparalleled ML inference and training performance.The Inference Enablement and...  ...running a wide range of models and supporting novel architecture...  ...unit is fine tuned for optimal performance for our...  ...review, and communicate with internal and external stakeholders.... 
    Internship
    Performance
    Work experience placement
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $184k - $287.5k

     ...revolution! The Algorithmic Model Optimization Team specifically focuses...  ...diffusion models for maximal inference efficiency using techniques...  ...Our software is used both internally across NVIDIA and externally...  ...and improving high-performance kernel implementations in CUDA... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    22 hours ago
  •  ...is offering an entry-level engineer or intern position in Santa Clara, California, focused on optimizing and deploying multimodal models for autonomous driving. Candidates will...  ...tools in C++ and Python, and analyze model performance metrics. Ideal applicants will have a... 
    Internship
    Performance

    XPENG

    Santa Clara, CA
    5 days ago
  •  ...located in Santa Clara, California, is seeking a talented individual to design and optimize software for large-scale AI models in modern vehicles. The role involves improving system performance and reducing latency on embedded platforms. The ideal candidate will have a... 
    Internship
    Performance

    XPENG

    Santa Clara, CA
    5 days ago
  • $193.3k - $261.5k

     ...looking for a Senior Inference Engineer to own inference...  ...production — shaping model architecture so it is...  ...arelocked in• Implement and optimize the inference path for...  ...Develop and tune high-performance kernels for critical...  .../multimodal serving internals (e.g., vLLM, TensorRT-... 
    Internship
    Performance
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    1 day ago
  •  ...is seeking an engineering leader to build and scale the Inference Model Scaling organization. You will define the technical...  ...on Cerebras hardware, leading ML model compilation and optimization and high-performance kernel development. Work across compiler, runtime, cloud... 
    Performance

    Cerebras

    Sunnyvale, CA
    2 days ago
  • $38 - $94 per hour

     ...level scheduling Power, performance, and energy-efficiency...  ...LLM training and inference Systems/AI algorithms...  ...hardware design, code optimization, etc.) ML for EDA Programming...  ...and programming models for parallel computing...  ...hourly rate for our interns is 38 USD - 94 USD. You... 
    Internship
    Performance
    Hourly pay
    Summer internship

    Nvidia Corporation in

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis...  ...benchmarking infrastructure, model-level performance projection tooling... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $38 - $94 per hour

     ...level scheduling Power, performance, and energy‑efficiency...  ...LLM training and inference Systems/AI algorithms...  ...hardware design, code optimization, etc.) ML for EDA Programming...  ...and programming models for parallel computing...  ...The hourly rate for our interns is 38 USD - 94 USD. You... 
    Internship
    Performance
    Hourly pay
    Summer internship

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  •  ...are heavily focused on inference . Backed by hundreds...  ...Architecture interns to join our team and contribute...  ...on developing and optimizing compute architectures...  ...that deliver exceptional performance and efficiency for inference...  ...and performance modeling over the course of your... 
    Internship
    Performance
    Summer internship
    Work at office
    Relocation

    Etched

    San Jose, CA
    a month ago
  • $207k - $300k

    Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically...  ...complex model inference performance bottlenecks across the stack....  ...(e.g., PyTorch profiler) and internal trace analysis tools.Experience... 
    Performance

    Google

    Mountain View, CA
    3 days ago
  • $212.7k - $287.7k

     ...scalemachine learning accelerators. Join us to optimize LLMs to run really fast on the Trainium hardware.As an SDM for the LLM Inference Model Enablement team, you will lead a team of...  ...in LLM model architectures, model performance optimizations, and inference techniques,... 
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  •  ...build the Customer World Model (CWM) : a continuously...  ...of graduate research interns to help build the...  ...Preference and goal inference Learning when intervention...  ...predict real-world performance, trust, and customer value...  ...deep learning, optimization, representation learning... 
    Internship
    Performance
    Remote job
    Local area
    Flexible hours

    Block

    San Jose, CA
    4 hours ago
  • $182.5k - $260.5k

     ...visibility and control without performance trade-offs.At Netskope,...  ...Scientist, you own the inference and optimization layer that makes AI in agentic...  ...fine-tune and evaluate models, push latency and throughput...  ...grasp of transformer internals and the levers that move real... 
    Performance

    Netskope

    Santa Clara, CA
    2 days ago
  •  ...generation algorithms for optimization, machine learning, and...  ...and large language models. The Algorithms &...  ...computing architectures. Interns will contribute to...  ...decoding algorithms and inference methods for next‑generation...  ...to improve decoding performance and scalability.... 
    Internship
    Performance
    Flexible hours

    NTT Research Inc.

    Sunnyvale, CA
    5 days ago
  • $117.7k - $221.4k

     ...depends not only on stronger models, but also on better infrastructure...  ...-aware approach that first performs the cheapest reusable work,...  ..., speed, and cost instead of optimizing any one of them in isolation....  ..., featurization, and inference foundations that power scalable... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  •  ...machine learning, and smart connectivity. Key Responsibilities Design and optimize software for deploying large-scale AI models in production vehicles. Profile and improve inference performance across compute, memory, and I/O systems. Reduce latency and improve power... 
    Internship
    Performance

    Xpengmotors

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

     ...accelerating it. The TensorRT inference platform is the backbone of...  ...cutting-edge deep learning models on every NVIDIA GPU. With demand...  ...the global standard for AI performance. What You'll Be Doing: Lead...  ...kernel development, runtime optimizations, and frameworks for LLM... 
    Performance

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  •  ...by design. Small, efficient models run directly on the user's...  ...hardware people actually own. You optimize inference engines (llama.cpp, vLLM)...  ...tiers and publish honest performance comparisons Upstream fixes...  ...’ll learn / grow into The internals of modern inference engines... 
    Performance
    Local area
    Shift work

    Intel

    Santa Clara, CA
    5 days ago
  •  ...Vision‑Language‑Action (VLA) models and foundation models are...  ...entry‑level engineer or intern to support the optimization and deployment of...  ...model deployment, and edge inference for real‑world autonomous...  ...systems. Help analyze model performance, memory usage, latency,... 
    Internship
    Performance

    Xpengmotors

    Santa Clara, CA
    5 days ago
  •  ...‑leading training and inference speeds, empowering machine...  ...to support high‑performance, low‑latency inference...  ...inference workloads. Optimize resource allocation and...  ...Debug issues related to model deployment, container...  ...ensuring clarity for internal teams and external customers... 
    Internship
    Performance
    Full time
    Part time

    Cerebras Systems

    Sunnyvale, CA
    5 days ago
  •  ...deliver industry-leading training and inference speeds; over 10 times faster than...  ...Cerebras works with the leading model labs, global enterprises, and...  ...model transformation pipeline, graph optimization infrastructure, high-performance kernel enablement, and runtime integration... 
    Performance

    Cerebras

    Sunnyvale, CA
    2 days ago
  • $250k - $350k

     ...RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In...  ..., GPU parallelism, advanced model deployment, and video generation technologies...  ...at scale.You will design and optimize inference pipelines, implement... 
    Performance
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    2 days ago
  •  ...flexible, partner-led business model, Nuro is working toward...  ...a software engineering intern, you will work closely...  ...a reliable and high-performance platform that allows our...  ...development in Nuro and optimize on-cloud training and onboard inference. Our solutions include a... 
    Internship
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    4 days ago
  • $160.5k - $240.7k

     ...developersoptimizeand deploy machine learning models on edge and mobile hardware....  ...supportcutting-edgemodel optimization workflows — pushing the...  ..., or similar families) for inference optimization Familiarity...  ...Experience with C++ for performance-critical components is a... 
    Performance
    Work experience placement
    Immediate start
    Work from home

    Socket.dev

    Santa Clara, CA
    2 days ago
  •  ...are heavily focused on inference. Backed by hundreds of...  ...'27, and Summer '27 interns. This role requires business...  ...gaps, and driving performance improvements across...  ...identifying opportunities for optimization. You may be a good...  ...and on increasing model sizes. Our addressable... 
    Internship
    Performance
    Summer internship
    Work at office
    Relocation
    Shift work

    The Consensus

    San Jose, CA
    5 days ago
  • $120k - $250k

     ...develop proprietary AI models to transform...  ...and beyond) into high-performance, AI-ready data engines...  ...developing SOTA visual models, optimizing AI-driven data pipelines...  ...systems, and batch inference at an internet scale,...  ...Software Engineer: Fullstack Intern Opportunities for... 
    Internship
    Performance
    Full time
    Flexible hours

    Orbifold AI

    Los Altos, CA
    1 day ago
  • $45 - $65 per hour

     ...Internship experience from some of our TRI interns! You’ll be joining a multidisciplinary...  ...on developing a world foundation model for driving—a unified, transferable representation...  ...and, time permitting, on high-performance autonomous driving hardware. Present research... 
    Internship
    Performance
    Hourly pay
    Full time
    Work at office
    Local area
    Shift work

    Toyota Research Institute

    Los Altos, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Inference Optimization Intern - Performance Modeling. Be the first to apply!