Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Inference Intern

Etched

Job Description

Job Description

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

We are seeking talented Fall '26, Spring '27, and Summer '27 Inference Architecture interns to join our team and contribute to the design of next-generation AI accelerators. This role focuses on developing and optimizing compute architectures that deliver exceptional performance and efficiency for inference workloads. You will work on cutting-edge architectural problems and performance modeling over the course of your internship.

Key responsibilities

  • Support porting state-of-the-art models to our architecture. Help build programming abstractions and testing capabilities to rapidly iterate on model porting.

  • Assist in building, enhancing, and scaling our runtime, including multi-node inference, intra-node execution, state management, and robust error handling.

  • Contribute to optimizing routing and communication layers using our collectives.

  • Utilize performance profiling and debugging tools to identify bottlenecks and correctness issues.

  • Develop and leverage a deep understanding of our architecture to co-design both HW instructions and model architecture operations to maximize model performance

  • Implement high-performance software components for the Model Toolkit

You may be a good fit if you have

  • Progress towards a Bachelor’s, Master’s, or PhD degree in computer science, computer engineering, applied mathematics, or a related field

  • Proficiency in Python, C++

  • Understanding of performance-sensitive or complex distributed software systems, e.g. Linux internals, accelerator architectures (e.g. GPUs, TPUs), Compilers, or high-speed interconnects (e.g. NVLink, InfiniBand).

  • Ported applications to non-standard accelerator hardware or hardware platforms.

  • Deep knowledge of transformer model architectures and/or inference serving stacks (vLLM, SGLang, etc.)

Strong candidates may have some experience with

  • Proficiency in Rust

  • Low-latency, high-performance applications using both kernel-level and user-space networking stacks.

  • Deep understanding of distributed systems concepts, algorithms, and challenges, including consensus protocols, consistency models, and communication patterns.

  • Solid grasp of Transformer architectures, particularly Mixture-of-Experts (MoE).

  • Built applications with extensive SIMD (Single Instruction, Multiple Data) optimizations for performance-critical paths.

  • Familiarity with PyTorch or JAX.

  • Math competitions (AIME, AMC, etc)

We encourage you to apply even if you do not believe you meet every qualification.

Program details

  • 12-week paid internship

  • Generous housing support for those relocating

  • Daily lunch and dinner in our office

  • Based at our office in San Jose, CA

  • Direct mentorship from industry leaders and world-class engineers

  • Opportunity to work on one of the most important problems of our time

For any questions, contact View email address on ziprecruiter.com.

How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.

 

We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Vacancy posted 10 days ago
Similar jobs that could be interesting for youBased on the Inference Intern in San Jose, CA vacancy
  • $195.2k - $361.2k

     ...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments...  ...learn / grow intoCuriosity is required. You will develop:The internals of modern inference engines and where the milliseconds... 
    Internship
    Full time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    1 day ago
  •  ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and...  ...industry in history. Job Summary Our electrical engineering interns will work on both hands-on product design and the development... 
    Internship
    Summer internship
    Work at office
    Relocation

    Etched

    San Jose, CA
    17 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack:...  ...out of them at inference time.• Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM... 
    Suggested

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $193.3k - $261.5k

     ...accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.As a Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    17 hours ago
  •  ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and...  ...looking for Summer '26, Fall '26, Spring '27, and Summer '27 interns. You may be a good fit if you have Progress towards... 
    Internship
    Summer work
    Summer internship
    Work at office
    Relocation

    ETCHED LLC

    San Jose, CA
    3 days ago
  •  ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and...  ...are looking for Summer '26, Fall '26, Spring '27, and Summer '27 interns. Key responsibilities: Interview Scheduling & Logistics: Schedule... 
    Internship
    Summer work
    Summer internship
    Work at office
    Relocation

    The Consensus

    San Jose, CA
    1 day ago
  • $165.2k - $223.6k

     ...popular ML frameworks like PyTorch and JAX enabling unparalleled ML inference and training performance.The Inference Enablement and...  ...participate in design discussions, code review, and communicate with internal and external stakeholders. You will work cross-functionally to... 
    Internship
    Work experience placement
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  •  ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and...  ...strategies, and supply chain management, we are looking for a Finance intern to tackle strategic financial challenges and execute on our day... 
    Internship
    Summer internship
    Work at office
    Relocation

    ETCHED LLC

    San Jose, CA
    3 days ago
  • $182.5k - $260.5k

     ...roleAs a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast...  ...C++ for low-level interop is a plus.Solid grasp of transformer internals and the levers that move real inference performance and cost:... 

    Netskope

    Santa Clara, CA
    2 days ago
  • $21.69 per hour

     ...of Santa Clara and Mission College are partnering for a Student Internship Program for Summer 2026. The City is looking for Student Intern II's to perform work in various departments Citywide. As part of this program, Student Intern II's are required to be currently... 
    Internship
    Hourly pay
    Summer work
    Summer internship

    GovernmentJobs.com

    Santa Clara, CA
    1 day ago
  • $20 - $24 per hour

     ...successful InsurTech companies, certified as a Great Place to Work®! We’re looking for an eager and motivated Office Administrator Intern to join our team and gain hands-on experience in a fast-paced, professional environment. This role provides exposure to day-to-day... 
    Internship
    Hourly pay
    Full time
    Part time
    Work experience placement
    Summer internship
    Work at office
    Local area
    Visa sponsorship
    Monday to Thursday
    Flexible hours

    Visitorscoverage Inc.

    Santa Clara, CA
    17 hours ago
  • $25 - $35 per hour

     ...opportunity. Program dates are: ~ May 26, 2026 - July 31, 2026 ~ June 15, 2026 - August 21, 2026 ~*We will not be able to accommodate interns outside of these two program dates. This role will be based at Archer HQ in San Jose, CA Relocation and housing will not... 
    Internship
    Hourly pay
    Full time
    Summer internship
    Local area
    Relocation

    Archer

    San Jose, CA
    17 hours ago
  • $152k - $241.5k

     ...applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise will help shape the performance, functionality,... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $60 per hour

     ...Ellis Technologies, Inc. is seeking a Researcher Intern for Summer 2026 in San Jose, California. This role involves building impactful features and conducting cutting-edge research in vision and graphics. PhD candidates in Computer Science or related fields with a strong... 
    Internship
    Hourly pay
    Summer work

    Ellis Technologies, Inc.

    San Jose, CA
    17 hours ago
  • $124k - $195.5k

    NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks...  ....Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...optimizations for next-generation NVIDIA GPUs.Advance the state-of-the-art: Solve complex compilation problems for AI workloads (both inference and training) and successfully transition these breakthroughs into enterprise and consumer products.Collaborate on hardware/... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose,... 

    AMD

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...in a fast-growing technology company that leads the AI revolution.What you will be doing:Implement language and multimodal model inference as part of NVIDIA Inference Microservices (NIMs).Contribute new features, fix bugs and deliver production code to TRT-LLM, NVIDIA’... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and...  ...history. Job Summary Our mechanical/thermal engineering interns will work within the mechanical and thermal engineering... 
    Internship
    Summer work
    Summer internship
    Work at office
    Relocation

    Etched

    San Jose, CA
    17 days ago
  •  ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and...  ...team. We are looking for Fall '26, Spring '27, and Summer '27 interns. This role requires business acumen, analytical skills, strong... 
    Internship
    Summer internship
    Work at office
    Relocation
    Shift work

    Etched

    San Jose, CA
    10 days ago
  • $152k - $241.5k

     ...are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery...  ...ML layers affect execution timeFamiliarity with PyTorch internals (custom ops, autograd, export) or equivalent frameworkExperience... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...problems. We're seeking talented and motivated engineers to join our TensorRT team in developing the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the TensorRT team, you will be responsible for designing and... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a state-of-the-art inference framework for accelerating Deep Learning models, especially Large Language Models, on NVIDIA... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...accelerating deep learning models, and enabling RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node systems. You will collaborate across internal GPU software teams and engage with open-source communities to integrate and optimize... 

    AMD

    Santa Clara, CA
    6 days ago
  • $163k - $253k

     ...technical presentations and documentation. • Occasional domestic and international travel (<10%).What You BringPh.D. in Computer Science,...  ..., including LLMs, DLRMs, and large-scale training and inference systems. • Demonstrated ability to translate workload characteristics... 
    Work at office
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    2 days ago
  • $148.7k - $201.2k

     ...organization, the team is tasked with developing the next generation of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are looking for an experienced System development engineer to drive development for new EC2 machine learning... 
    Internship
    Local area
    Work from home
    Worldwide
    Flexible hours

    Amazon Development Center U.S., Inc.

    Santa Clara, CA
    17 hours ago
  •  ...and decode workloads. Our first products are heavily focused on inference . Backed by hundreds of millions from top-tier investors and...  ...industry in history. Job Summary We’re hiring for a GTM intern - someone who will help us build the operational backbone of... 
    Internship
    Summer internship
    Work at office
    Relocation

    Etched

    San Jose, CA
    10 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Inference Intern. Be the first to apply!