Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Performance Engineer - GPU & Inference

Modal

Modal is building an infrastructure layer for AI and is seeking strong engineers to optimize ML systems for performance at scale. You will contribute to Modal’s container runtime and open-source projects, pushing language and diffusion models toward higher throughput and lower latency. The role emphasizes working with Torch, TensorRT, CUDA, and NVIDIA GPU architectures to maximize efficiency, while exploring low-level OS foundations to improve performance and reliability. #J-18808-Ljbffr Modal

Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Performance Engineer - GPU & Inference. Be the first to apply!