Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Systems Engineer

Career Techniques

As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes. The role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention. Responsibilities: Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation. Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them. Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups. Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing. Own infrastructure projects end to end, from scope and design through implementation and long-term support. Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues. Qualifications: 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments. Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it. Hands-on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance. Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not. Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on. Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef. Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer. Clear communication. You will work daily with researchers, engineers, and vendors. Comp: 200-300K + Bonus #J-18808-Ljbffr Career Techniques

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the GPU Systems Engineer in New York, NY vacancy
  • $150k - $300k

    Hudson River Trading (HRT) is looking for GPU Systems Engineers to help scale and evolve our exceptionally sophisticated HPC/AI research environment. Joining our Research and Development team, you will collaborate with experts responsible for the compute, storage, operating... 
    Suggested
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    4 days ago
  • $200k - $300k

     ...the world’s best systematic trading and engineering talent. We empower portfolio managers to...  ...responsible for the compute, storage, operating systems, and automation behind that work at...  ...of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes. The... 
    Suggested
    Work experience placement
    Casual work
    Work at office

    Socket

    New York, NY
    3 days ago
  •  ...skilled professional in New York to design and operate large-scale GPU infrastructure for model inference and reinforcement learning. The role demands several years of experience in deploying GPU systems, optimizing model performance, and working with frameworks like... 
    Suggested

    Reflection

    New York, NY
    1 day ago
  •  ...focusing on the compute, storage, and automation behind large-scale trading workloads. You will join engineers responsible for hundreds of petabytes of storage and thousands of GPU-accelerated nodes, shaping architectures for AI clusters and performance profiling.... 
    Suggested

    Tower Research Capital

    New York, NY
    4 days ago
  •  ...Senior GPU Systems / AI Infrastructure Engineer (NYC) Location: New York City (Hybrid / On-site preferred) Comp: Competitive + equity (Series A-C / high-growth AI infra) About the Role We’re hiring a senior-level engineer to build and optimise next-generation... 
    Suggested
    Full time
    New York, NY
    more than 2 months ago
  • Tower Research Capital seeks an accomplished engineer to design, deploy, and scale distributed GPU clusters, building out hardware selection, production operations, and monitoring across thousands of nodes. You will diagnose bottlenecks across compute, storage, and network... 

    Socket.dev

    New York, NY
    4 days ago
  • Obsidian is seeking an MLOps Engineer to join their AI lab's cutting-edge GenAI team in New York. This full-time role involves training and...  ...years of experience in ML infrastructure, with skills in custom GPU kernel optimization. This position offers a chance to be at the... 
    Full time

    Obsidian

    New York, NY
    2 days ago
  • $150k - $300k

    Hudson River Trading (HRT) is looking for Systems Engineers to join our growing Research & Development team. This team builds and maintains exceptionally...  ...software, and development tools. We have incredibly large GPU and CPU compute clusters, larger than most national labs. We... 
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Worldwide

    Hudson River Trading

    New York, NY
    4 days ago
  • $165k - $242k

     ...You’ll Do:CoreWeave is seeking a highly skilled and motivated Systems Kernel Engineer to join our HAVOCK Team, reporting into the Manager of...  ...them where applicable (networking, storage, virtualization, GPU/DPU enablement).Stack-Wide Support - Ensure kernel support and... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    4 days ago
  •  ...leading financial services organization is seeking a senior Storage Systems Engineer to support and modernize its core infrastructure environment....  ...environmentAzure infrastructure experienceExposure to AI or GPU-backed environmentsAutomation or scripting experienceFinancial... 

    Madison Davis

    New York, NY
    4 days ago
  • $182k - $242k

     ...workloads and bare metal. We own the operating system, virtualization, runtime, and hardware interfaces that allow thousands of GPU servers to securely execute customer...  ...performance.About the role:As a Senior Software Engineer on HAVOCK's Runtime & Virtualization team, you... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    4 days ago
  •  ...Fluidstack Production Engineering TeamExamples of key exciting problems the team is working on...  ...company, not a hundred scripts.Make the system's view of itself always match reality: integrate...  ...-facing platforms, so every new site and GPU generation lands cleanly from day zero.... 
    Local area

    Fluidstack

    New York, NY
    2 days ago
  •  ...OS / K8s Systems EngineerBaseten powers mission-critical inference for the world's most dynamic...  .... Join us and help build the platform engineers turn to to ship AI products.As an OS / K8...  ...the automation and systems that turn raw GPU hardware into production-ready compute.... 
    Flexible hours

    Baseten

    New York, NY
    3 days ago
  •  ...platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and...  ...to hire someone to help us with Platform Engineering work. We're working on lots of fun problems:Low-level systems development: working with container runtimes, OCI... 

    BEAM inc.

    New York, NY
    5 days ago
  •  ...platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and...  ...to hire someone to help us with Platform Engineering work. We're working on lots of fun problems:Low-level systems development: working with container runtimes, OCI... 

    Beam

    New York, NY
    5 days ago
  • $182k - $242k

     ...publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at .What You'll Do:The Systems Engineering team owns the Linux kernel and host software stack underneath one of the largest GPU fleets in the world. When something breaks at the Kubernetes layer — a pod stuck... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    4 days ago
  • $112.9k - $155.24k

     ...individuals for their own use. SCOPE OF POSITION   The Senior Systems Engineer is responsible for working directly with our customers and...  ...AI/ML fabric architectures (scale-up / scale-out), including GPU/XPU cluster designs and optical I/O requirements   This... 
    Full time
    Work experience placement

    Corning

    New York, NY
    26 days ago
  • DS creates systems that power the next generation of radio spectrum intelligence. We collect...  ...embedding model architectures, custom GPU kernels, and much more. Joining DS means...  ...higher or equivalent experience Electrical Engineering or Physics or similar Nice to Have Experience... 
    Permanent employment
    Temporary work
    Work at office
    Flexible hours

    Distributed Spectrum

    New York, NY
    3 days ago
  •  ...Description ThisWay Global is looking for a Distributed Systems Engineer in a remote role within the United States. ThisWay Global, Inc. is...  ...supporting accelerated deployment timelines and NVIDIA NVL72/GB300 GPU clusters. Amalgamy.ai — AI Orchestration Software: An... 
    Full time
    Remote work

    GrabJobs

    New York, NY
    3 days ago
  • $112.9k - $155.24k

    Senior Systems Engineer Co-Packaged Optics Locations: Santa Clara and Bay Area (West Coast) Company: Corning Requisition Number: 75530 The company...  ...AI/ML fabric architectures (scale-up / scale-out), including GPU/XPU cluster designs and optical I/O requirements This position... 
    Full time
    Work experience placement

    Corning Incorporated

    New York, NY
    3 days ago
  •  ...ofenabling human life on Mars. SR. HIGH PERFORMANCE COMPUTING (HPC) SYSTEMS ENGINEER SpaceX is looking for an HPC Systems Engineer with strong...  ...) Familiarity with large scale AI training Familiarity with GPU usage in a compute cluster and Cuda Experience with... 
    Permanent employment
    Flexible hours
    Weekend work

    SpaceX

    New York, NY
    4 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its...  ...level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. Our team... 
    Full time

    Nvidia

    New York, NY
    1 day ago
  • $105.4k - $124k

     ...career. Try new things, learn new skills and discover what you excel at—all from Day One.Job DescriptionU.S. Bank is seeking a z/OS Systems Programmer that has extensive experience in mainframe technologies with an emphasis on installation/upgrading, tuning, and... 
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    New York, NY
    3 days ago
  •  ...On SiteJob Reference 0000020103Salary Type AnnuallyIndustry Managed Services ProviderSelling Points Advance your career as a Systems Engineer in a dynamic environment. Collaborate on innovative IT solutions and enhance system performance. Gain exposure to cutting-edge... 

    Green Key Resources

    New York, NY
    3 days ago
  • TypeContract Qualifications: Good understanding of application, middleware and database interactions and connectivity Experience supporting Web servers and applications Knowledge of scripting (ex: powershell, Unix Shell Scripting,ansible,puppet) Knowledge of LDAP concepts...
    Work experience placement
    Shift work

    Intelliswift

    New York, NY
    4 days ago
  • $150k - $225k

     ...offices throughout the world. These servers are highly automated and finely tuned to meet the growing demands of our business. Our Systems Engineers are responsible for designing, building, and supporting these servers, working with engineers across our global groups to... 
    Worldwide
    Shift work

    Virtu Financial

    New York, NY
    4 days ago
  •  ...for advertisers to reach deeply engaged audiences.Our TeamThe Decisioning & Optimization engineering team sits within the Ad Serving & Decisioning at Netflix Ads. We own the systems that power real-time ad decisioning, delivering relevant, high-quality ads while... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    New York, NY
    5 days ago
  • $86k - $115k

    Datadog’s Technical Solutions organization includes 1,200+ sales engineers, support engineers, post-sales experts, and solution architects...  ...“TSO”) owns that ecosystem.We manage the full lifecycle of the systems TS depends on: Zendesk, Jira, Confluence, and a growing... 
    Work at office

    Datadog

    New York, NY
    4 days ago
  • $119k - $180k

     ...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-Life...  ...driven networking and automation.What You’ll DoAs an Advisory Systems Engineer (ASE), you are positioned to sell Arista products in... 

    Arista Networks

    New York, NY
    2 days ago
  • $136.8k - $163.5k

     ...to do the most important work of your career, come join us!The Systems & Networking team is responsible for designing and maintaining...  ...comprehensive & secure self service functionalities that lets engineering teams move faster.Design and build our tools for monitoring... 
    Temporary work
    Shift work

    Yext

    New York, NY
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Systems Engineer. Be the first to apply!