Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - GPU reliability

$200k - $300k
Full-time

Hudson River Trading


Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale.

Responsibilities

This role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:


  • Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configuration

  • Troubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issues

  • Working across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently 

  • Analyzing GPU job statistics to identify trends and areas for improvement

Qualifications

Required:


  • BS and/or MS in computer science or a related field

  • 2+ years of relevant experience, including programming in Python and managing GPUs

  • Experience using automation to solve problems and improve process efficiency

  • Experience working with, troubleshooting, tuning, and deploying various types of GPU hardware

  • Strong grasp of computer science fundamentals and software design patterns

  • Solid understanding of Linux/UNIX operating systems 

  • Familiarity with open-source software

  • Ability to debug and analyze problems quickly

  • Skilled at balancing multiple tasks while maintaining meticulous attention to detail

  • Ability to operate effectively as a team player and also work independently 

  • Ability to learn at a fast pace and apply new skills effectively

Preferred:


  • Understanding of Debian operating system

  • Familiarity with systems configuration management and monitoring technologies 

  • Familiarity with continuous integration and continuous deployment tools and processes

  • Understanding of networking protocols

The estimated base salary range for this position is 200,000 to 300,000 USD per year (or local equivalent). The base pay offered may vary depending on multiple individualized factors, including location, job-related knowledge, skills, and experience. 

This role will also be eligible for discretionary performance-based bonuses and a competitive benefits package which includes medical, dental, vision, basic life insurance, and enrollment in our company’s retirement savings plans. Employees will receive sick and parental leave, as well as other paid time off (including 20 vacation days and 10 paid holidays in the US). Please note that benefits and time off policies will vary across non-US locations.

Culture

Hudson River Trading (HRT) brings a scientific approach to trading financial products. We have built one of the world's most sophisticated computing environments for research and development. Our researchers are at the forefront of innovation in the world of algorithmic trading.

At HRT we welcome a variety of expertise: mathematics and computer science, physics and engineering, media and tech. We’re a community of self-starters who are motivated by the excitement of being at the cutting edge of automation in every part of our organization—from trading, to business operations, to recruiting and beyond. We value openness and transparency, and celebrate great ideas from HRT veterans and new hires alike. At HRT we’re friends and colleagues – whether we are sharing a meal, playing the latest board game, or writing elegant code. We embrace a culture of togetherness that extends far beyond the walls of our office.

Feel like you belong at HRT? Our goal is to find the best people and bring them together to do great work in a place where everyone is valued. HRT is proud of our diverse staff; we have offices all over the globe and benefit from our varied and unique perspectives. HRT is an equal opportunity employer; so whoever you are we’d love to get to know you.

Please be advised: Use of AI tools during interviews or assessments is strictly prohibited, unless otherwise instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that AI assistance was used during any stage of the hiring process, we reserve the right to immediately disqualify your candidacy or rescind any job offers extended.

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer - GPU reliability in Remote vacancy
  •  ...Join the engineering teams that bring OpenAI’s ideas safely to the world...  ...that they are performant and reliable. You will work in a deeply...  ...-functional teams, including software engineers, product managers,...  ...the platform for CPU/storage, GPU, and network lifecycle management... 
    Suggested
    Full time
    Work experience placement
    Relocation package

    OpenAI

    Remote
    11 hours ago
  • $175k - $300k

     ...frontier forward. The Production Engineering Team Key exciting problems...  ...10 GW fleet: at our scale, a GPU failure isn't a ticket. It's...  ...tooling depend on. Own end-to-end reliability, scalability, and operation...  ...silicon level, not just the software stack above it. You move... 
    Suggested
    Local area

    FluidStack

    Seattle, WA
    3 days ago
  • $109k - $160k

     ...Learn more at  . About the role A Software Engineer contributes to the design,...  ...role focuses on improving the efficiency, reliability, and scalability of systems that power...  ...product, and hardware teams to evolve our GPU performance testing platform to ensure... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Remote
    11 hours ago
  • $250k

     ...AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference...  ...States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... 
    Suggested
    Permanent employment
    Remote work
    San Francisco, CA
    a month ago
  • $204k - $259k

     ...responsible for running the fully autonomous vehicle’s software stack. To achieve our mission, we architect and...  ...hybrid role, you will report to a Senior Software Engineer. You will: Develop high-performance GPU primitives and abstractions to enable Waymo to scale... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    11 hours ago
  • $170k - $216k

     ...U.S. states. The Planner/Perception Reliability team builds out architectures, tools, and...  ...reliability and is accountable for onboard software health while ensuring high development...  ...you will report to a Staff Software Engineer / Tech Lead Manager. You will: Architect... 
    Full time
    Immediate start
    Remote work

    Waymo

    Remote
    11 hours ago
  • $125k - $145k

     ...developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER (FLIGHT RELIAIBLITY) The Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate... 
    Permanent employment
    Full time
    Temporary work
    Remote work
    Worldwide
    Weekend work

    Spacex

    Remote
    11 hours ago
  • $150k - $176k

     ...Checkr is recognized on Forbes Cloud 100 2025 List and is a Y Combinator 2024 Breakthrough Company . As a Software Engineer II on the Site Reliability Engineering team within the Platform Engineering group at Checkr, you will identify reliability challenges impacting... 
    Full time
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours
    3 days per week

    Checkr

    San Francisco, CA
    11 hours ago
  •  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. At...  ...for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class...  ...network configuration to architect the software fabric that unifies thousands of GPUs... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    11 hours ago
  •  ...online without adding hardware, installing software, or changing a line of code. Internet...  ...aspect of Cloudflare's systems: enable more reliable network connectivity for Cloudflare’s...  .... You will work closely with various Engineering teams to translate their requirements... 
    Full time
    Local area

    Cloudflare

    Oklahoma
    11 hours ago
  • $175k - $250k

     ...customized and developed by our expert team of lawyers, engineers and research scientists. We’ve found product market...  ...Competitive compensation. Role Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability... 
    Full time
    Relocation package

    Harvey

    Remote
    11 hours ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...across engineering: You will partner closely with software engineering teams to elevate our system architecture, streamline... 
    Full time

    Zoox

    Remote
    11 hours ago
  • $102.8k - $190.2k

     ...Online Products Job Title: Senior Site Reliability Engineer, Data & Analytics Requisition ID: R...  ...operating critical systems, writing software, and building tools that improve...  ...pipelines and inference services, including GPU workloads  ~ Help define how... 
    Full time
    Temporary work
    Part time
    Local area
    Remote work
    Relocation package

    Blizzard Entertainment

    Albany, NY
    5 days ago
  •  ...Lead Site Reliability Engineer Bridge Defense is redefining how modern defense technology is delivered...  ..., cleared talent, and mission-ready software to meet evolving defense challenges....  ...physical systems, including high-density GPU servers, networking gear, and storage... 
    Contract work
    Remote work
    Relocation

    Bridge Defense

    Washington DC
    2 days ago
  • $95k - $171k

     ...minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Sacramento, CA
    2 days ago
  • $130k - $165k

     ...Job Title: Senior Software Engineer Company: Snapsheet Job Location: USA, Remote Job Type: Full-time, direct hire Job Department: Technology Team: Site Reliability Engineering About Snapsheet: Snapsheet exists to simplify claims. We leverage our expertise... 
    Full time
    Temporary work
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Snapsheet

    United States
    1 day ago
  • $165k - $225k

     ...Sr. Site Reliability Engineer (SRE) Chicago, IL or Remote Moonlite delivers high-performance...  ...solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation...  ...engineers, network engineers, and software developers. Preferred Qualifications... 
    Remote work
    Flexible hours

    Moonlite AI

    United States
    3 days ago
  •  ...Striveworks Site Reliability Engineer Opportunity Striveworks is a leader in Machine Learning Operations...  ...to cloud-based, on-prem, and hybrid software environments, and to resolve any...  ...Pipelines Administration/development of GPU-enabled servers Interview Process... 
    Remote work

    1872 Consulting

    United States
    1 day ago
  • $170k - $290k

     ...intelligence. This requires a massive, reliable, and performant GPU infrastructure that pushes the...  ...looking for a hands-on, first-principles engineer who is fluent in Linux, comfortable operating...  ...toil. Debug Complex Hardware/Software Failures: Serve as the final... 
    Remote work

    Luma AI

    United States
    3 days ago
  •  ...GPU Server Test Software Engineer This role focuses on GPU server testing diagnostic integration, test automation development, and execute comprehensive test plan and test cases to ensure product quality. Key Responsibilities 1. Test Planning & Execution:... 
    Full time
    Casual work
    Work at office
    Remote work
    Flexible hours

    Employee Magnets

    Remote
    11 hours ago
  • $150k - $200k

     ...Site Reliability Engineer at Triomics (W21) $150K - $200K AI Agents for Oncology EHRs Triomics is building the agentic AI layer for oncology electronic...  ...in customer as well as Triomics cloud environments, with GPU infrastructure serving AI extraction models. We need someone... 
    Full time
    Work at office
    Remote work
    Day shift

    Triomics

    New York, NY
    1 day ago
  • $184k - $356.5k

     ...Developer Tools team and empower engineers throughout the world...  ...of the team that brings new GPU technologies to market with sophisticated...  ...is looking for a senior software engineer to join our efforts...  ...fast, effective, maintainable, reliable and well-documented code.... 
    Full time

    NVIDIA

    Remote
    6 days ago
  • $170k - $216k

     ...developer productivity of onboard engineers to enable the entire...  ...role you will report to a Staff Software Engineer / Tech Lead Manager. You will: Develop reliable, scalable, and maintainable systems...  ...should happen on CPU, GPU, and TPU. Collaborate to resolve... 
    Full time
    Remote work

    Waymo

    Remote
    11 hours ago
  •  ...Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco • Full-Time About Andromeda Andromeda Cluster...  ...response. Nice to Have Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.). Familiarity... 
    Full time
    Remote work

    Andromeda Cluster, Inc

    United States
    4 days ago
  •  ...TENEX Staff Site Reliability Engineer TENEX is an AI-native, automation-first, built-for-scale...  ...years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing...  ...for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization).... 
    Work from home

    TenEx

    Washington DC
    2 days ago
  •  ...and implement Infrastructure Services for a GPU-as-a-Service platform, the full-time remote Full-Stack Software Engineer will build control planes that translate high...  ...metal server provisioning, and ensure system reliability. Key Responsibilities Design, build, and maintain... 
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  • $90k - $190k

     ...Tutor Intelligence. As an AI software company that deploys its inventions...  ....   Full Stack Software Engineer (XR Applications) As a...  ..., you will write performant, GPU-enabled frontend interfaces for...  ...solutions Improve system reliability, scalability, and... 
    Full time
    Work at office

    Tutor Intelligence

    Remote
    11 hours ago
  • $175k - $250k

    Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality...  ...ensuring scalability, performance, and reliability across environments. What You’ll Do...  ...workloads at scale Manage and automate GPU compute clusters using tools such as Python... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  • $166k - $244k

    # Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them... 
    Temporary work

    Epic Games (Portuguese)

    Seattle, WA
    1 day ago
  • $149.4k - $202k

    Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the... 
    Remote work

    Noctua Technology

    Virginia, MN
    16 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!