Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - GPU reliability

$200k - $300k
Full-time

Hudson River Trading


Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale.

Responsibilities

This role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:


  • Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configuration

  • Troubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issues

  • Working across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently 

  • Analyzing GPU job statistics to identify trends and areas for improvement

Qualifications

Required:


  • BS and/or MS in computer science or a related field

  • 2+ years of relevant experience, including programming in Python and managing GPUs

  • Experience using automation to solve problems and improve process efficiency

  • Experience working with, troubleshooting, tuning, and deploying various types of GPU hardware

  • Strong grasp of computer science fundamentals and software design patterns

  • Solid understanding of Linux/UNIX operating systems 

  • Familiarity with open-source software

  • Ability to debug and analyze problems quickly

  • Skilled at balancing multiple tasks while maintaining meticulous attention to detail

  • Ability to operate effectively as a team player and also work independently 

  • Ability to learn at a fast pace and apply new skills effectively

Preferred:


  • Understanding of Debian operating system

  • Familiarity with systems configuration management and monitoring technologies 

  • Familiarity with continuous integration and continuous deployment tools and processes

  • Understanding of networking protocols

The estimated base salary range for this position is 200,000 to 300,000 USD per year (or local equivalent). The base pay offered may vary depending on multiple individualized factors, including location, job-related knowledge, skills, and experience. 

This role will also be eligible for discretionary performance-based bonuses and a competitive benefits package which includes medical, dental, vision, basic life insurance, and enrollment in our company’s retirement savings plans. Employees will receive sick and parental leave, as well as other paid time off (including 20 vacation days and 10 paid holidays in the US). Please note that benefits and time off policies will vary across non-US locations.

Culture

Hudson River Trading (HRT) brings a scientific approach to trading financial products. We have built one of the world's most sophisticated computing environments for research and development. Our researchers are at the forefront of innovation in the world of algorithmic trading.

At HRT we welcome a variety of expertise: mathematics and computer science, physics and engineering, media and tech. We’re a community of self-starters who are motivated by the excitement of being at the cutting edge of automation in every part of our organization—from trading, to business operations, to recruiting and beyond. We value openness and transparency, and celebrate great ideas from HRT veterans and new hires alike. At HRT we’re friends and colleagues – whether we are sharing a meal, playing the latest board game, or writing elegant code. We embrace a culture of togetherness that extends far beyond the walls of our office.

Feel like you belong at HRT? Our goal is to find the best people and bring them together to do great work in a place where everyone is valued. HRT is proud of our diverse staff; we have offices all over the globe and benefit from our varied and unique perspectives. HRT is an equal opportunity employer; so whoever you are we’d love to get to know you.

Please be advised: Use of AI tools during interviews or assessments is strictly prohibited, unless otherwise instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that AI assistance was used during any stage of the hiring process, we reserve the right to immediately disqualify your candidacy or rescind any job offers extended.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer - GPU reliability in Remote vacancy
  • $250k

     ...AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference...  ...States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... 
    Suggested
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...Join the engineering teams that bring OpenAI’s ideas safely to the world...  ...that they are performant and reliable. You will work in a deeply...  ...-functional teams, including software engineers, product managers,...  ...the platform for CPU/storage, GPU, and network lifecycle management... 
    Suggested
    Full time
    Work experience placement
    Relocation package

    OpenAI

    Remote
    1 day ago
  • $152k - $241.5k

     ...Performance Computing and Visualization. The GPU, our invention, serves as the visual...  ...looking for highly motivated Senior Software Engineers to join our Fabric Networking team with...  ...NVLink Rack-Scale Systems Stability & Reliability. In this role, you will partner closely... 
    Suggested
    Full time
    Remote work

    Nvidia

    Illinois
    3 days ago
  •  ...teams spanning hardware and software. Speed and scale are our key...  ...frontier forward. The Production Engineering Team Examples of key...  ...00s of GWs: at our scale, a GPU failure isn't a ticket. It's...  ...depend on. Own end-to-end reliability, scalability, and operation of... 
    Suggested
    Local area

    Fluidstack

    San Francisco, CA
    4 days ago
  • $109k - $160k

     ...Learn more at  . About the role A Software Engineer contributes to the design,...  ...role focuses on improving the efficiency, reliability, and scalability of systems that power...  ...product, and hardware teams to evolve our GPU performance testing platform to ensure... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Remote
    1 day ago
  • $193.8k - $285k

    About the TeamThe Reliability Platform role is a key pillar of DoorDash...  ...and repetitive tasks. We use software and agents to “keep the...  ...!About the RoleAs a Software Engineer on the Reliability Platform team...  ...Kafka topics, Databases, CPU/GPU Pools, Service Scaffolding, etc... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    2 days ago
  • $125k - $145k

     ...developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER (FLIGHT RELIAIBLITY) The Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate... 
    Permanent employment
    Full time
    Temporary work
    Remote work
    Worldwide
    Weekend work

    Spacex

    Remote
    1 day ago
  • $150k - $176k

     ...Checkr is recognized on Forbes Cloud 100 2025 List and is a Y Combinator 2024 Breakthrough Company . As a Software Engineer II on the Site Reliability Engineering team within the Platform Engineering group at Checkr, you will identify reliability challenges impacting... 
    Full time
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours
    3 days per week

    Checkr

    San Francisco, CA
    1 day ago
  • $204k - $259k

     ...responsible for running the fully autonomous vehicle’s software stack. To achieve our mission, we architect and...  ...hybrid role, you will report to a Senior Software Engineer. You will: Develop high-performance GPU primitives and abstractions to enable Waymo to scale... 
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  •  ...need you, a highly experienced Senior Software Engineer, to join the technical team located in...  ...codebase for enhance availability, increased reliability, and an improved user experience-...  ...Extensive hands-on experience with DIO/DAQ and GPU/GPGPU-Demonstrable experience in all... 
    Visa sponsorship

    Technology Navigators

    Austin, TX
    9 hours ago
  • $200k - $300k

     ...abundance for all. About the Team Our team owns the reliability and testing plan for every software system that runs on the robot or talks to it. That...  ...customers can trust their robots, and whether our engineers can move fast, comes down to this layer. Key... 
    Full time
    Temporary work
    Local area
    Work from home
    Flexible hours

    1x

    Remote
    1 day ago
  • $170k - $216k

     ...U.S. states. The Planner/Perception Reliability team builds out architectures, tools, and...  ...reliability and is accountable for onboard software health while ensuring high development...  ...you will report to a Staff Software Engineer / Tech Lead Manager. You will: Architect... 
    Full time
    Immediate start
    Remote work

    Waymo

    Remote
    1 day ago
  •  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. At...  ...for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class...  ...network configuration to architect the software fabric that unifies thousands of GPUs... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with...  .../ hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware... 
    Full time
    Remote work

    Nvidia

    Austin, TX
    2 days ago
  • The roleWe’re looking to hire our first data software engineer at Gridmatic! Looking for a startup-minded eng who works closely with our ML...  ...area to both make sure the data is ingested and transformed reliably, and also be able to build the tooling/abstractions to make... 
    Work at office
    Home office
    Flexible hours
    3 days per week

    Gridmatic

    Cupertino, CA
    3 days ago
  • $102.8k - $190.2k

     ...& Online ProductsJob Title:Senior Site Reliability Engineer, Data & AnalyticsRequisition ID:R027436...  ...comfortable operating critical systems, writing software, and building tools that improve...  ...and inference services, including GPU workloads Help define how data and ML services... 
    Full time
    Temporary work
    Part time
    Local area
    Remote work
    Relocation package

    Blizzard Entertainment

    Irvine, CA
    1 day ago
  •  ...of superintelligence. One person, one GPU.If you'd like to build the world's best...  ...centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability...  ..., distributed systems, or production software engineering.Have deep experience operating... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • Role Description We are looking for a Site Reliability Engineer with a strong software development background to join our Scrum team maintaining and improving an Enterprise Generative AI platform. This role focuses on platform stability, performance optimization, uptime... 
    Full time
    Remote work
    Flexible hours

    Corning

    Remote
    4 days ago
  •  ...online without adding hardware, installing software, or changing a line of code. Internet...  ...aspect of Cloudflare's systems: enable more reliable network connectivity for Cloudflare’s...  .... You will work closely with various Engineering teams to translate their requirements... 
    Full time
    Local area

    Cloudflare

    Oklahoma
    1 day ago
  • The Software Reliability Engineer (SRE) will play a critical role in ensuring that our Warehouse Management Software (WMS) runs seamlessly across both automated and manual facilities. This role focuses on investigating, diagnosing, and resolving operational software issues... 
    Full time
    Local area
    Remote work
    Rotating shift

    Lineage Logistics

    Novi, MI
    4 days ago
  • $175k - $250k

     ...customized and developed by our expert team of lawyers, engineers and research scientists. We’ve found product market...  ...Competitive compensation. Role Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability... 
    Full time
    Relocation package

    Harvey

    Remote
    1 day ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...across engineering: You will partner closely with software engineering teams to elevate our system architecture, streamline... 
    Full time

    Zoox

    Remote
    1 day ago
  • $152k - $241.5k

    NVIDIA's GPU Architecture Group is looking for a software engineer to further modernize and scale GPU development. As GPU designs become more complex, our hardware models, testbenches, build scripts, and code generation flows need to keep adapting to this complexity. The... 
    Full time
    Remote work

    Nvidia

    Hillsboro, OR
    1 day ago
  •  ...cloud infrastructure (AWS/GCP/Azure) for performance, cost, and reliability. Improve observability across the platform through...  ...and integrations across systems and tools. Collaborate with engineering, operations, and data teams to understand and support their infrastructure... 
    Full time

    Qdrant

    Remote
    22 days ago
  • $152k - $241.5k

     ...Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern...  ...artificial intelligence.We are looking for highly motivated Senior Software Engineers to work on our GPU Fabric Networking team. Our team develops... 
    Full time
    Remote work

    Nvidia

    Illinois
    2 days ago
  • $80k - $107k

     ...GPU Software Engineer (CUDA) – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Troy, MI
    4 days ago
  •  ...superintelligence. One person, one GPU.If you'd like to build the...  ...day is currently Tuesday.Engineering at Lambda is responsible for...  ...plane services and dataplane software running on SmartNICsDevelop tooling...  ...teams to improve service reliability and deployment workflowsDeploy... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $139k - $257.55k

     ...is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine...  ...energy with the resources of a large software company.What you'll doThis is a role...  ...inference infrastructure — model serving, GPU workloads, language model gateway and... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    1 day ago
  •  ...Solution IT Inc. is looking for Sr. GPU AI Solution Cloud Architect for one of its...  ...Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field...  ...engineering, solutions architecture, site reliability engineering, HPC, or a similar technical... 
    Work experience placement
    Immediate start
    Remote work

    SolutionIT

    Santa Clara, CA
    1 day ago
  • $267k - $356k

     ...superintelligence. One person, one GPU.If you'd like to build the...  ...Tuesday.Lambda's Storage Engineering team is the backbone behind our...  ...in the industry, which means reliability and performance aren't just...  ...operating behind Lambda's own software-defined data plane.Build and... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!