Software Engineer - GPU reliability
$200k - $300kHudson River Trading
Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale.
Responsibilities
This role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:
- Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configuration
- Troubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issues
- Working across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently
- Analyzing GPU job statistics to identify trends and areas for improvement
Qualifications
Required:
- BS and/or MS in computer science or a related field
- 2+ years of relevant experience, including programming in Python and managing GPUs
- Experience using automation to solve problems and improve process efficiency
- Experience working with, troubleshooting, tuning, and deploying various types of GPU hardware
- Strong grasp of computer science fundamentals and software design patterns
- Solid understanding of Linux/UNIX operating systems
- Familiarity with open-source software
- Ability to debug and analyze problems quickly
- Skilled at balancing multiple tasks while maintaining meticulous attention to detail
- Ability to operate effectively as a team player and also work independently
- Ability to learn at a fast pace and apply new skills effectively
Preferred:
- Understanding of Debian operating system
- Familiarity with systems configuration management and monitoring technologies
- Familiarity with continuous integration and continuous deployment tools and processes
- Understanding of networking protocols
The estimated base salary range for this position is 200,000 to 300,000 USD per year (or local equivalent). The base pay offered may vary depending on multiple individualized factors, including location, job-related knowledge, skills, and experience.
This role will also be eligible for discretionary performance-based bonuses and a competitive benefits package which includes medical, dental, vision, basic life insurance, and enrollment in our company’s retirement savings plans. Employees will receive sick and parental leave, as well as other paid time off (including 20 vacation days and 10 paid holidays in the US). Please note that benefits and time off policies will vary across non-US locations.
Culture
Hudson River Trading (HRT) brings a scientific approach to trading financial products. We have built one of the world's most sophisticated computing environments for research and development. Our researchers are at the forefront of innovation in the world of algorithmic trading.
At HRT we welcome a variety of expertise: mathematics and computer science, physics and engineering, media and tech. We’re a community of self-starters who are motivated by the excitement of being at the cutting edge of automation in every part of our organization—from trading, to business operations, to recruiting and beyond. We value openness and transparency, and celebrate great ideas from HRT veterans and new hires alike. At HRT we’re friends and colleagues – whether we are sharing a meal, playing the latest board game, or writing elegant code. We embrace a culture of togetherness that extends far beyond the walls of our office. Feel like you belong at HRT? Our goal is to find the best people and bring them together to do great work in a place where everyone is valued. HRT is proud of our diverse staff; we have offices all over the globe and benefit from our varied and unique perspectives. HRT is an equal opportunity employer; so whoever you are we’d love to get to know you.Please be advised: Use of AI tools during interviews or assessments is strictly prohibited, unless otherwise instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that AI assistance was used during any stage of the hiring process, we reserve the right to immediately disqualify your candidacy or rescind any job offers extended.
$250k
...AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference... ...States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments...SuggestedFull timeRemote work- ...Join the engineering teams that bring OpenAI’s ideas safely to the world... ...that they are performant and reliable. You will work in a deeply... ...-functional teams, including software engineers, product managers,... ...the platform for CPU/storage, GPU, and network lifecycle management...SuggestedFull timeWork experience placementRelocation package
$152k - $241.5k
...Performance Computing and Visualization. The GPU, our invention, serves as the visual... ...looking for highly motivated Senior Software Engineers to join our Fabric Networking team with... ...NVLink Rack-Scale Systems Stability & Reliability. In this role, you will partner closely...SuggestedFull timeRemote work- ...teams spanning hardware and software. Speed and scale are our key... ...frontier forward. The Production Engineering Team Examples of key... ...00s of GWs: at our scale, a GPU failure isn't a ticket. It's... ...depend on. Own end-to-end reliability, scalability, and operation of...SuggestedLocal area
$109k - $160k
...Learn more at . About the role A Software Engineer contributes to the design,... ...role focuses on improving the efficiency, reliability, and scalability of systems that power... ...product, and hardware teams to evolve our GPU performance testing platform to ensure...SuggestedPermanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$193.8k - $285k
About the TeamThe Reliability Platform role is a key pillar of DoorDash... ...and repetitive tasks. We use software and agents to “keep the... ...!About the RoleAs a Software Engineer on the Reliability Platform team... ...Kafka topics, Databases, CPU/GPU Pools, Service Scaffolding, etc...Hourly payWork at officeLocal areaRemote workFlexible hours$125k - $145k
...developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER (FLIGHT RELIAIBLITY) The Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate...Permanent employmentFull timeTemporary workRemote workWorldwideWeekend work$150k - $176k
...Checkr is recognized on Forbes Cloud 100 2025 List and is a Y Combinator 2024 Breakthrough Company . As a Software Engineer II on the Site Reliability Engineering team within the Platform Engineering group at Checkr, you will identify reliability challenges impacting...Full timeWork at officeLocal areaRemote workRelocationFlexible hours3 days per week$204k - $259k
...responsible for running the fully autonomous vehicle’s software stack. To achieve our mission, we architect and... ...hybrid role, you will report to a Senior Software Engineer. You will: Develop high-performance GPU primitives and abstractions to enable Waymo to scale...Full timeRemote work- ...need you, a highly experienced Senior Software Engineer, to join the technical team located in... ...codebase for enhance availability, increased reliability, and an improved user experience-... ...Extensive hands-on experience with DIO/DAQ and GPU/GPGPU-Demonstrable experience in all...Visa sponsorship
$200k - $300k
...abundance for all. About the Team Our team owns the reliability and testing plan for every software system that runs on the robot or talks to it. That... ...customers can trust their robots, and whether our engineers can move fast, comes down to this layer. Key...Full timeTemporary workLocal areaWork from homeFlexible hours$170k - $216k
...U.S. states. The Planner/Perception Reliability team builds out architectures, tools, and... ...reliability and is accountable for onboard software health while ensuring high development... ...you will report to a Staff Software Engineer / Tech Lead Manager. You will: Architect...Full timeImmediate startRemote work- ...Conviction. Join us and help build the platform engineers turn to to ship AI products. At... ...for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class... ...network configuration to architect the software fabric that unifies thousands of GPUs...Full timeFlexible hours
$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with... .../ hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware...Full timeRemote work- The roleWe’re looking to hire our first data software engineer at Gridmatic! Looking for a startup-minded eng who works closely with our ML... ...area to both make sure the data is ingested and transformed reliably, and also be able to build the tooling/abstractions to make...Work at officeHome officeFlexible hours3 days per week
$102.8k - $190.2k
...& Online ProductsJob Title:Senior Site Reliability Engineer, Data & AnalyticsRequisition ID:R027436... ...comfortable operating critical systems, writing software, and building tools that improve... ...and inference services, including GPU workloads Help define how data and ML services...Full timeTemporary workPart timeLocal areaRemote workRelocation package- ...of superintelligence. One person, one GPU.If you'd like to build the world's best... ...centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability... ..., distributed systems, or production software engineering.Have deep experience operating...Work at officeLocal areaWork from homeFlexible hours
- Role Description We are looking for a Site Reliability Engineer with a strong software development background to join our Scrum team maintaining and improving an Enterprise Generative AI platform. This role focuses on platform stability, performance optimization, uptime...Full timeRemote workFlexible hours
- ...online without adding hardware, installing software, or changing a line of code. Internet... ...aspect of Cloudflare's systems: enable more reliable network connectivity for Cloudflare’s... .... You will work closely with various Engineering teams to translate their requirements...Full timeLocal area
- The Software Reliability Engineer (SRE) will play a critical role in ensuring that our Warehouse Management Software (WMS) runs seamlessly across both automated and manual facilities. This role focuses on investigating, diagnosing, and resolving operational software issues...Full timeLocal areaRemote workRotating shift
$175k - $250k
...customized and developed by our expert team of lawyers, engineers and research scientists. We’ve found product market... ...Competitive compensation. Role Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability...Full timeRelocation package$140k - $230k
...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development... ...across engineering: You will partner closely with software engineering teams to elevate our system architecture, streamline...Full time$152k - $241.5k
NVIDIA's GPU Architecture Group is looking for a software engineer to further modernize and scale GPU development. As GPU designs become more complex, our hardware models, testbenches, build scripts, and code generation flows need to keep adapting to this complexity. The...Full timeRemote work- ...cloud infrastructure (AWS/GCP/Azure) for performance, cost, and reliability. Improve observability across the platform through... ...and integrations across systems and tools. Collaborate with engineering, operations, and data teams to understand and support their infrastructure...Full time
$152k - $241.5k
...Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern... ...artificial intelligence.We are looking for highly motivated Senior Software Engineers to work on our GPU Fabric Networking team. Our team develops...Full timeRemote work$80k - $107k
...GPU Software Engineer (CUDA) – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...superintelligence. One person, one GPU.If you'd like to build the... ...day is currently Tuesday.Engineering at Lambda is responsible for... ...plane services and dataplane software running on SmartNICsDevelop tooling... ...teams to improve service reliability and deployment workflowsDeploy...Work at officeLocal areaWork from homeFlexible hours
$139k - $257.55k
...is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine... ...energy with the resources of a large software company.What you'll doThis is a role... ...inference infrastructure — model serving, GPU workloads, language model gateway and...Full timeTemporary workLocal areaRemote workWorldwide- ...Solution IT Inc. is looking for Sr. GPU AI Solution Cloud Architect for one of its... ...Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field... ...engineering, solutions architecture, site reliability engineering, HPC, or a similar technical...Work experience placementImmediate startRemote work
$267k - $356k
...superintelligence. One person, one GPU.If you'd like to build the... ...Tuesday.Lambda's Storage Engineering team is the backbone behind our... ...in the industry, which means reliability and performance aren't just... ...operating behind Lambda's own software-defined data plane.Build and...Work experience placementWork at officeLocal areaWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!
- agile software developer Remote
- software developer internship no experience Remote
- intermediate software engineer Remote
- software engineer staff Remote
- experienced software developer Remote
- work from home software developer Remote
- software developer no experience Remote
- software developer fintech Remote
- software data engineer Remote
- financial software developer Remote




