Software Engineer - GPU reliability
$200k - $300kHudson River Trading
Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale.
Responsibilities
This role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:
- Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configuration
- Troubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issues
- Working across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently
- Analyzing GPU job statistics to identify trends and areas for improvement
Qualifications
Required:
- BS and/or MS in computer science or a related field
- 2+ years of relevant experience, including programming in Python and managing GPUs
- Experience using automation to solve problems and improve process efficiency
- Experience working with, troubleshooting, tuning, and deploying various types of GPU hardware
- Strong grasp of computer science fundamentals and software design patterns
- Solid understanding of Linux/UNIX operating systems
- Familiarity with open-source software
- Ability to debug and analyze problems quickly
- Skilled at balancing multiple tasks while maintaining meticulous attention to detail
- Ability to operate effectively as a team player and also work independently
- Ability to learn at a fast pace and apply new skills effectively
Preferred:
- Understanding of Debian operating system
- Familiarity with systems configuration management and monitoring technologies
- Familiarity with continuous integration and continuous deployment tools and processes
- Understanding of networking protocols
The estimated base salary range for this position is 200,000 to 300,000 USD per year (or local equivalent). The base pay offered may vary depending on multiple individualized factors, including location, job-related knowledge, skills, and experience.
This role will also be eligible for discretionary performance-based bonuses and a competitive benefits package which includes medical, dental, vision, basic life insurance, and enrollment in our company’s retirement savings plans. Employees will receive sick and parental leave, as well as other paid time off (including 20 vacation days and 10 paid holidays in the US). Please note that benefits and time off policies will vary across non-US locations.
Culture
Hudson River Trading (HRT) brings a scientific approach to trading financial products. We have built one of the world's most sophisticated computing environments for research and development. Our researchers are at the forefront of innovation in the world of algorithmic trading.
At HRT we welcome a variety of expertise: mathematics and computer science, physics and engineering, media and tech. We’re a community of self-starters who are motivated by the excitement of being at the cutting edge of automation in every part of our organization—from trading, to business operations, to recruiting and beyond. We value openness and transparency, and celebrate great ideas from HRT veterans and new hires alike. At HRT we’re friends and colleagues – whether we are sharing a meal, playing the latest board game, or writing elegant code. We embrace a culture of togetherness that extends far beyond the walls of our office. Feel like you belong at HRT? Our goal is to find the best people and bring them together to do great work in a place where everyone is valued. HRT is proud of our diverse staff; we have offices all over the globe and benefit from our varied and unique perspectives. HRT is an equal opportunity employer; so whoever you are we’d love to get to know you.Please be advised: Use of AI tools during interviews or assessments is strictly prohibited, unless otherwise instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that AI assistance was used during any stage of the hiring process, we reserve the right to immediately disqualify your candidacy or rescind any job offers extended.
$200k - $300k
...Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server...SuggestedFull timeWork at officeLocal areaImmediate start$250k
...AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference... ...States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments...SuggestedFull timeRemote work- ...teams spanning hardware and software. Speed and scale are our key... ...forward. The Production Engineering Team Examples of key exciting... ...100s of GWs: at our scale, a GPU failure isn't a ticket. It's... ...depend on. Own end-to-end reliability, scalability, and operation of...SuggestedLocal area
- ...Join the engineering teams that bring OpenAI’s ideas safely to the world... ...that they are performant and reliable. You will work in a deeply... ...-functional teams, including software engineers, product managers,... ...the platform for CPU/storage, GPU, and network lifecycle management...SuggestedFull timeWork experience placementRelocation package
$150k - $200k
...Software Engineer, Reliability (Avionics / Compute Systems) Cowboy Space Corp. is building the infrastructure to power and connect the orbital economy... ..., or embedded compute platforms Familiarity with GPU/accelerator-based systems and performance validation (e.g....SuggestedPermanent employmentFull timeWork at officeLocal areaRelocation package$182k - $242k
...AI-cloud for high-performance GPU infrastructure across AI/ML,... ...time inference. Our stack is engineered for speed, scale, and cost-efficiency... ...to latency, throughput, and reliability across our inference stack.... ...computing, GPU/accelerator software, or performance-critical...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$215k - $285k
...success for both clients and candidates. Senior Software Engineer, GPU Sandboxes Location: San Francisco, CA Company Stage... ...through deployment, observability, on-call, and ongoing reliability. Work closely with the GPU Sandboxes Engineering Lead and...Full timeWork at office- Lambda is building a GPU-accelerated Kubernetes-based AI cloud orchestration platform. The Senior Software Engineer will shape architecture, reliability, and automation for Kubernetes-based infrastructure powering AI workloads at scale. This role requires four days in the...Work at officeWork from home
- ...role carries a dual mandate: ensuring the reliability, security, and continuity of production... ...of self-hosted LLM deployment, GPU compute requirements, model serving frameworks... ...enterprise vendor management including software license tracking, renewal calendars, and...For contractorsWork at office
- ...need you, a highly experienced Senior Software Engineer, to join the technical team located in... ...codebase for enhance availability, increased reliability, and an improved user experience-... ...Extensive hands-on experience with DIO/DAQ and GPU/GPGPU-Demonstrable experience in all...Visa sponsorship
$165.8k - $308k
...but we are changing lives. Our software teams are laying the... ...team builds highly scalable, reliable software and secure systems for... ...OpportunityAs a Bioinformatics Software Engineer, you will design and develop... ...techniques using GPU hardware. Support software development...Full timeLocal areaRelocation package$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with... .../ hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware...Full timeRemote work$125k - $145k
...developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SOFTWARE ENGINEER (FLIGHT RELIAIBLITY)The Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate launch...Permanent employmentTemporary workRemote workWorldwideWeekend work- The roleWe’re looking to hire our first data software engineer at Gridmatic! Looking for a startup-minded eng who works closely with our ML... ...area to both make sure the data is ingested and transformed reliably, and also be able to build the tooling/abstractions to make...Work at officeHome officeFlexible hours3 days per week
$109k - $160k
...Learn more at . About the role A Software Engineer contributes to the design,... ...role focuses on improving the efficiency, reliability, and scalability of systems that power... ...product, and hardware teams to evolve our GPU performance testing platform to ensure...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours- ...of superintelligence. One person, one GPU.If you'd like to build the world's best... ...centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability... ..., distributed systems, or production software engineering.Have deep experience operating...Work at officeLocal areaWork from homeFlexible hours
- ...Ultimate Staffing is seeking a GPU Server Test Software Engineer to join a client in Fort Worth. This is a full-time, direct hire position. The... ...executing thorough test plans to ensure product quality and reliability. The ideal candidate will have hands-on experience with...Full timeLocal areaRemote work
- Role Description We are looking for a Site Reliability Engineer with a strong software development background to join our Scrum team maintaining and improving an Enterprise Generative AI platform. This role focuses on platform stability, performance optimization, uptime...Full timeRemote workFlexible hours
- The Software Reliability Engineer (SRE) will play a critical role in ensuring that our Warehouse Management Software (WMS) runs seamlessly across both automated and manual facilities. This role focuses on investigating, diagnosing, and resolving operational software issues...Full timeLocal areaRemote workRotating shift
$100k - $175k
...GPU Software Engineer (CUDA)- Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$204k - $259k
...Software Engineer, GPU Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced...Full timeRemote work$98.9k - $148.3k
...Qualcomm Technologies, Inc. Job Area: Engineering Group, Engineering Group Graphics Software Engineering General Summary: As a leading technology... ..., and optimize the structure and performance of GPU hardware, drivers, features, applications, and tools...Work from home- ...power of superintelligence. One person, one GPU.If you'd like to build the world's best... ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building... ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,...Work at officeLocal areaWork from homeFlexible hours
- ...superintelligence. One person, one GPU.If you'd like to build the... ...day is currently Tuesday.Engineering at Lambda is responsible for... ...plane services and dataplane software running on SmartNICsDevelop tooling... ...teams to improve service reliability and deployment workflowsDeploy...Work at officeLocal areaWork from homeFlexible hours
$139k - $257.55k
...is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine... ...energy with the resources of a large software company.What you'll doThis is a role... ...inference infrastructure — model serving, GPU workloads, language model gateway and...Full timeTemporary workLocal areaRemote workWorldwide$267k - $356k
...superintelligence. One person, one GPU.If you'd like to build the... ...Tuesday.Lambda's Storage Engineering team is the backbone behind our... ...in the industry, which means reliability and performance aren't just... ...operating behind Lambda's own software-defined data plane.Build and...Work experience placementWork at officeLocal areaWork from homeFlexible hours$174k - $252k
...as system design consulting, developing software platforms and frameworks, capacity... ...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...- ...cloud infrastructure (AWS/GCP/Azure) for performance, cost, and reliability. Improve observability across the platform through... ...and integrations across systems and tools. Collaborate with engineering, operations, and data teams to understand and support their infrastructure...Full time
$92.3k - $153.9k
AVEVA is creating software trusted by over 90% of leading industrial companies.Salary Range... ...and/or training. Software Development Engineer - Performance & ReliabilityLocation: Lake... ..., ensuring the platform performs reliably at scale across distributed microservices...Full timeWork at officeLocal areaRemote workFlexible hours- ...to creating category-leading enterprise software that unleashes that power.To make that... ...part of this journey?At UiPath's Site Reliability team, we build the platforms and systems... ...each, powered by AI.This is a software engineering role. You will not be the person who identifies...Work at officeImmediate startRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!
- software engineer internship Remote
- software development engineer aws Remote
- software developer internship no experience Remote
- real time software engineer Remote
- financial software developer Remote
- oracle software engineer Remote
- part time software developer Remote
- graduate software developer Remote
- software engineer travel Remote
- experienced software developer Remote


