Software Engineer - GPU reliability
$200k - $300kHudson River Trading
Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale.
Responsibilities
This role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:
- Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configuration
- Troubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issues
- Working across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently
- Analyzing GPU job statistics to identify trends and areas for improvement
Qualifications
Required:
- BS and/or MS in computer science or a related field
- 2+ years of relevant experience, including programming in Python and managing GPUs
- Experience using automation to solve problems and improve process efficiency
- Experience working with, troubleshooting, tuning, and deploying various types of GPU hardware
- Strong grasp of computer science fundamentals and software design patterns
- Solid understanding of Linux/UNIX operating systems
- Familiarity with open-source software
- Ability to debug and analyze problems quickly
- Skilled at balancing multiple tasks while maintaining meticulous attention to detail
- Ability to operate effectively as a team player and also work independently
- Ability to learn at a fast pace and apply new skills effectively
Preferred:
- Understanding of Debian operating system
- Familiarity with systems configuration management and monitoring technologies
- Familiarity with continuous integration and continuous deployment tools and processes
- Understanding of networking protocols
The estimated base salary range for this position is 200,000 to 300,000 USD per year (or local equivalent). The base pay offered may vary depending on multiple individualized factors, including location, job-related knowledge, skills, and experience.
This role will also be eligible for discretionary performance-based bonuses and a competitive benefits package which includes medical, dental, vision, basic life insurance, and enrollment in our company’s retirement savings plans. Employees will receive sick and parental leave, as well as other paid time off (including 20 vacation days and 10 paid holidays in the US). Please note that benefits and time off policies will vary across non-US locations.
Culture
Hudson River Trading (HRT) brings a scientific approach to trading financial products. We have built one of the world's most sophisticated computing environments for research and development. Our researchers are at the forefront of innovation in the world of algorithmic trading.
At HRT we welcome a variety of expertise: mathematics and computer science, physics and engineering, media and tech. We’re a community of self-starters who are motivated by the excitement of being at the cutting edge of automation in every part of our organization—from trading, to business operations, to recruiting and beyond. We value openness and transparency, and celebrate great ideas from HRT veterans and new hires alike. At HRT we’re friends and colleagues – whether we are sharing a meal, playing the latest board game, or writing elegant code. We embrace a culture of togetherness that extends far beyond the walls of our office. Feel like you belong at HRT? Our goal is to find the best people and bring them together to do great work in a place where everyone is valued. HRT is proud of our diverse staff; we have offices all over the globe and benefit from our varied and unique perspectives. HRT is an equal opportunity employer; so whoever you are we’d love to get to know you.Please be advised: Use of AI tools during interviews or assessments is strictly prohibited, unless otherwise instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that AI assistance was used during any stage of the hiring process, we reserve the right to immediately disqualify your candidacy or rescind any job offers extended.
- ...Join the engineering teams that bring OpenAI’s ideas safely to the world... ...that they are performant and reliable. You will work in a deeply... ...-functional teams, including software engineers, product managers,... ...the platform for CPU/storage, GPU, and network lifecycle management...SuggestedFull timeWork experience placementRelocation package
$175k - $300k
...frontier forward. The Production Engineering Team Key exciting problems... ...10 GW fleet: at our scale, a GPU failure isn't a ticket. It's... ...tooling depend on. Own end-to-end reliability, scalability, and operation... ...silicon level, not just the software stack above it. You move...SuggestedLocal area$109k - $160k
...Learn more at . About the role A Software Engineer contributes to the design,... ...role focuses on improving the efficiency, reliability, and scalability of systems that power... ...product, and hardware teams to evolve our GPU performance testing platform to ensure...SuggestedPermanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$250k
...AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference... ...States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments...SuggestedPermanent employmentRemote work$204k - $259k
...responsible for running the fully autonomous vehicle’s software stack. To achieve our mission, we architect and... ...hybrid role, you will report to a Senior Software Engineer. You will: Develop high-performance GPU primitives and abstractions to enable Waymo to scale...SuggestedFull timeRemote work$170k - $216k
...U.S. states. The Planner/Perception Reliability team builds out architectures, tools, and... ...reliability and is accountable for onboard software health while ensuring high development... ...you will report to a Staff Software Engineer / Tech Lead Manager. You will: Architect...Full timeImmediate startRemote work$125k - $145k
...developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER (FLIGHT RELIAIBLITY) The Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate...Permanent employmentFull timeTemporary workRemote workWorldwideWeekend work$150k - $176k
...Checkr is recognized on Forbes Cloud 100 2025 List and is a Y Combinator 2024 Breakthrough Company . As a Software Engineer II on the Site Reliability Engineering team within the Platform Engineering group at Checkr, you will identify reliability challenges impacting...Full timeWork at officeLocal areaRemote workRelocationFlexible hours3 days per week- ...Conviction. Join us and help build the platform engineers turn to to ship AI products. At... ...for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class... ...network configuration to architect the software fabric that unifies thousands of GPUs...Full timeFlexible hours
- ...online without adding hardware, installing software, or changing a line of code. Internet... ...aspect of Cloudflare's systems: enable more reliable network connectivity for Cloudflare’s... .... You will work closely with various Engineering teams to translate their requirements...Full timeLocal area
$175k - $250k
...customized and developed by our expert team of lawyers, engineers and research scientists. We’ve found product market... ...Competitive compensation. Role Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability...Full timeRelocation package$140k - $230k
...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development... ...across engineering: You will partner closely with software engineering teams to elevate our system architecture, streamline...Full time$102.8k - $190.2k
...Online Products Job Title: Senior Site Reliability Engineer, Data & Analytics Requisition ID: R... ...operating critical systems, writing software, and building tools that improve... ...pipelines and inference services, including GPU workloads ~ Help define how...Full timeTemporary workPart timeLocal areaRemote workRelocation package- ...Lead Site Reliability Engineer Bridge Defense is redefining how modern defense technology is delivered... ..., cleared talent, and mission-ready software to meet evolving defense challenges.... ...physical systems, including high-density GPU servers, networking gear, and storage...Contract workRemote workRelocation
$95k - $171k
...minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$130k - $165k
...Job Title: Senior Software Engineer Company: Snapsheet Job Location: USA, Remote Job Type: Full-time, direct hire Job Department: Technology Team: Site Reliability Engineering About Snapsheet: Snapsheet exists to simplify claims. We leverage our expertise...Full timeTemporary workLocal areaRemote workVisa sponsorshipWork visaFlexible hours$165k - $225k
...Sr. Site Reliability Engineer (SRE) Chicago, IL or Remote Moonlite delivers high-performance... ...solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation... ...engineers, network engineers, and software developers. Preferred Qualifications...Remote workFlexible hours- ...Striveworks Site Reliability Engineer Opportunity Striveworks is a leader in Machine Learning Operations... ...to cloud-based, on-prem, and hybrid software environments, and to resolve any... ...Pipelines Administration/development of GPU-enabled servers Interview Process...Remote work
$170k - $290k
...intelligence. This requires a massive, reliable, and performant GPU infrastructure that pushes the... ...looking for a hands-on, first-principles engineer who is fluent in Linux, comfortable operating... ...toil. Debug Complex Hardware/Software Failures: Serve as the final...Remote work- ...GPU Server Test Software Engineer This role focuses on GPU server testing diagnostic integration, test automation development, and execute comprehensive test plan and test cases to ensure product quality. Key Responsibilities 1. Test Planning & Execution:...Full timeCasual workWork at officeRemote workFlexible hours
$150k - $200k
...Site Reliability Engineer at Triomics (W21) $150K - $200K AI Agents for Oncology EHRs Triomics is building the agentic AI layer for oncology electronic... ...in customer as well as Triomics cloud environments, with GPU infrastructure serving AI extraction models. We need someone...Full timeWork at officeRemote workDay shift$184k - $356.5k
...Developer Tools team and empower engineers throughout the world... ...of the team that brings new GPU technologies to market with sophisticated... ...is looking for a senior software engineer to join our efforts... ...fast, effective, maintainable, reliable and well-documented code....Full time$170k - $216k
...developer productivity of onboard engineers to enable the entire... ...role you will report to a Staff Software Engineer / Tech Lead Manager. You will: Develop reliable, scalable, and maintainable systems... ...should happen on CPU, GPU, and TPU. Collaborate to resolve...Full timeRemote work- ...Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco • Full-Time About Andromeda Andromeda Cluster... ...response. Nice to Have Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.). Familiarity...Full timeRemote work
- ...TENEX Staff Site Reliability Engineer TENEX is an AI-native, automation-first, built-for-scale... ...years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing... ...for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization)....Work from home
- ...and implement Infrastructure Services for a GPU-as-a-Service platform, the full-time remote Full-Stack Software Engineer will build control planes that translate high... ...metal server provisioning, and ensure system reliability. Key Responsibilities Design, build, and maintain...Full timeRemote work
$90k - $190k
...Tutor Intelligence. As an AI software company that deploys its inventions... .... Full Stack Software Engineer (XR Applications) As a... ..., you will write performant, GPU-enabled frontend interfaces for... ...solutions Improve system reliability, scalability, and...Full timeWork at office$175k - $250k
Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality... ...ensuring scalability, performance, and reliability across environments. What You’ll Do... ...workloads at scale Manage and automate GPU compute clusters using tools such as Python...Full timeRemote workRelocationRelocation package$166k - $244k
# Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them...Temporary work$149.4k - $202k
Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!
- software engineer full time Remote
- software system engineer Remote
- consulting software engineer Remote
- software engineer travel Remote
- software engineer mainframe Remote
- software engineer unity Remote
- real time software engineer Remote
- network software engineer Remote
- senior software engineer remote Remote
- entry level software engineer remote Remote



