Software Engineer - GPU reliability
$200k - $300kHudson River Trading
Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale.ResponsibilitiesThis role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configurationTroubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issuesWorking across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently Analyzing GPU job statistics to identify trends and areas for improvementQualificationsRequired:BS and/or MS in computer science or a related field2+ years of relevant experience, including programming in Python and managing GPUsExperience using automation to solve problems and improve process efficiencyExperience working with, troubleshooting, tuning, and deploying various types of GPU hardwareStrong grasp of computer science fundamentals and software design patternsSolid understanding of Linux/UNIX operating systems Familiarity with open-source softwareAbility to debug and analyze problems quicklySkilled at balancing multiple tasks while maintaining meticulous attention to detailAbility to operate effectively as a team player and also work independently Ability to learn at a fast pace and apply new skills effectivelyPreferred:Understanding of Debian operating systemFamiliarity with systems configuration management and monitoring technologies Familiarity with continuous integration and continuous deployment tools and processesUnderstanding of networking protocolsThe estimated base salary range for this position is 200,000 to 300,000 USD per year (or local equivalent). The base pay offered may vary depending on multiple individualized factors, including location, job-related knowledge, skills, and experience. This role will also be eligible for discretionary performance-based bonuses and a competitive benefits package which includes medical, dental, vision, basic life insurance, and enrollment in our company’s retirement savings plans. Employees will receive sick and parental leave, as well as other paid time off (including 20 vacation days and 10 paid holidays in the US). Please note that benefits and time off policies will vary across non-US locations.CultureHudson River Trading (HRT) brings a scientific approach to trading financial products. We have built one of the world's most sophisticated computing environments for research and development. Our researchers are at the forefront of innovation in the world of algorithmic trading. At HRT we welcome a variety of expertise: mathematics and computer science, physics and engineering, media and tech. We’re a community of self-starters who are motivated by the excitement of being at the cutting edge of automation in every part of our organization—from trading, to business operations, to recruiting and beyond. We value openness and transparency, and celebrate great ideas from HRT veterans and new hires alike. At HRT we’re friends and colleagues – whether we are sharing a meal, playing the latest board game, or writing elegant code. We embrace a culture of togetherness that extends far beyond the walls of our office.Feel like you belong at HRT? Our goal is to find the best people and bring them together to do great work in a place where everyone is valued. HRT is proud of our diverse staff; we have offices all over the globe and benefit from our varied and unique perspectives. HRT is an equal opportunity employer; so whoever you are we’d love to get to know you.Please be advised: Use of AI tools during interviews or assessments is strictly prohibited, unless otherwise instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that AI assistance was used during any stage of the hiring process, we reserve the right to immediately disqualify your candidacy or rescind any job offers extended.
$160k - $240k
Senior Software Engineer - BQL Reliability Engineering Location New York Business Area Engineering and CTO Ref # 10053945 Description & Requirements What You’ll Do: As part of the BQL (Bloomberg Query Language) Reliability Engineering team, you will...SuggestedTemporary workFor contractorsWork experience placement$207k - $300k
Design, develop, test, and deploy scalable software solutions that maintain and enhance... ...providing feedback to ensure best practices in reliability, security, and efficiency.Triage and... ...development initiatives. Mentor other engineers and contribute to the engineering...SuggestedFull timeWork at office$139k - $257.55k
...is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine... ...energy with the resources of a large software company.What you'll doThis is a role... ...inference infrastructure — model serving, GPU workloads, language model gateway and...SuggestedFull timeTemporary workLocal areaRemote workWorldwide$190k - $260k
...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...distributed systems with Kubernetes, and GPU workloads on those clustersExperience with...SuggestedFull timeWork experience placementWork at officeLocal areaRemote workHome office$139k - $257.55k
...experience and stronger capabilitiesBuild out next-generation GPU-accelerated features and modernize existing featuresCollaborate... ...usersCollaborate with technology and platform teamsWrite and review engineering documents and design specsEnsure quality in all phases of...SuggestedFull timeTemporary workLocal areaWorldwideShift work$117.5k - $157.5k
Job Posting Title:Software Engineer IIReq ID:10148551Job Description:Technology is at the heart... ...Disney Streaming’s distributed systems are reliable, performant, and transparent. We build... ...on time-series data, model training on GPU clusters, real-time inference pipelines...Full time$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ...while embedding directly with product engineering teams and their leadership to drive... ...the open source community; evangelize software engineering best practices, especially...Full timeWork experience placementWork at officeLocal area2 days per week3 days per week- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership...
$139k - $242k
...Learn more at .What You’ll Do:The Runtime & GPU Systems team builds and operates secure,... ..., GPU infrastructure, and Linux systems engineering. We partner closely with security,... ...diagnosing and resolving complex performance, reliability, or isolation issues across containers,...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$153k - $204k
...managed storage products. We build reliable, scalable storage solutions... .... Object Storage works with engineering teams across infrastructure,... ...technologies such as RDMA, GPU Direct Storage, and distributed... ...AI tools to augment software development.Familiarity with...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$165.3k - $219.68k
...insights to improve their business. Founded by engineers — and customer obsessed — we leap at... ...Model APIs, to state a few.Improve reliability, latency, and efficiency of distributed... ...real-time serving, ML infrastructure, or GPU orchestrationExposure to platforms like...Local areaWorldwide$150k - $160k
Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs...Work at officeLocal area$182k - $250.8k
...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great... ...standards that embed observability, resilience, and software engineering rigor into all engineering...Permanent employmentLocal areaRemote workWorldwideFlexible hoursWeekend workWeekday work$207k - $300k
...deployment, operation, and refinement.Manage availability, latency, scalability, and efficiency of Google services by engineering reliability into software and systems.Respond to and resolve emergent service problems; write software and build automation to prevent problem...$152k - $241.5k
...world working for us. If you're passionate about building reliable systems software for cloud-scale GPU infrastructure, we encourage you to contribute to our team. We are looking for a Senior Software Engineer to join our DGX Cloud / Fleet Intelligence team and build agent...Full timeLocal areaRemote work$182k - $242k
...seeking a passionate and innovative Senior Software Engineer of Network Services to lead the... ...roadmap, drive innovation, and ensure the reliability, security, and scalability of the CoreWeave... ...services infrastructure for our GPU cloud services, including networking cloud...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...Saragossa seeks a senior infrastructure engineer to own and evolve the compute platform for... ...environments, high-performance storage, GPU workloads, cloud infrastructure, and... ...role emphasizes performance, automation, reliability, and scale, with significant ownership over...
$195k - $275k
...Management, Capital Markets Application & Data Services, Deployment Planning & Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the...Temporary workWork at officeWorldwideNight shift$158.5k - $172k
...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will... ...high-impact position driving continuous reliability, deep system optimization, and... ...ensuring fast, secure, and friction-free software delivery workflows.Secure and Standardize...Full timeTemporary workWork at officeFlexible hours3 days per week$141k - $216.6k
...and justice issues with our ecosystem of devices and cloud software. Like our products, we work better together. We connect with... ...building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational...Work experience placementWork at office$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team... ...this stack, and are the principal drivers of growth for software and infrastructure practice within our larger Enterprise Technology...Work at officeLocal areaImmediate start- ...world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank... ...best practices within your teamCollaborates with other software engineers and teams to design, develop, test, and implement...Shift work
$120k - $150k
...allows each person to achieve personal success and add value to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About the RoleJoin our Platform Engineering team as a Site Reliability...Full time$182.8k - $247.3k
...learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering... ...in production)Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews and root...Work experience placement- The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also...Full timeWork experience placementRemote work
$160k - $240k
Senior Software Engineer - AI Inference Location New York Business Area Engineering... ...platform, balancing latency, throughput, reliability, and cost. Partner across engineering... ...serving. Familiarity with PyTorch and GPU software stacks such as CUDA and NCCL....Temporary workFor contractorsWork experience placement$184.9k - $250.2k
We are seeking a Senior Software Development Engineer to build and scale the software systems that power... ...ensuring that perception systems are reliable, performant, and maintainable at scale... ...for embedded hardware (e.g., ARM, GPU-accelerated edge devices), including...InternshipFlexible hours$145k - $182k
...re looking for a Senior AI/ML Engineer to design, build, and... ...connector ecosystem - Ensure reliability: Monitor pipeline performance... ...efficiency.Collaborate with software engineers to integrate machine... ...distributed training orchestration, GPU/TPU resource allocation, and...Full timeTemporary workWork at officeShift work3 days per week$153k - $242k
...Learn more at .About the RoleAs a Senior Software Engineer within our Compute Architecture... ...lifecycle management across large-scale GPU data centers. The METALDEV team builds... ...GPU servers and rack-scale systems with reliability and confidence. This is a software-first...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$184k - $287.5k
...computing. An era in which our GPU acts as the brains of... ...you excited about open-source software, developer tools, and the future... ...looking for a Senior Software Engineer to help build the core libraries... ..., and proving quality, reliability, cost, and latency gains across...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!
- summer intern software engineer New York, NY
- software developer fintech New York, NY
- startup software engineer New York, NY
- financial software developer New York, NY
- junior software developer remote New York, NY
- software engineer New York, NY
- software data engineer New York, NY
- freelance software developer New York, NY
- software developer internship no experience New York, NY
- part time software developer New York, NY

