Software Engineer - GPU reliability
$200k - $300kHudson River Trading
Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale. Responsibilities This role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:
- Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configuration
- Troubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issues
- Working across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently
- Analyzing GPU job statistics to identify trends and areas for improvement
- BS and/or MS in computer science or a related field
- 2+ years of relevant experience, including programming in Python and managing GPUs
- Experience using automation to solve problems and improve process efficiency
- Experience working with, troubleshooting, tuning, and deploying various types of GPU hardware
- Strong grasp of computer science fundamentals and software design patterns
- Solid understanding of Linux/UNIX operating systems
- Familiarity with open-source software
- Ability to debug and analyze problems quickly
- Skilled at balancing multiple tasks while maintaining meticulous attention to detail
- Ability to operate effectively as a team player and also work independently
- Ability to learn at a fast pace and apply new skills effectively
- Understanding of Debian operating system
- Familiarity with systems configuration management and monitoring technologies
- Familiarity with continuous integration and continuous deployment tools and processes
- Understanding of networking protocols
$159.8k - $235k
About the TeamThe Reliability Platform role is a key pillar of DoorDash... ...and repetitive tasks. We use software and agents to “keep the... ...!About the RoleAs a Software Engineer on the Reliability Platform team... ...Kafka topics, Databases, CPU/GPU Pools, Service Scaffolding, etc...SuggestedHourly payWork at officeLocal areaRemote workFlexible hours$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI... ...latency and output variance, and token/GPU cost anomalies. Translate reliability... ...Product Engineering has modernized software and product delivery, creating a scalable...SuggestedWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week$139k - $257.55k
...is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine... ...energy with the resources of a large software company.What you'll doThis is a role... ...inference infrastructure — model serving, GPU workloads, language model gateway and...SuggestedFull timeTemporary workLocal areaRemote workWorldwide- ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...distributed systems with Kubernetes, and GPU workloads on those clustersExperience with...SuggestedFull timeWork experience placementWork at officeLocal areaRemote workHome office
- ...infrastructure foundation for AI teams. With instant GPU access, sub-second container startups,... ...olympiad medalists, and experienced engineering and product leaders with decades of... ...company, we seek to improve our reliability dramatically while scaling the size of our...Suggested
$325k
...Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems... ...growing group of committed researchers, engineers, policy experts, and business leaders... ...-- we're looking for reliability-minded software engineers and SREs. Are curious and...Full timeWork at officeVisa sponsorshipFlexible hours$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating complex operational challenges, improving observability...- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management....Flexible hours
- ...innovators in this way. The RoleThe Director of Platform & Reliability Engineering will lead a critical engineering organization responsible for... ..., and post-incident learning.Partner with product and software engineering leaders to create paved-road solutions that improve...Work at officeLocal area2 days per week3 days per week
$140k - $225k
...blackstone on LinkedIn, X, and Instagram.Role: Blackstone's Site Reliability Engineering team is responsible for improving the reliability of... ...professional experience with either, Infrastructure Engineering, Software Engineering, DevOps Engineering or Platform Engineering....Full timeLocal areaFlexible hours$109k - $145k
...This team enables both internal engineers and customers to monitor,... ...optimize AI workloads running on GPU-dense infrastructure at... ...About the role: As a Software Engineer on the Observability... ...Python, while improving system reliability through enhanced monitoring,...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$139k - $242k
...Learn more at .What You’ll Do:The Runtime & GPU Systems team builds and operates secure,... ..., GPU infrastructure, and Linux systems engineering. We partner closely with security,... ...diagnosing and resolving complex performance, reliability, or isolation issues across containers,...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...development out of the field and into software, Antioch drastically reduces... ...for? We're hiring Platform Engineers to build the cloud... ...: orchestration, scheduling, GPU pools, storage, and observability... ...loops agents need to iterate reliably. Own performance,...Full timeRelocation
- As a Site Reliability Engineering at JPMorgan Chase within the Enterprise technology, liquidity risk team, you are the non-functional requirement... ..., monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and...
$165.3k - $219.68k
...insights to improve their business. Founded by engineers — and customer obsessed — we leap at... ...Model APIs, to state a few.Improve reliability, latency, and efficiency of distributed... ...real-time serving, ML infrastructure, or GPU orchestrationExposure to platforms like...Local areaWorldwide$150k - $160k
Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs...Work at officeLocal area$197.3k - $313.7k
...of Salesforce.Job Title: Director, Site Reliability EngineeringLocation: New York, NY; San... ...looking for a Director of Site Reliability Engineering to spearhead the evolution of our... ...requirements are incorporated throughout the software development lifecycle rather than...Full timeImmediate start- ..., driven by pride in ownership.As a Senior Manager of Site Reliability Engineering at JPMorgan Chase within the Corporate Investment Bank, Markets... ..., monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and...Bank staffShift work
$150k - $190k
Senior Site Reliability Engineer, VPAt Morgan Stanley, we advise, originate, trade, manage and distribute capital for governments, institutions... ...team or independently.Test and tune network, hardware, and software configurations to maximize performance needs.Troubleshoot...Temporary workWorldwideFlexible hoursWeekend work$182k - $242k
...seeking a passionate and innovative Senior Software Engineer of Network Services to lead the... ...roadmap, drive innovation, and ensure the reliability, security, and scalability of the CoreWeave... ...services infrastructure for our GPU cloud services, including networking cloud...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$195k - $275k
...Management, Capital Markets Application & Data Services, Deployment Planning & Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the...Temporary workWork at officeWorldwideNight shift$123k - $165k
Job Summary:Department/Group OverviewOur engineering fleet is a horizontal set of teams... ...organization. Our specific team provides reliability engineering and operational support to... ...operational workflows.Collaborate with software engineering teams to implement SRE best...$165k - $241.4k
...functional and very effective.We’re looking for talented engineers with a software or operations background, experienced in designing and operating... ...with our application development teams to ensure the reliability, performance and security of our infrastructure.ResponsibilitiesJoin...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$138.1k - $198.2k
...technology that simply works. The SRE Engineering Enablement Team supports our CI Platforms... ...day-to-day work. We support the entire software development lifecycle (SDLC), including... ...at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter...Permanent employmentFull timeTemporary workWork experience placementLocal areaRemote workFlexible hours$141k - $216.6k
...and justice issues with our ecosystem of devices and cloud software. Like our products, we work better together. We connect with... ...building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational...Work experience placementWork at office$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team... ...this stack, and are the principal drivers of growth for software and infrastructure practice within our larger Enterprise Technology...Work at officeLocal areaImmediate start$120k - $200k
...NY, United StatesSalary: USD 120000 - 200000 (Yearly)Job Type: PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery Participate in the design and implementation of the company’s...Overseas$120k - $142k
...making journalism so good that it’s worth paying for. Mission Overview & Responsibilities: At The New York Times, our Site Reliability Engineering (SRE) team is central to how we design, test, and operate the systems that support our most critical customer experiences....Local areaFlexible hours- ...world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank... ...best practices within your teamCollaborates with other software engineers and teams to design, develop, test, and implement...Shift work
$110k - $120k
...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading... ...on SS&C for expertise, scale, and technology.Kick off your software engineering career on our Quality & Automation team. You...Ongoing contractFull timeCasual workRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!
- ngo software engineer New York, NY
- software data engineer New York, NY
- graduate software engineer New York, NY
- software system engineer New York, NY
- associate software engineer New York, NY
- software developer no experience New York, NY
- graduate software developer New York, NY
- software engineer - early career New York, NY
- entry level software engineer remote New York, NY
- software engineer intern New York, NY


