Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - GPU reliability

$200k - $300k

Hudson River Trading

Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale.ResponsibilitiesThis role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configurationTroubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issuesWorking across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently Analyzing GPU job statistics to identify trends and areas for improvementQualificationsRequired:BS and/or MS in computer science or a related field2+ years of relevant experience, including programming in Python and managing GPUsExperience using automation to solve problems and improve process efficiencyExperience working with, troubleshooting, tuning, and deploying various types of GPU hardwareStrong grasp of computer science fundamentals and software design patternsSolid understanding of Linux/UNIX operating systems Familiarity with open-source softwareAbility to debug and analyze problems quicklySkilled at balancing multiple tasks while maintaining meticulous attention to detailAbility to operate effectively as a team player and also work independently Ability to learn at a fast pace and apply new skills effectivelyPreferred:Understanding of Debian operating systemFamiliarity with systems configuration management and monitoring technologies Familiarity with continuous integration and continuous deployment tools and processesUnderstanding of networking protocolsThe estimated base salary range for this position is 200,000 to 300,000 USD per year (or local equivalent). The base pay offered may vary depending on multiple individualized factors, including location, job-related knowledge, skills, and experience. This role will also be eligible for discretionary performance-based bonuses and a competitive benefits package which includes medical, dental, vision, basic life insurance, and enrollment in our company’s retirement savings plans. Employees will receive sick and parental leave, as well as other paid time off (including 20 vacation days and 10 paid holidays in the US). Please note that benefits and time off policies will vary across non-US locations.CultureHudson River Trading (HRT) brings a scientific approach to trading financial products. We have built one of the world's most sophisticated computing environments for research and development. Our researchers are at the forefront of innovation in the world of algorithmic trading. At HRT we welcome a variety of expertise: mathematics and computer science, physics and engineering, media and tech. We’re a community of self-starters who are motivated by the excitement of being at the cutting edge of automation in every part of our organization—from trading, to business operations, to recruiting and beyond. We value openness and transparency, and celebrate great ideas from HRT veterans and new hires alike. At HRT we’re friends and colleagues – whether we are sharing a meal, playing the latest board game, or writing elegant code. We embrace a culture of togetherness that extends far beyond the walls of our office.Feel like you belong at HRT? Our goal is to find the best people and bring them together to do great work in a place where everyone is valued. HRT is proud of our diverse staff; we have offices all over the globe and benefit from our varied and unique perspectives. HRT is an equal opportunity employer; so whoever you are we’d love to get to know you.Please be advised: Use of AI tools during interviews or assessments is strictly prohibited, unless otherwise instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that AI assistance was used during any stage of the hiring process, we reserve the right to immediately disqualify your candidacy or rescind any job offers extended.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Software Engineer - GPU reliability in New York, NY vacancy
  • $159.8k - $235k

    About the TeamThe Reliability Platform role is a key pillar of DoorDash...  ...and repetitive tasks. We use software and agents to “keep the...  ...!About the RoleAs a Software Engineer on the Reliability Platform team...  ...Kafka topics, Databases, CPU/GPU Pools, Service Scaffolding, etc... 
    Suggested
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    New York, NY
    2 days ago
  •  ...teams spanning hardware and software. Speed and scale are our key...  ...forward. The Production Engineering Team Examples of key exciting...  ...100s of GWs: at our scale, a GPU failure isn't a ticket. It's...  ...depend on. Own end-to-end reliability, scalability, and operation of... 
    Suggested
    Local area

    FluidStack

    New York, NY
    2 days ago
  • $204k - $259k

     ...responsible for running the fully autonomous vehicle’s software stack. To achieve our mission, we architect and...  ...hybrid role, you will report to a Senior Software Engineer. You will: Develop high-performance GPU primitives and abstractions to enable Waymo to scale... 
    Suggested
    Full time
    Remote work

    Waymo

    New York, NY
    17 hours ago
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI...  ...latency and output variance, and token/GPU cost anomalies. Translate reliability...  ...Product Engineering has modernized software and product delivery, creating a scalable... 
    Suggested
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    New York, NY
    3 days ago
  •  ...GPU Kernel Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable... 
    Suggested
    Flexible hours

    Baseten

    New York, NY
    4 days ago
  • $139k - $257.55k

     ...is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine...  ...energy with the resources of a large software company.What you'll doThis is a role...  ...inference infrastructure — model serving, GPU workloads, language model gateway and... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    13 hours ago
  •  ...infrastructure foundation for AI teams. With instant GPU access, sub-second container startups,...  ...olympiad medalists, and experienced engineering and product leaders with decades of...  ...company, we seek to improve our reliability dramatically while scaling the size of our... 

    Modal Labs

    New York, NY
    4 days ago
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...distributed systems with Kubernetes, and GPU workloads on those clustersExperience with... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    4 days ago
  • $325k

     ...Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems...  ...growing group of committed researchers, engineers, policy experts, and business leaders...  ...-- we're looking for reliability-minded software engineers and SREs. Are curious and... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    17 hours ago
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating complex operational... 

    Ridge Line Services

    New York, NY
    3 days ago
  •  ...innovators in this way. The RoleThe Director of Platform & Reliability Engineering will lead a critical engineering organization responsible for...  ..., and post-incident learning.Partner with product and software engineering leaders to create paved-road solutions that improve... 
    Work at office
    Local area
    2 days per week
    3 days per week

    Forge Global

    New York, NY
    13 hours ago
  • $140k - $170k

     ...optimize business interactions. Role Description: As a Site Reliability Engineer, you will work with Agile engineering teams to provide production insight into running and operating software at-scale in a globally distributed and highly available cloud based system... 
    Full time
    Local area

    Symphony Communication Services

    New York, NY
    10 hours ago
  • $140k - $225k

     ...blackstone on LinkedIn, X, and Instagram.Role: Blackstone's Site Reliability Engineering team is responsible for improving the reliability of...  ...professional experience with either, Infrastructure Engineering, Software Engineering, DevOps Engineering or Platform Engineering.... 
    Full time
    Local area
    Flexible hours

    Blackstone Group

    New York, NY
    4 days ago
  • $109k - $145k

     ...This team enables both internal engineers and customers to monitor,...  ...optimize AI workloads running on GPU-dense infrastructure at...  ...About the role: As a Software Engineer on the Observability...  ...Python, while improving system reliability through enhanced monitoring,... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Coreweave

    New York, NY
    17 hours ago
  • $139k - $242k

     ...Learn more at .What You’ll Do:The Runtime & GPU Systems team builds and operates secure,...  ..., GPU infrastructure, and Linux systems engineering. We partner closely with security,...  ...diagnosing and resolving complex performance, reliability, or isolation issues across containers,... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    4 days ago
  • As a Site Reliability Engineering at JPMorgan Chase within the Enterprise technology, liquidity risk team, you are the non-functional requirement...  ..., monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and... 

    JP Morgan Chase

    New York, NY
    3 days ago
  •  ...development out of the field and into software, Antioch drastically reduces...  ...for? We're hiring Platform Engineers to build the cloud...  ...: orchestration, scheduling, GPU pools, storage, and observability...  ...loops agents need to iterate reliably. Own performance,... 
    Full time
    Relocation

    Antioch

    New York, NY
    17 hours ago
  • $165.3k - $219.68k

     ...insights to improve their business. Founded by engineers — and customer obsessed — we leap at...  ...Model APIs, to state a few.Improve reliability, latency, and efficiency of distributed...  ...real-time serving, ML infrastructure, or GPU orchestrationExposure to platforms like... 
    Local area
    Worldwide

    DataBricks

    New York, NY
    13 hours ago
  • $131k - $164k

    Position Overview We are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and...  ...smart, and creative people who not only want to help build the software company of the future, but who want to make the world a more... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Diligent Corporation

    New York, NY
    2 days ago
  • $150k - $160k

    Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs... 
    Work at office
    Local area

    Haymarket Media Group

    New York, NY
    4 days ago
  • $197.3k - $313.7k

     ...of Salesforce.Job Title: Director, Site Reliability EngineeringLocation: New York, NY; San...  ...looking for a Director of Site Reliability Engineering to spearhead the evolution of our...  ...requirements are incorporated throughout the software development lifecycle rather than... 
    Full time
    Immediate start

    Salesforce

    New York, NY
    3 days ago
  •  ..., driven by pride in ownership.As a Senior Manager of Site Reliability Engineering at JPMorgan Chase within the Corporate Investment Bank, Markets...  ..., monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and... 
    Bank staff
    Shift work

    JP Morgan Chase

    New York, NY
    1 day ago
  • $150k - $190k

    Senior Site Reliability Engineer, VPAt Morgan Stanley, we advise, originate, trade, manage and distribute capital for governments, institutions...  ...team or independently.Test and tune network, hardware, and software configurations to maximize performance needs.Troubleshoot... 
    Temporary work
    Worldwide
    Flexible hours
    Weekend work

    Morgan Stanley

    New York, NY
    4 days ago
  • $182k - $242k

     ...seeking a passionate and innovative Senior Software Engineer of Network Services to lead the...  ...roadmap, drive innovation, and ensure the reliability, security, and scalability of the CoreWeave...  ...services infrastructure for our GPU cloud services, including networking cloud... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    4 days ago
  • $165k - $241.4k

     ...functional and very effective.We’re looking for talented engineers with a software or operations background, experienced in designing and operating...  ...with our application development teams to ensure the reliability, performance and security of our infrastructure.ResponsibilitiesJoin... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    3 days ago
  • $138.1k - $198.2k

     ...technology that simply works.  The SRE Engineering Enablement Team supports our CI Platforms...  ...day-to-day work. We support the entire software development lifecycle (SDLC), including...  ...at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    New York, NY
    3 days ago
  • $141k - $216.6k

     ...and justice issues with our ecosystem of devices and cloud software. Like our products, we work better together. We connect with...  ...building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational... 
    Work experience placement
    Work at office

    Axon

    New York, NY
    2 days ago
  • $88k - $158k

     ...Security (AIS) is looking for talented Software Engineers to join our Cyber AI team. AIS’s Cyber...  ...embedded or resource-constrained devices (GPU/CPU/NPU optimization, quantization,...  ...reasoning and multimodal sensing, and ensure reliable operation in real-world environments.... 

    Assured Information Security

    New York, NY
    4 days ago
  • $158.5k - $172k

     ...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will...  ...high-impact position driving continuous reliability, deep system optimization, and...  ...ensuring fast, secure, and friction-free software delivery workflows.Secure and Standardize... 
    Full time
    Work at office
    3 days per week

    GrubHub

    New York, NY
    3 days ago
  • $110k - $120k

     ...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading...  ...on SS&C for expertise, scale, and technology.Kick off your software engineering career on our Quality & Automation team. You... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    New York, NY
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!