Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - GPU reliability

$200k - $300k

Jobleads-US

Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server and network infrastructure. In this role, your main focus will be to develop tools in Python to analyze the performance of GPU hardware and build creative solutions to improve observability, reliability, and efficiency of the fleet. You’ll work closely with other engineering teams to deeply understand research and trading workflows and ensure that GPU infrastructure is utilized optimally. Strong Python skills and development experience are required, along with Unix experience and a background of managing GPU hardware at scale.

Responsibilities

This role offers a unique opportunity to make a significant impact on a critical part of our existing and growing infrastructure. Your responsibilities may vary day to day, but will include:

  • Building and maintaining tools and software features to automate systems engineering workflows related to GPU management, monitoring, metrics collection, maintenance, and network configuration
  • Troubleshooting software and hardware bugs on a fleet of GPU devices, including application, network, operating system, and/or kernel issues
  • Working across HRT’s engineering teams to tune workloads and processes to use GPUs more efficiently
  • Analyzing GPU job statistics to identify trends and areas for improvement

Qualifications

Required:

  • BS and/or MS in computer science or a related field
  • 2+ years of relevant experience, including programming in Python and managing GPUs
  • Experience using automation to solve problems and improve process efficiency
  • Experience working with, troubleshooting, tuning, and deploying various types of GPU hardware
  • Strong grasp of computer science fundamentals and software design patterns
  • Solid understanding of Linux/UNIX operating systems
  • Familiarity with open-source software
  • Ability to debug and analyze problems quickly
  • Skilled at balancing multiple tasks while maintaining meticulous attention to detail
  • Ability to operate effectively as a team player and also work independently
  • Ability to learn at a fast pace and apply new skills effectively

Preferred:

  • Understanding of Debian operating system
  • Familiarity with systems configuration management and monitoring technologies
  • Familiarity with continuous integration and continuous deployment tools and processes
  • Understanding of networking protocols

The estimated base salary range for this position is 200,000 to 300,000 USD per year (or local equivalent). The base pay offered may vary depending on multiple individualized factors, including location, job-related knowledge, skills, and experience.

This role will also be eligible for discretionary performance-based bonuses and a competitive benefits package which includes medical, dental, vision, basic life insurance, and enrollment in our company’s retirement savings plans. Employees will receive sick and parental leave, as well as other paid time off (including 20 vacation days and 10 paid holidays in the US). Please note that benefits and time off policies will vary across non-US locations.

Culture

Hudson River Trading (HRT) brings a scientific approach to trading financial products. We have built one of the world's most sophisticated computing environments for research and development. Our researchers are at the forefront of innovation in the world of algorithmic trading.
At HRT we welcome a variety of expertise: mathematics and computer science, physics and engineering, media and tech. We’re a community of self-starters who are motivated by the excitement of being at the cutting edge of automation in every part of our organization—from trading, to business operations, to recruiting and beyond. We value openness and transparency, and celebrate great ideas from HRT veterans and new hires alike. At HRT we’re friends and colleagues – whether we are sharing a meal, playing the latest board game, or writing elegant code. We embrace a culture of togetherness that extends far beyond the walls of our office.
Feel like you belong at HRT? Our goal is to find the best people and bring them together to do great work in a place where everyone is valued. HRT is proud of our diverse staff; we have offices all over the globe and benefit from our varied and unique perspectives. HRT is an equal opportunity employer; so whoever you are we’d love to get to know you.

Please be advised: Use of AI tools by an applicant during interviews or assessments is strictly prohibited, unless otherwise instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that AI assistance was used by an applicant during an interview or an assessment, we reserve the right to immediately end the interview or assessment, disqualify the applicant’s candidacy and/or rescind any job offers extended.

Voluntary Self-Identification

HRT is committed to providing equal employment opportunities for all groups. We care deeply about cultivating diversity and inclusion within our organization. For this reason, we invite you to voluntarily self-identify by answering the questions below. We do not discriminate based on any of these factors and any information collected will be kept completely confidential during your recruitment process. This information is used solely to support our commitment to equal employment opportunity and to comply with applicable laws that require this information to be summarized and reported to State and Federal Governments for civil rights enforcement purposes.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer - GPU reliability in New York, NY vacancy
  • $200k - $300k

    Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds and maintains the platform that is shared by all Systems teams to provision, monitor, and manage HRT’s server... 
    Suggested
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    2 days ago
  •  ...Hudson River Trading seeks a Software Engineer focused on GPU reliability to join the Systems Development team. You will build Python tools to analyze GPU performance, improve observability, and optimize GPU infrastructure, working with multiple engineering teams to ensure... 
    Suggested

    Jobleads-US

    New York, NY
    1 day ago
  • $120k - $160k

     ...at .About the roleASoftware Engineer contributes to the design, implementation...  ...on improving the efficiency, reliability, and scalability of systems...  ...hardware teams to evolve our GPU performance testing platform...  ...in Go and/or Python software development.Hands-on experience... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    3 days ago
  • $204k - $259k

     ...Software Engineer, GPU Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced... 
    Suggested
    Full time
    Remote work

    Waymo

    New York, NY
    1 day ago
  • $160k - $240k

    Senior Software Engineer - BQL Reliability Engineering Location New York Business Area Engineering and CTO Ref # 10053945 Description & Requirements What You’ll Do: As part of the BQL (Bloomberg Query Language) Reliability Engineering team, you will... 
    Suggested
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    3 days ago
  • $174k - $252k

     ...as system design consulting, developing software platforms and frameworks, capacity...  ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... 

    Google

    New York, NY
    15 hours ago
  • $190k - $260k

     ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...distributed systems with Kubernetes, and GPU workloads on those clustersExperience with... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  •  ...wealth management with you. The Role We're hiring a Senior Software Engineer for the Infrastructure team, the group that owns the AWS...  ...a growing engineering team Set platform, security, and reliability standards, document them, and bring the broader engineering... 

    Jobleads-US

    New York, NY
    2 days ago
  •  ...infrastructure company delivering high-performance GPU compute, inference services, and...  ...We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure...  ...current with emerging GPU hardware and software technologies, integrating improvements... 

    GMI Cloud

    New York, NY
    5 hours ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management.... 
    Flexible hours

    Baseten

    New York, NY
    1 day ago
  •  ...in the United States is seeking an experienced Infrastructure GPU Engineer to build and support high-performance cloud infrastructure....  ...optimizing resource allocation for GPU workloads, ensuring system reliability, and collaborating with cross-functional teams. The position... 
    Remote job

    DevOpsChat

    New York, NY
    4 days ago
  •  ...digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver...  ...that lead to real fixes.QUALIFICATIONS6+ years in SRE, platform reliability or observability engineering, with strong hands-on AWS... 

    Appnovation

    New York, NY
    3 days ago
  • $174k - $252k

     ...Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...experience. 5 years of experience with software development in one or more programming...  ...or Engineering. About the job Site Reliability Engineering (SRE) is what you get when... 

    Google

    New York, NY
    23 hours ago
  •  ...Senior Service Reliability Engineer Fitch Group is currently seeking a Senior Service Reliability...  ...includes the Chief Data Office, Chief Software Office, Chief Technology Office, Emerging...  ...at scale: SageMaker endpoints, GPU node groups, autoscaling, and Kubernetes... 
    Temporary work
    Work at office
    Immediate start
    2 days per week
    3 days per week

    Fitch Group

    New York, NY
    23 hours ago
  • $130k - $190k

     ...quantitative disciplines to deliver high-impact results for our clients. About the Role: Summary Responsible for the operational reliability, observability, and stability of the Strategic Full Revaluation Capability (SFRC) batch platform. This role acts as the first... 
    Temporary work
    Remote work
    Worldwide

    BIP US

    New York, NY
    3 days ago
  • $140k - $215k

     ...intersection of our Core Platform and Embedded Reliability charters: building the foundational...  ...while embedding directly with product engineering teams and their leadership to drive...  ...the open source community; evangelize software engineering best practices, especially... 
    Full time
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    New York, NY
    2 days ago
  • $143k - $210k

     ...managed storage products. We build reliable, scalable storage solutions...  ...leading performance. Storage engine works with engineering teams...  ...with technologies such as RDMA, GPU Direct Storage, and...  ...leveraging AI tools to augment software development.Familiarity with storage... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    1 day ago
  • $153k - $204k

     ...Cloud secure by design, from data centers and GPU fleets to the platform layers powering our customers...  ...cryptographic services that are secure, reliable, and easy to use at scale.About the RoleAs a Senior Software Engineer on the PKI & Secrets team, you will shape how... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    3 days ago
  • $182k - $242k

     ...Learn more at .What You'll Do:The Runtime & GPU Systems team builds and operates secure,...  ..., GPU infrastructure, and Linux systems engineering. We partner closely with security,...  ...diagnosing and resolving complex performance, reliability, or isolation issues across containers,... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    2 days ago
  • $182k - $242k

     ...at .About the RoleWe are seeking Senior Software Engineers who specialize across the pillars of...  ...infrastructure, including highly scalable and reliable logging, metrics, and tracing platforms...  ...., large-scale training and inference, GPU-based infrastructure, MLOps tooling) is... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    15 hours ago
  • $165.3k - $219.68k

     ...insights to improve their business. Founded by engineers — and customer obsessed — we leap at...  ...Model APIs, to state a few.Improve reliability, latency, and efficiency of distributed...  ...real-time serving, ML infrastructure, or GPU orchestrationExposure to platforms like... 
    Local area
    Worldwide

    DataBricks

    New York, NY
    3 days ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership... 

    JP Morgan Chase

    New York, NY
    23 hours ago
  • $300k - $400k

     ...want to build it. About the Role We're hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our...  ...with capacity planning and cost optimization for GPU or TPU infrastructure Familiarity with inference-specific... 
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    New York, NY
    1 day ago
  • $150k - $160k

    Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join the Engineering team. This position is located in our New York, NY office; three (3) days in office depending on business needs... 
    Work at office
    Local area

    Haymarket Media Group

    New York, NY
    2 days ago
  • $182k - $250.8k

     ...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great...  ...standards that embed observability, resilience, and software engineering rigor into all engineering... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    New York, NY
    3 days ago
  •  ...Responsibilities Design and architect scalable, highly available, and resilient AWS infrastructure. Lead cloud infrastructure engineering and reliability initiatives across AWS environments. Develop and maintain Infrastructure as Code (IaC) using Terraform and AWS... 
    Contract work

    2T Consulting

    New York, NY
    3 days ago
  •  ...Lightning AI is seeking a Senior Backend Engineer for the Managed Services team to design, build, and operate backend services that power our GPU infrastructure platform. You will develop control planes and automation to provision and scale Kubernetes and Slurm clusters... 
    Work at office
    2 days per week

    Jobleads-US

    New York, NY
    1 day ago
  • $300k - $400k

     ...for generalist infrastructure and systems engineers to help build the systems that power our...  ...underlying infrastructure for the clusters to reliably and safely train frontier models....  ...and running large Kubernetes clusters with GPU workloads, or building infrastructure to... 
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours

    Thinking Machines Lab

    New York, NY
    2 days ago
  •  ...the banking core to support customer accounts at scale. As an engineering manager, you will lead a team of 4 to 8 senior engineers to...  ...and partner across product, policy, and operations to ensure reliability and usability. You will translate direction into measurable... 

    Jobleads-US

    New York, NY
    2 days ago
  •  ...Skip To Main ContentBack to SearchRemote in United States of America: New YorkSite Reliability Engineering& 4 othersLooking for something else?Find a vacancy that works for you. Send us your CV to receive a personalized offer.Find me a jobWe are seeking a specialized Observability... 

    EPAM Systems Inc

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - GPU reliability. Be the first to apply!