GPU Distributed System Researcher
AMD
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE:AMD Research and Advanced Development is seeking a researcher to invent and implement novel programming models and runtime environments for high performance computing systems comprised of heterogeneous processors, accelerators, and scale-up/scale-out networks.THE PERSON:You love inventing and implementing new features, APIs and abstractions to increase efficiency and optimize performance. You love bringing new systems to life and developing software that shows the full power of novel hardware architectures. You want to help define the software stack of AMD’s future hardware accelerators. Are you ready for the challenge? Then join us!KEY RESPONSIBILITIES:Invent, design and implement novel programming models and runtime environments for high performance computing systems comprised of heterogeneous processors and accelerators. Implement innovative software solutions and demonstrate their effectiveness on prototype hardware.Work on software enhancements that aim to potentially improve programmer productivity.Drive the results of the research projects into product roadmaps.Collaborate with teams within AMD Research and Advanced Development, product groups, and vendor partners to help improve AMD’s ML/HPC ecosystem.Submit patentable inventions.Clearly communicate research findings to academic, product and commercial audiences.PREFERRED EXPERIENCE:Must have experience in developing and debugging GPU and/or multi-threaded CPU applications. Distributed system experience is desired as well.Prior exposure to ML/HPC/datacenter software, middleware and device driver development, and familiarity with ML frameworks, OpenSHMEM, or MPI programming models are preferred.The position involves investigating networking architectures with an emphasis on optimizing cluster-scale applications that are bounded by memory and/or inter-GPU communication performance.Demonstrated experience in ideation, evaluation, and optimization of a research project.Excellent written and oral communication skills, ability to organize and present complex technical information.Strong analytical skills.Strong skills in C/C++, Python and/or GPU programming.Use of networking simulation platforms (NS-3, OMNeT++, etc.).Experience with High Performance Computing networking.Experience with the AMD ROCm software stack.Version control systems such as Git. ACADEMIC CREDENTIALS:PhD degree in Computer Engineering / Electrical Engineering preferred.LOCATION: Bellevue, WA preferred; San Jose, CA or Austin, TX are possibilitiesThis role is not eligible for visa sponsorship.#LI-MR1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$150k - $290k
...company building integrated task systems — fusing hardened hardware... ...We are seeking a talented ML Researcher / Research Engineer to... ...Knowledge of CUDA programming and GPU optimization Experience with... ...CD pipelines Experience with distributed training and large-scale data...SuggestedPermanent employmentContract work$153.9k - $300.96k
...portfolio includes relational databases, distributed caches, key-value stores, document... ...latency, and highly available distributed systems to support mission-critical business operations... ...with strong database systems or AI research experience. Proficiency in one or more...SuggestedTemporary workLocal area- A leading technology company is seeking a Senior Research Scientist for System Software and I/O Architecture in Seattle. The role involves developing and analyzing scalable multi-GPU platforms, contributing to innovative software solutions, and collaborating across diverse...Suggested
$173k
...Machine Learning Scientist, Agentic AI & GenAI Systems (Growth Marketing)We create and deliver... ..., and marketing interactions, leveraging distributed representation learning and real-time... ...Claude).Strong background in distributed GPU training and inference, cloud infrastructure...SuggestedFull timeTemporary work$184k - $287.5k
...computing. An era in which our GPU acts as the brains of... ...building the next generation of AI systems that can perceive, reason about... ...role operates at the applied-research boundary: developing and... ...multi-node environments and distributed training workflows Practical...Suggested$171.5k - $245k
...devices, and applications in any location. Distributed across 160+ public exchanges globally... ...We are looking for a Principal Threat Researcher to join our team. This is a remote role,... ...data sources across at least one operating system (Windows, macOS, Linux) and major Cloud/...Full timeLive inWork at officeLocal areaRemote work$150k - $190k
...hypergrowth company. About The Role We’re hiring an Executive Researcher to build a modern, AI-powered research and strategic... ...maps for CoreWeave’s critical functions — AI infrastructure, GPU cloud, ML systems, datacenter operations, GTM, product, and G&A leadership....Full timeTemporary workWork at officeFlexible hoursShift work$263.7k - $333.4k
...San Jose office.Meet the TeamThe Cisco AI Research team is a dynamic group of scientists... ..., agentic workflows, and self-improving systems. You will bridge the gap between theoretical... ...models, dictating dataset curation, distributed training strategies, and alignment execution...Full timeTemporary workWork at officeLocal areaFlexible hours$192k - $304.75k
We are now looking for a Research Scientist with a focus in System Software and I/O!NVIDIA is seeking Research Scientists with a focus in System Software... ...development of future fast, scalable storage accesses by GPU threads. Scalable systems in a post-Moore world require...Full timeWork experience placement- ...invent that future.You will lead the research and development of frontier AI systems that combine foundation models,... ...learningRetrieval-Augmented Generation (RAG)Distributed training and inferenceLarge-scale... ...AIHugging FaceDistributed GPU training infrastructureYou'll...Work at officeImmediate startRemote work
$70k - $95k
...s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic... ...-starting, action-oriented, and highly motivated Security Researcher to join our Counter Adversary Operations Team. This...Work experience placementWork at officeLocal area- ...About the Role As a Distribution Center Associate, you'll be a crucial part of our logistics operations, ensuring that products are efficiently... ...products in designated locations, using inventory management systems to track stock levels. Order Picking: Accurately select...
- ...’s database development team is seeking PhD-qualified researchers to design and build distributed NoSQL databases, including caches, key-value, document... ...databases. You will deliver high-performance, low-latency systems for global infrastructure in a cloud-native...
$145.46k - $193.94k
...Black Lotus Labs is seeking a remote Threat Intelligence Researcher on the Research & Analysis team focused on tracking advanced... ...incident response. Software development experience with Docker, distributed data technologies such as Hadoop or Spark, cloud AI services,...Full timeTemporary workRemote work$125.5k - $190k
...Research Scientist At Chewy, we believe great work starts with great people. Here, you... ...research, data science, and supply chain systems to solve complex fulfillment and... ...Experience working with large datasets and distributed computing environments ~ Strong problem...Local areaFlexible hours- ...AMD is looking for a strategic research and software engineering... ...models, large-scale computing systems, hardware architectures, and... ...to model training, inference, distributed computing, algorithms, memory... ...based AI and HPC environments, GPU computing, heterogeneous systems...
$80 - $85 per hour
...currently seeking a highly motivated Perception/Machine Learning Researcher for an onsite role at Redmond, WA Position Title: Perception/... ...focuses on developing and deploying advanced machine learning systems for egocentric human behavior understanding in AR and VR —...Daily paidTemporary workWork experience placementImmediate start- ...devices, and applications in any location. Distributed across 160+ public exchanges globally... ...Role We are looking for a Staff Threat Researcher to join our team. This is a remote role... ...data sources across at least one operating system (Windows, macOS, Linux) and major Cloud/...Full timeRemote work
$142.8k - $193.2k
...a critical role in driving the design, research, and development of these science initiatives... ...learning feedback* Build recommender systems to generate personalized learning... ...optimization, data mining, parallel and distributed computing, high-performance computingPreferred...Flexible hours$167.1k - $226.1k
...dollars at Amazon Scale worldwide! The systems we build are entirely in-house, and are... ...forefront of both academic and applied research in large scale supply chain planning, optimization... ...per seconds, to large scale distributed system that optimize Amazon’s fulfilment...Temporary workWorldwideFlexible hours$192.2k - $260k
...To make that possible, we're creating a system that can automatically generate a simulation... ...robotics.Role OverviewThis is applied research with a short distance from idea... ...scipy etc.- Experience with large scale distributed systems such as Hadoop, Spark etc.Amazon...Local areaFlexible hours$112k - $151k
...Siemens builds the systems the physical world runs on: factories, power grids, buildings... ...at the intersection of machine learning research, real world data, and production systems... ...transfers Experience working with globally distributed research, product, or engineering...Local area- ...Applied Scientist, you will lead the architecture, research, and productization of these next-generation ML systems, bridging deep research with deployment at scale... ...evaluation.• Experience with large datasets, distributed training or inference, scalable architecture design...Work at officeImmediate startRemote work
- ...machine learning models and systems to protect our users from the... ...for this direction and produce research outcomes with significant... ....).- Strong understanding of distributed computing framework & performance... ...Have a deep understanding of GPU and/or other AI accelerators,...Flexible hoursShift work
$173k - $248.4k
...At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.As a Senior User Researcher at Snowflake, you’ll shape what we build and how customers experience it. You’ll work alongside product, design, engineering, and...$152.2k - $205.9k
...role sits at the forefront of that transformation. You will lead research that shapes agentic products and experiences across AWS —... ...Researcher with deep experience in agentic tools and large-scale systems. You will deliver insights that inform Experience Health metrics...Flexible hoursShift work$192k - $304.75k
...looking for a passionate AI research scientist with deep quantum computing... ...of fault-tolerant quantum systems powered by machine learning.... ...with CUDA and NVIDIA GPU programming for accelerating... ...computing (HPC) environments and distributed training frameworks (e.g., PyTorch...Full time$143k - $286k
...from reactive to proactive by building systems that operate in the real world. As a Staff... .... Familiarity with MLOps practices, distributed training, or inference optimization (especially... ...). Ability to balance cutting edge research with real world operational needs. A track...Full timeTemporary workPart timeWorldwide$117.8k - $160k
...experience to delight hundreds of millions of customers around the globe. We are seeking an AI-fluent UX Designer to join our Design Systems team and drive the day-to-day evolution of the Prime Design Library (PDL). In this role, you will be a builder and maintainer of...Contract workWorldwideFlexible hours$137.8k - $186.4k
...Engagement (PE) experiences made up of forward-thinking UX designers, researchers and writers. PEX is revolutionizing the user experience for... ....As a Senior UX Designer on the employee-facing design systems team, you are responsible for leading design initiatives, including...Immediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Distributed System Researcher. Be the first to apply!




