Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist-Model Efficiency

Full-time

Bitdeer Technologies Group

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About Bitdeer AI Lab:

Bitdeer AI Lab is a frontier AI lab under Bitdeer, a global-leading computing power solutions provider. Guided by long-termism, we are committed to exploring the frontiers of artificial intelligence with the ambition, courage, and determination to build technologies that can truly change the world.

Our mission is to turn energy into intelligence that people can actually afford to use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving decide what gets built at all. We work on this from the ground up, from the power and datacenters we own to the software that turns them into tokens — and we continue to invest in and expand the infrastructure behind it.

What you will be responsible for:

  • This role makes models cheaper and faster to serve without giving up quality that matters. We are not prescribing the technique — quantization, sparsity and pruning, speculative decoding and MTP, and serving-time attention and KV-cache methods are all in scope. You will implement and adapt published methods on our models and hardware, and develop your own optimizations where they fall short. You will also build the evaluation discipline that makes a claim like “lossless at 2× throughput” defensible.

How you will stand out:

  • Bachelor's, Master's, or PhD in Computer Science, Electrical Engineering, or a related field, with hands-on experience in LLM inference, model optimization, or ML systems
  • Strong programming ability in Python and deep familiarity with PyTorch; experience with C++, CUDA, or Triton is a plus
  • Genuine implementation-level depth in at least one area of model efficiency, such as quantization, sparsity and pruning, speculative decoding and MTP, or serving-time attention and KV-cache methods
  • Strong understanding of transformer internals and where accuracy loss actually shows up in model behaviour
  • Rigorous evaluation practice — task-level metrics, controlled comparisons, honest baselines — and the ability to state honestly what a number does and does not prove
  • Experience taking efficiency methods into production serving, or equivalent research depth, is highly preferred
  • Familiarity with inference engines such as vLLM, SGLang, or TensorRT-LLM and their efficiency features is highly preferred
  • Publications at top-tier systems or ML venues, or substantial open-source contributions, are welcome
  • Deep enthusiasm for cutting-edge AI infrastructure and efficient inference, with a strong ownership mentality and solid engineering discipline

What you will experience working with us:

  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

Vacancy posted 29 days ago
Similar jobs that could be interesting for youBased on the Research Scientist-Model Efficiency in San Jose, CA vacancy
  • $187.04k

    Research Scientist - Multimodal Interaction and World Model - Pre-Training Location: San Jose Responsibilities About Seed Team: Established in 2023, the ByteDance...  ...that optimize both the model's performance and efficiency. Establish scaling laws, design and conduct... 
    Suggested
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    2 days ago
  • $244.8k

     ...toward artificial general intelligence, with research spanning MLLM, GenMedia, AI for Science,...  ...industry-leading general foundation models and multimodal capabilities, powering...  ...optimize both the model\'s performance and efficiency. Establish scaling laws, design and... 
    Suggested

    ByteDance

    San Jose, CA
    1 day ago
  • $162k - $316.8k

    World Model Research Scientist (Intelligent Creation) - Global Frontier Tech Recruitment Program - 2027 Start (PhD) Location: San Jose Employment...  ...Image/video generation and editing; VLM/LLM fine-tuning; Efficient model design and optimization; Reinforcement learning... 
    Suggested
    Temporary work
    Local area

    TikTok

    San Jose, CA
    1 day ago
  • $156k - $316.8k

    Research Scientist — Privacy-Preserving Large-Scale Model Training & Architecture Optimization Location: San Jose Employment Type: Regular Job Code: DW1L Responsibilities...  .... Improve training ETTR / MFU / utilization efficiency under real-world production constraints. Optimize... 
    Suggested
    Temporary work
    Local area

    Ellis Technologies, Inc.

    San Jose, CA
    1 day ago
  • $244.8k

     ...Seed Vision team focuses on foundational models for visual generation, developing...  ...generative models, and carrying out leading research and application development to solve fundamental...  ...Optimize model architectures, training efficiency, and evaluation frameworks. Explore... 
    Suggested
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    1 day ago
  • $187.04k - $359.72k

    Senior Research Scientist, Foundation Model (LLM/ VLLM), TikTok - Trust and Safety Senior Research Scientist, Foundation Model (LLM/ VLLM), TikTok...  ...large-scale training pipelines, improve model serving efficiency, and leverage centralized GPU resources effectively. 5.... 
    Full time
    Temporary work
    Local area

    TikTok

    San Jose, CA
    1 day ago
  • $190k - $250k

     ...transportation committed to a safer and more efficient future for all. The company has...  ...developing large-scale generative world models that learn to predict realistic, physically...  ...trucks. We are looking for a research scientist to lead the design and development of world... 
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    17 hours ago
  • $192.2k - $260k

     ...at Amazon's Delivery Foundation Model team, where you'll work alongside world-class scientists and engineers to pioneer the...  ...Amazon customer, and improving efficiency of Amazon delivery network. - Guide...  ...direction for specific research initiatives, ensuring robust performance... 
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    3 days ago
  • $192k - $304.75k

    We are now looking for an Applied Deep Learning Research Scientist, Efficiency!Join our ADLR - Efficiency team to make deep learning faster and consume...  ...make AI more efficient; we work on the Nemotron series of models to make our state-of-the-art deep learning models the most... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $192k - $304.75k

    NVIDIA is searching for an outstanding Senior Researcher working on efficient deep learning to join our learning and perception research team. We are...  ...are particularly excited about methods for post-training model optimization (pruning, quantization, NAS), efficient architecture... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using...  ...world-class researchers, data scientists, and engineers, tackling the most fundamental...  ...tool-use capabilities. Research efficient multimodal learning techniques,... 

    Institute of Foundation Models

    Sunnyvale, CA
    2 days ago
  • $212.8k

     ...low-power consumption deployment of AI models, strive for excellence and aim to become...  ...round‑the‑clock environmental awareness and efficiently capture real‑time user intentions is...  ...models. Topic Value: Breakthroughs in this research direction will enable intelligent... 
    Temporary work
    Internship
    Local area

    ByteDance

    San Jose, CA
    1 day ago
  • $244.8k

    Responsibilities About the team The Seed Multimodal Interaction and World Model team is dedicated to developing models that have human-level...  ...or simulation environments. Preferred Qualifications Strong research track record in relevant areas. Strong problem-solving and... 
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    1 day ago
  • $212.8k

    Senior Research Scientist (Multimodal Large Language Model) - PICO Location: San Jose Team: Technology Employment Type: Regular Job Code: A57637 About the Team PICO-MR team is dedicated to pioneering core technologies for intelligent human-computer interaction in MR... 
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    3 days ago
  • $254.4k

     ...society. With a long-term vision for the AI sector, the Seed team's research spans MLLM, GenMedia, AI for Science, and Robotics. We maintain...  ...To date, we have launched industry-leading general foundation models and cutting-edge multimodal capabilities. Our technology powers... 
    Temporary work
    Local area
    Worldwide

    ByteDance

    San Jose, CA
    3 days ago
  • $60 per hour

     ...responsible for building machine learning models and systems to protect our users from...  ...GRPO/PPO, overcoming bottlenecks in sample efficiency and training stability Context...  ...paradigm for this direction and produce research outcomes with significant industry influence... 
    Hourly pay
    Summer work
    Internship
    Local area
    Flexible hours
    Shift work

    TikTok

    San Jose, CA
    1 day ago
  • TikTok is seeking a World Model Research Scientist in San Jose to advance generative AI and multimodal model development. You will own research from conception to production, collaborating across teams to deliver scalable foundation models and innovative creative experiences... 

    TikTok

    San Jose, CA
    1 day ago
  • $212.8k - $387.6k

     ...fundamental challenges in LLM development. Our areas of focus include model pretraining, posttraining, inference, memory capabilities,...  ...model capabilities, such as reasoning, code, math. In-depth research and exploration of future use cases. Minimum Qualifications Currently... 
    Temporary work
    Internship

    ByteDance

    San Jose, CA
    2 days ago
  • $168k - $264.5k

    NVIDIA is looking for a passionate researcher in efficient deep learning to join their research team in Santa Clara, California. The role involves researching novel methods for deep learning optimization, publishing original research, and mentoring interns. Candidates should... 

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $165k - $185k

    Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale, California, Pittsburgh, Pennsylvania...  ..., our AI research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language... 
    Work experience placement
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    3 days ago
  • $55 per hour

     ...compilation technologies for AI foundation models. Contribute to infrastructure and...  ...large-scale models. Support development of efficient and scalable components across areas...  ...tooling, or frameworks. Collaborate with researchers and engineers to support research and system... 
    Hourly pay
    Internship
    Local area

    ByteDance

    San Jose, CA
    4 days ago
  • $85 per hour

    Responsibilities The Seed-LLM-Model team is dedicated to foundational algorithm research for LLM models, focusing on issues such as model architecture, optimization...  ..., and stability. This ensures the performance and efficiency of large model training and inference, providing a... 
    Hourly pay
    Full time
    Summer work
    Internship
    Local area

    ByteDance

    San Jose, CA
    3 days ago
  • $84.13 - $91.34 per hour

    AI Researcher - Efficient AI (Contractor) Step into the innovative world of LG Electronics. As a global leader in technology, LG Electronics is...  ...developing technologies that make modern LLMs, VLMs, multimodal models, and AI agents faster, smaller, and more deployable in real-... 
    Full time
    Contract work
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    Santa Clara, CA
    3 days ago
  • LG Electronics is seeking a Contract AI Researcher for the Emerging Technology Lab in Santa Clara, CA. The role focuses on making LLMs,...  ...deployable on edge devices. Work spans compression, quantization, efficient inference, and novel architectures, with collaboration across... 
    Contract work

    LG Electronics North America

    Santa Clara, CA
    2 days ago
  • LG Electronics is seeking a Contract AI Researcher - Efficient AI in Santa Clara, CA (hybrid) to advance on-device AI efficiency for LLMs, VLMs, and multimodal systems. You will bridge research and implementation, transforming cutting-edge ideas into working prototypes... 
    Contract work

    LG Electronics

    Santa Clara, CA
    4 days ago
  • The Role We are looking for a Research Scientist to join the Multi-Embodiment Generalist Agent (MEGA) team within Wayve Science as a founding member. MEGA is building foundation models for general-purpose robots beyond not self-driving vehicles. Our goal is to create intelligent... 
    Full time
    Work from home

    Wayve

    Sunnyvale, CA
    2 days ago
  • $212.8k - $387.6k

    ByteDance is looking for a talented candidate to develop and scale vision foundation models, focusing on image and video modalities. Applicants should be pursuing a Bachelor's or Master's degree in a relevant field, have excellent coding skills in languages like C/C++... 

    ByteDance

    San Jose, CA
    1 day ago
  • ByteDance in San Jose is looking for talented PhD graduates to join their Engineering Architecture team focused on integrating large model technology with real-world applications. This role includes designing technical solutions, collaborating with various teams, and... 

    ByteDance

    San Jose, CA
    1 day ago
  • $126k - $423k

     ...About the role and team We are looking for multiple passionate Research Scientists to join the Research Group at Applied Intuition. The mission...  ...: Conduct research on pretraining world‑action foundation model with various world modalities including vision and physics associated... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Immediate start
    Remote work
    Day shift

    Decisive Point

    Sunnyvale, CA
    3 days ago
  • LG Electronics USA's Emerging Technology Lab in Santa Clara, CA seeks a Contract AI Researcher focused on Efficient AI. You will explore model compression, quantization, and on-device inference to make LLMs and multimodal models faster and lighter for real-world applications... 
    Contract work

    LG Electronics USA

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist-Model Efficiency. Be the first to apply!