Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist-Model Efficiency

Full-time

Bitdeer Technologies Group

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About Bitdeer AI Lab:

Bitdeer AI Lab is a frontier AI lab under Bitdeer, a global-leading computing power solutions provider. Guided by long-termism, we are committed to exploring the frontiers of artificial intelligence with the ambition, courage, and determination to build technologies that can truly change the world.

Our mission is to turn energy into intelligence that people can actually afford to use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving decide what gets built at all. We work on this from the ground up, from the power and datacenters we own to the software that turns them into tokens — and we continue to invest in and expand the infrastructure behind it.

What you will be responsible for:

  • This role makes models cheaper and faster to serve without giving up quality that matters. We are not prescribing the technique — quantization, sparsity and pruning, speculative decoding and MTP, and serving-time attention and KV-cache methods are all in scope. You will implement and adapt published methods on our models and hardware, and develop your own optimizations where they fall short. You will also build the evaluation discipline that makes a claim like “lossless at 2× throughput” defensible.

How you will stand out:

  • Bachelor's, Master's, or PhD in Computer Science, Electrical Engineering, or a related field, with hands-on experience in LLM inference, model optimization, or ML systems
  • Strong programming ability in Python and deep familiarity with PyTorch; experience with C++, CUDA, or Triton is a plus
  • Genuine implementation-level depth in at least one area of model efficiency, such as quantization, sparsity and pruning, speculative decoding and MTP, or serving-time attention and KV-cache methods
  • Strong understanding of transformer internals and where accuracy loss actually shows up in model behaviour
  • Rigorous evaluation practice — task-level metrics, controlled comparisons, honest baselines — and the ability to state honestly what a number does and does not prove
  • Experience taking efficiency methods into production serving, or equivalent research depth, is highly preferred
  • Familiarity with inference engines such as vLLM, SGLang, or TensorRT-LLM and their efficiency features is highly preferred
  • Publications at top-tier systems or ML venues, or substantial open-source contributions, are welcome
  • Deep enthusiasm for cutting-edge AI infrastructure and efficient inference, with a strong ownership mentality and solid engineering discipline

What you will experience working with us:

  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

Vacancy posted 13 days ago
Similar jobs that could be interesting for youBased on the Research Scientist-Model Efficiency in San Jose, CA vacancy
  • $187.04k - $359.72k

     ...Senior Research Scientist, Foundation Model (LLM/ VLLM), TikTok - Trust and Safety Senior Research Scientist, Foundation Model (LLM/ VLLM), TikTok...  ...large-scale training pipelines, improve model serving efficiency, and leverage centralized GPU resources effectively.... 
    Suggested
    Full time
    Temporary work
    Local area

    TikTok

    San Jose, CA
    3 hours ago
  • $190k - $250k

     ...transportation committed to a safer and more efficient future for all. The company has...  ...developing large-scale generative world models that learn to predict realistic, physically...  ...trucks. We are looking for a research scientist to lead the design and development of world... 
    Suggested
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    23 hours ago
  • $192.2k - $260k

     ...at Amazon's Delivery Foundation Model team, where you'll work alongside world-class scientists and engineers to pioneer the...  ...Amazon customer, and improving efficiency of Amazon delivery network. - Guide...  ...direction for specific research initiatives, ensuring robust performance... 
    Suggested
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $192k - $304.75k

    We are now looking for an Applied Deep Learning Research Scientist, Efficiency!Join our ADLR - Efficiency team to make deep learning faster and consume...  ...make AI more efficient; we work on the Nemotron series of models to make our state-of-the-art deep learning models the most... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $192k - $304.75k

    NVIDIA is searching for an outstanding Senior Researcher working on efficient deep learning to join our learning and perception research team. We are...  ...are particularly excited about methods for post-training model optimization (pruning, quantization, NAS), efficient architecture... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $165k - $185k

     ...Company Description The Bosch Research and Technology Center North America with offices in Sunnyvale, California, Pittsburgh, Pennsylvania...  ..., our AI research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural... 
    Work experience placement
    Worldwide

    Bosch Group Inc

    Santa Clara, CA
    1 day ago
  •  ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using...  ...world-class researchers, data scientists, and engineers, tackling the most fundamental...  ...tool-use capabilities. Research efficient multimodal learning techniques,... 

    Institute of Foundation Models

    Sunnyvale, CA
    more than 2 months ago
  • $174k - $252k

     ...checking code in, accuracy, testability, and efficiency).Contribute to existing documentation or...  ...GOLang, Rust, or Java.Experience in ML model coding languages (e.g., Python).Google's...  ...technology forward.The Google Cloud AI Research team addresses AI challenges motivated... 

    Google

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

     ...modern AI, delivering the industry's fastest and most efficient deployment of cutting-edge deep learning models on every NVIDIA GPU. With demand for AI exploding,...  ...collaborative, interfacing directly with NVIDIA Researchers, GPU Architects, and other teams across the... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $120.8k - $193.3k

    DescriptionJob Title: Sr. Simulation & Model Training Engineer - Drone Autonomy Job Location: San Jose, CA (This position requires a...  ...flight.Edge Optimization — Optimize deep learning models to run efficiently on specialized, power-constrained edge hardware.Validation —... 
    Full time
    Work at office

    SiMa Technologies

    San Jose, CA
    3 days ago
  • $224k - $356.5k

     ...cutting-edge infrastructure for large-scale foundation model training in the Generalist Embodied Agent Research (GEAR) group. Our team is leading Project GR00T,...  ...robotics.Optimize GPU and cluster utilization for efficient model training and fine-tuning on massive datasets... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $165.2k - $223.6k

     ...the forefront of running a wide range of models and supporting novel architecture...  ...ensure highest performance and maximize the efficiency of them running on the customer AWS Trainium...  ...with a cross-functional team of applied scientists, system engineers, and product managers... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  •  ...and beyond. Together, we advance your career. THE ROLE:The AI Models and Applications team at AMD is looking for a specialized Sr. Staff...  ...level engineer who is passionate about enabling innovative and efficient Generative AI training/inferencing at scale. You will be part... 

    AMD

    San Jose, CA
    1 day ago
  • $192k - $304.75k

    We are now looking for Senior Research Scientist, Security and Privacy. NVIDIA is seeking exceptional security and privacy researchers to contribute...  ...protection while still maximizing performance and energy efficiency. We are seeking candidates that have a proven track record... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $192k - $304.75k

    We are now looking for a Research Scientist with a focus in System Software and I/O!NVIDIA is seeking Research Scientists with a focus in System...  ..., to achieve high throughput while improving energy efficiency. We are seeking candidates with a proven track record of research... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $192k - $304.75k

     ...'re looking for a passionate scientist at the intersection of quantum...  ...of intelligent, real-time models for fault-tolerant quantum hardware...  .... As a Sr. Quantum Applied Research Scientist, you will help...  ...fine-tuning—including parameter-efficient methods (LoRA, QLoRA, adapters... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $192k - $304.75k

     ...computing. We're looking for a passionate AI research scientist with deep quantum computing expertise...  ...You will research and develop open AI models, curated datasets, and rigorous...  ...training and fine-tuning—including parameter-efficient fine-tuning (LoRA, QLoRA, adapters) and... 
    Full time

    Nvidia

    Santa Clara, CA
    23 hours ago
  •  ...career. THE ROLE:We are hiring an AI Research Scientist, Recursive Self Improvement, AI Safety...  ...engineering-first sense: systems where models, data generators, or toolchains participate...  ...RSI concepts to concrete metrics—data efficiency, robustness, regression rates—not open... 
    Shift work

    AMD

    Santa Clara, CA
    4 days ago
  •  ...building a Data-AML-Engine Orchestration platform to power online model serving across TikTok and other ByteDance products. You will...  ...scalable infrastructure—enabling low latency, high availability, and efficient GPU usage while tackling challenging distributed systems... 

    ByteDance

    San Jose, CA
    3 hours ago
  •  ...leader for deep learning applications in Advertising. This role involves applying cutting-edge research to develop methodologies for complex problems, focusing on conversion modeling and optimization. Ideal candidates will have a PhD and extensive experience in statistical... 

    Roku, Inc.

    San Jose, CA
    4 hours ago
  • $150k - $300k

     ...scale. The Silicon Valley Research Lab focuses on developing...  ...new knowledge and skills as efficiently as humans. We are looking for...  ...) , etc.   As a Research Scientist in the team, you will conduct...  ...implement and evaluate algorithms, models and prototypes of AI systems... 
    Full time
    H1b
    Work at office
    3 days per week

    Horizon Robotics

    Cupertino, CA
    more than 2 months ago
  • $136.8k - $259.2k

     ...Research Scientist Graduate (3D/4D Reconstruction/Generation/Relighting) - 2026 Start (PHD) Location...  ...in MR/XR scenarios; Develop efficient and scalable 3D/4D reconstruction/generation...  ...formulation, dataset construction, model training and optimization, inference acceleration... 
    Temporary work
    Local area

    Pangleglobal

    San Jose, CA
    4 hours ago
  • $163k - $236k

    Analyze efficiency across the multi-modal product family (e.g., TPS/QPS per chip, per GSU) and turn it into data-driven capacity and pricing...  ...and set Provisioned Throughput (PT) burndown rates for new model launches to reflect actual capacity consumed.Introduce spend-based... 

    Google

    Sunnyvale, CA
    5 days ago
  • $272k - $431.25k

     ...FlashDreams are the foundation for a new generation of interactive world-model systems. With this release, NVIDIA has set the standard for...  ...OmniDreams useful in real workflows, and helps customers and researchers succeed with the stack. Success means more than shipping... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $128k - $256k

     ...Responsibilities Volcano Ark is an all-in-one large model service platform launched by Volcano Engine. It is a leading platform...  .... Design reinforcement learning systems to improve training efficiency and provide user-friendly reinforcement learning training interfaces... 
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    4 hours ago
  • $272k - $431.25k

     ...future, we’re generating it! Our world model team is pushing the boundaries of multimodal...  ...AI. We are looking for a Senior Research Manager to lead world-model evaluation and...  ...you’ll be doing:Lead a team of Research Scientists focused on world-model evaluation, benchmarking... 
    Full time

    Nvidia

    Santa Clara, CA
    23 hours ago
  • $224k - $356.5k

     ...performance computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in...  ...technical challenges and communicate effectively across research, engineering, and product teams.Ways to stand out from the crowd... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $174.72k - $295.68k

     ...edge R&D in AI, machine learning, and smart connectivity.We are looking for a full-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-Language-Action (VLA) Foundation Model — the core... 
    Full time

    XPENG Motors

    Santa Clara, CA
    1 day ago
  • $174.72k - $295.68k

     ...seeking Machine Learning Engineers with strong expertise in generative modeling and large-scale deep learning systems, along with solid software development skills. In this role, you will research, implement, and evaluate world models that learn the dynamics of the physical... 
    Full time

    XPENG Motors

    Santa Clara, CA
    1 day ago
  • $125k - $201.25k

     ...DevelopmentJob Sub Function: Clinical Development & Research - Non-MDJob Category:Scientific/...  ...team as a Staff Clinical Research Scientist located in Irvine, CA or Milpitas, CAFueled...  ...critical thinking, who are inventive, efficient and methodical, with a desire to work with... 
    Full time
    Work experience placement
    Local area
    Immediate start

    Johnson & Johnson

    Milpitas, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist-Model Efficiency. Be the first to apply!