Applied Scientist, Agent Evaluation & Adaptive Model Routing
Bitdeer Technologies Group
About Bitdeer:
Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.About Bitdeer AI Lab:
Bitdeer AI Lab is a frontier AI lab under Bitdeer, a global-leading computing power solutions provider. Guided by long-termism, we are committed to exploring the frontiers of artificial intelligence with the ambition, courage, and determination to build technologies that can truly change the world.
Our mission is to turn energy into intelligence that people can actually afford to use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving decide what gets built at all. We work on this from the ground up, from the power and datacenters we own to the software that turns them into tokens — and we continue to invest in and expand the infrastructure behind it
What you will be responsible for:
- This role builds the evaluation and decision systems that make agentic inference measurable, reliable, and adaptive. You will own and extend our LLM and agent evaluation pipeline, develop representative task suites, and measure behavior at the trajectory level across task success, tool use, quality, cost, latency, and token consumption. You will research and prototype adaptive model-routing strategies inside a shared agent harness, including rule-based baselines, cascades, stage-aware routing, uncertainty-aware selection, and escalation and recovery policies. You will take promising methods from offline evaluation through traffic replay, shadow testing, and internal pilots, while working closely with the MaaS and platform engineering teams on production integration. The role owns evaluation methodology, routing policy, and research prototypes; production gateway reliability, billing, access control, and SLA engineering are collaborative responsibilities with the MaaS and platform teams
How you will stand out:
- Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Statistics, Electrical Engineering, or a related field, with substantial hands-on experience in LLM evaluation, agentic systems, applied machine learning, or adaptive inference
- Strong programming ability in Python and practical experience with PyTorch and modern data and evaluation tooling; able to independently build reliable experimental pipelines and internal research prototypes
- Demonstrated experience designing or operating evaluation pipelines for LLMs or agents, including task-level and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis
- In addition to hands-on LLM or agent evaluation experience, candidates should have implementation-level depth in at least one core area: model selection and routing, uncertainty estimation and calibration, cascading and escalation, or stage-aware agent inference. Experience with preference modeling, contextual bandits, online learning, or broader adaptive inference methods is a plus.
- Hands-on experience evaluating multi-turn or tool-using agents, including task completion, tool-call correctness, planning failures, recovery behavior, and cost and latency trade-offs
- Rigorous experimental practice — controlled comparisons, honest baselines, statistical analysis, and the ability to explain clearly what an evaluation result does and does not prove
- Ability to analyze per-request and per-trajectory quality, cost, latency, and token usage, and construct cost-quality Pareto frontiers rather than relying only on aggregate model scores
- Experience with multi-model APIs, agent harnesses, traffic replay, shadow evaluation, A/B testing, or production model monitoring is highly preferred
- Familiarity with model-specific differences in tool calling, context windows, prompt caching, reasoning controls, and inference systems such as vLLM or SGLang is a plus
- Publications at top-tier ML, NLP, or systems venues, or substantial open-source contributions in evaluation, agents, routing, or inference, are welcome
- Strong ownership and product judgment, with a track record of taking ambiguous research questions from problem definition through a working 0-to-1 prototype and measurable internal validation
What you will experience working with us:
- A culture that values authenticity and diversity of thoughts and backgrounds;
- An inclusive and respectable environment with open workspaces and exciting start-up spirit;
- Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
- Ability to contribute directly and make an impact on the future of the digital asset industry;
- Involvement in new projects, developing processes/systems;
- Personal accountability, autonomy, fast growth, and learning opportunities;
- Attractive welfare benefits and developmental opportunities such as training and mentoring.
- ...based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical... ...an academic background in mathematics, have 2+ years of applied experience, and strong communication skills. This role is remote...SuggestedRemote work
- ...where that happens: every product built on a model is bounded by what it costs to run, so... ...are all in scope. You will implement and adapt published methods on our models and... ...they fall short. You will also build the evaluation discipline that makes a claim like “lossless...SuggestedFull timeInternship
$142.8k - $193.2k
...and we're looking for exceptional scientists to help lead the way.The Last Mile Routing & Planning organization develops... ...Learning methods- Proven experience applying these methods to large-scale,... ...business problems- Ability to translate models into production-ready code in...SuggestedFlexible hours$136k - $184k
...Amazon Security is seeking an Applied Scientist to work on GenAI... ...state-of-the-art large language models, development intelligent systems... ...frameworks including multi-agent orchestration, RAG pipelines... ...-built for third-party risk evaluation, security documentation processing...SuggestedFlexible hours- ...We are hiring an Applied Scientist with experience in performing hyperparameter optimizations, evaluating model performance, and deploying production natural language systems. Working with an elite group of professions, you will have the opportunity to work hands on with...Suggested
$114.6k - $234.6k
...Job Description The OCI AI Evaluation Science team builds the evidence behind model-selection, product-readiness, and... ...-augmented generation, AI agents, NL2SQL, multimodal understanding... ...responsible AI. As a Senior Applied Scientist on the team, you will independently...Temporary workFlexible hours- ...Applied Scientist Services at Apple help hundreds of millions of customers get the most out... ...observational testing frameworks, counterfactual modeling, and lifetime value estimation. As a... ...causal inference and AIML research, evaluating and integrating new frameworks where...
- Avride is seeking a data scientist to define and evaluate metrics for autonomous vehicle systems. You will analyze real-world and simulated driving... ...passenger comfort. You will work with cross-functional teams, apply statistical methods, and build scalable pipelines to...
- ...Manager to own AI inference and model serving for k0rdent AI, our... ...placement, autoscaling, routing, lifecycle management, observability... ...performance bottlenecks, evaluating system design trade-offs, and... ...employment decision for the position applied for. You also have the right...
$272k - $431.25k
...new generation of interactive world-model systems. With this release, NVIDIA has... ...need to see:Demonstrated research or applied innovation in world models, video diffusion... ....Hands-on experience building, adapting, or deeply evaluating world models, video generation systems...Full time$159.8k - $244.3k
...Artificial Intelligence Scientist to lead the... ...track record of taking models from problem... ...generative AI, and multi-agent solutions using complex... ...workflow engines, evaluation frameworks, guardrails, model routing, and human-in-the-... ...data movement.Apply strong software engineering...Full timeLocal areaWork from homeRelocation packageFlexible hours$93.75k - $133.2k
...hiring freeze. As an FBI special agent, your career is defined by... ...expand how your expertise is applied, encouraging you to think... ...brings new challenges that demand adaptability and resilience, but you’re... ...enforcement reflects how you evaluate information, navigate complexity...Work at officeLocal areaShift work$76k - $125.3k
...Organization (FSO). Our focused model and bold ambition have... ...working world by applying your knowledge, skills,... ...analysis Compile and evaluate moderately complex data... ...detail The ability to adapt your work style to work... ...an inquiry which will route you to EY’s Talent Shared...Summer holidayFlexible hours$99.6k - $174k
...deployment of AI agents. We also drive the... ...quality telemetry).Apply secure SDLC and privacy... ...practices (threat modeling, least privilege).... ...(RAG, retrieval, routing, tool-use, evals)... ...including safety evaluations, iterative testing... ...-tuning and model adaptation.Familiarity with...Full timeWork at officeRemote work2 days per week$175k - $200k
...adtech algorithms and supporting user acquisition or paid media modeling (highly desired) Strong modeling fundamentals: the ability to... ...Python and SQL Experience 5+ years in a hands‑on, in‑the‑weeds applied data science role delivering measurable business impact. Your Role...Remote work$200k - $240k
A technology company in Austin is looking for a Senior Applied ML Scientist to lead the development of fraud detection models and advance financial risk products. The role requires strong skills in machine learning and data analysis, with a firm understanding of the product...Remote jobFlexible hours$200k - $240k
...example, our engineering team in India works primarily from our Gurugram office. Role As a Senior Applied ML Scientist at SentiLink, you will build our core products: models that identify fraudsters and also advance our growing suite of products in financial risk. As an...Work experience placementLive inWork at officeRemote workHome officeFlexible hours- The Jr. Applied Scientist role is a unique employment opportunity for students seeking to gain on‑the‑job applied science and machine learning experience and receive excellent mentoring while completing their education. Jr. Applied Scientists are immersed in an Amazon team...Full timePart timeSummer workInternshipFlexible hours
- Protopia AI, Inc. is seeking an Applied Scientist to advance privacy-preserving NLP methods and production systems. You will work with LLMs... ...core NLP privacy tech while interface with state‑of‑the‑art models like GPT, BERT, LaMDA, and LLaMA. The role emphasizes fast R&...
$175k - $200k
Launch Potato is seeking a Data Scientist to own the data science engine for the Insurance vertical, focusing on delivering models that drive revenue and media efficiency. With a... ...per year, this role requires 5+ years of applied data science experience and proficiency...- A leading tech company in Austin, TX is offering a Jr. Applied Scientist internship aimed at Master's students. This role emphasizes hands-on applied science and machine learning experience, with one-on-one mentoring from experienced professionals. Candidates will conduct...Full timeInternshipFlexible hours
- ...needs for faculty specialized in computer science, statistics, applied mathematics, and artificial intelligence. The ideal candidate will... ...an environment of open inquiry in both teaching and research, modeling academic freedom and respectful debate. Qualifications - Ph....Full timeSummer work
$106.9k - $200.6k
...opportunity As a Agent Dev Architect, you’ll have... ...results. Evaluate and recommend technologies... ...and data architecture modeling. Business acumen in... ...future with confidence? Apply today. EY accepts applications... ...an inquiry which will route you to EY’s Talent...Full timeSummer holidayFlexible hours- Environmental Scientist, Transmission Line Routing & Siting Halff has an opening for an Environmental professional, specializing in Transmission Line Routing & Siting, to join our growing Environmental team as an Environmental Scientist. The ideal candidate will have at...Temporary workLocal areaFlexible hours
- ...for a full-time Professor of Computer Science, Statistics, or Applied Mathematics (any rank) to begin as early as Summer 2026. Applicants... ..., data science, computer science, and computational modeling. The ideal candidate will be comfortable with interdisciplinary...Full timeSummer work
- Halff Associates, Inc. is seeking an experienced Environmental Scientist specializing in Transmission Line Routing & Siting in Austin, Texas. The ideal candidate should have at least five years of relevant experience and possess a Bachelor's Degree in Environmental Science...
- A new university in Austin is seeking a full-time Professor of Computer Science, Statistics, or Applied Mathematics. The role involves teaching undergraduate courses, mentoring students, and contributing to a culture of open inquiry. Ideal candidates will hold a Ph.D. and...Full time
$60 per hour
...learn how modern AI systems are tested and evaluated, we want to hear from you. Project... ...are seeking QA experts for autonomous AI agents in a project focused on validating and improving... .... Opportunity to influence how future AI models understand and communicate in your field...FreelanceRemote workFlexible hours- ...Join Us as a Medical Laboratory Scientist/MT II - Make an Impact at the... ...role involves analyzing and evaluating individual and related... ...new employees • Ability to adapt to changes in workflow, unusual... ...truly make a difference. Apply today to help us deliver tomorrow...Full timeWork at officeShift work
$60 per hour
...attention to detail. Candidates will review AI evaluation tasks, identify inconsistencies, and define expected behaviors for agents. Ideal applicants have experience in policy... ...and project needs. A great opportunity to influence future AI models! #J-18808-Ljbffr...FreelanceRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied Scientist, Agent Evaluation & Adaptive Model Routing. Be the first to apply!
- materials scientist Austin, TX
- scientist assay development Austin, TX
- entry level research scientist Austin, TX
- health scientist Austin, TX
- quality control scientist Austin, TX
- deep learning scientist Austin, TX
- research associate scientist Austin, TX
- application scientist Austin, TX
- senior analytical scientist Austin, TX
- decision scientist Austin, TX


