Applied Scientist, Agent Evaluation & Adaptive Model Routing
Bitdeer Technologies Group
About Bitdeer:
Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.About Bitdeer AI Lab:
Bitdeer AI Lab is a frontier AI lab under Bitdeer, a global-leading computing power solutions provider. Guided by long-termism, we are committed to exploring the frontiers of artificial intelligence with the ambition, courage, and determination to build technologies that can truly change the world.
Our mission is to turn energy into intelligence that people can actually afford to use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving decide what gets built at all. We work on this from the ground up, from the power and datacenters we own to the software that turns them into tokens — and we continue to invest in and expand the infrastructure behind it
What you will be responsible for:
- This role builds the evaluation and decision systems that make agentic inference measurable, reliable, and adaptive. You will own and extend our LLM and agent evaluation pipeline, develop representative task suites, and measure behavior at the trajectory level across task success, tool use, quality, cost, latency, and token consumption. You will research and prototype adaptive model-routing strategies inside a shared agent harness, including rule-based baselines, cascades, stage-aware routing, uncertainty-aware selection, and escalation and recovery policies. You will take promising methods from offline evaluation through traffic replay, shadow testing, and internal pilots, while working closely with the MaaS and platform engineering teams on production integration. The role owns evaluation methodology, routing policy, and research prototypes; production gateway reliability, billing, access control, and SLA engineering are collaborative responsibilities with the MaaS and platform teams
How you will stand out:
- Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Statistics, Electrical Engineering, or a related field, with substantial hands-on experience in LLM evaluation, agentic systems, applied machine learning, or adaptive inference
- Strong programming ability in Python and practical experience with PyTorch and modern data and evaluation tooling; able to independently build reliable experimental pipelines and internal research prototypes
- Demonstrated experience designing or operating evaluation pipelines for LLMs or agents, including task-level and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis
- In addition to hands-on LLM or agent evaluation experience, candidates should have implementation-level depth in at least one core area: model selection and routing, uncertainty estimation and calibration, cascading and escalation, or stage-aware agent inference. Experience with preference modeling, contextual bandits, online learning, or broader adaptive inference methods is a plus.
- Hands-on experience evaluating multi-turn or tool-using agents, including task completion, tool-call correctness, planning failures, recovery behavior, and cost and latency trade-offs
- Rigorous experimental practice — controlled comparisons, honest baselines, statistical analysis, and the ability to explain clearly what an evaluation result does and does not prove
- Ability to analyze per-request and per-trajectory quality, cost, latency, and token usage, and construct cost-quality Pareto frontiers rather than relying only on aggregate model scores
- Experience with multi-model APIs, agent harnesses, traffic replay, shadow evaluation, A/B testing, or production model monitoring is highly preferred
- Familiarity with model-specific differences in tool calling, context windows, prompt caching, reasoning controls, and inference systems such as vLLM or SGLang is a plus
- Publications at top-tier ML, NLP, or systems venues, or substantial open-source contributions in evaluation, agents, routing, or inference, are welcome
- Strong ownership and product judgment, with a track record of taking ambiguous research questions from problem definition through a working 0-to-1 prototype and measurable internal validation
What you will experience working with us:
- A culture that values authenticity and diversity of thoughts and backgrounds;
- An inclusive and respectable environment with open workspaces and exciting start-up spirit;
- Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
- Ability to contribute directly and make an impact on the future of the digital asset industry;
- Involvement in new projects, developing processes/systems;
- Personal accountability, autonomy, fast growth, and learning opportunities;
- Attractive welfare benefits and developmental opportunities such as training and mentoring.
- ...where that happens: every product built on a model is bounded by what it costs to run, so... ...are all in scope. You will implement and adapt published methods on our models and... ...they fall short. You will also build the evaluation discipline that makes a claim like “lossless...SuggestedFull timeInternship
$142.8k - $193.2k
...and we're looking for exceptional scientists to help lead the way.The Last Mile Routing & Planning organization develops... ...Learning methods- Proven experience applying these methods to large-scale,... ...business problems- Ability to translate models into production-ready code in...SuggestedFlexible hours$183.8k - $248.7k
Twitch is looking for an Applied Scientist to lead computer vision and video manipulation work within... ...(APM). You will apply large language models (LLMs), vision-language models (VLMs),... ..., and you will have the autonomy to evaluate both first-party and third-party solutions...SuggestedFlexible hoursDay shiftAfternoon shift$114.6k - $234.6k
...Job Description The OCI AI Evaluation team builds the evidence behind model-selection, product-readiness, and... ...retrieval-augmented generation, AI agents, NL2SQL, multimodal... ...responsible AI. As a Senior Applied Scientist on the team, you will independently...SuggestedTemporary workFlexible hours- ...Applied Scientist Services at Apple help hundreds of millions of customers get the most out... ...observational testing frameworks, counterfactual modeling, and lifetime value estimation. As a... ...causal inference and AIML research, evaluating and integrating new frameworks where...Suggested
- ...Manager to own AI inference and model serving for k0rdent AI, our... ...placement, autoscaling, routing, lifecycle management, observability... ...performance bottlenecks, evaluating system design trade-offs, and... ...employment decision for the position applied for. You also have the right...
- We are hiring an Applied Scientist with experience in performing hyperparameter optimizations, evaluating model performance, and deploying production natural language systems. Working with an elite group of professions, you will have the opportunity to work hands on with...
$272k - $431.25k
...new generation of interactive world-model systems. With this release, NVIDIA has... ...need to see:Demonstrated research or applied innovation in world models, video diffusion... ....Hands-on experience building, adapting, or deeply evaluating world models, video generation systems...Full time$159.8k - $244.3k
...Artificial Intelligence Scientist to lead the... ...track record of taking models from problem... ...generative AI, and multi-agent solutions using complex... ...workflow engines, evaluation frameworks, guardrails, model routing, and human-in-the-... ...data movement.Apply strong software engineering...Full timeLocal areaWork from homeRelocation packageFlexible hours- ...standard business hours with a hybrid work model (4 days in-office, 1 day working from... ...external consultants on validation activities. Evaluate model performance monitoring and complete... ...engineering. Experience developing and applying machine learning models. 5+ years of...Full timeWork at officeWork from homeRelocation
$99.6k - $174k
...deployment of AI agents. We also drive the... ...quality telemetry).Apply secure SDLC and privacy... ...practices (threat modeling, least privilege).... ...(RAG, retrieval, routing, tool-use, evals)... ...including safety evaluations, iterative testing... ...-tuning and model adaptation.Familiarity with...Full timeWork at officeRemote work2 days per week- ...needs for faculty specialized in computer science, statistics, applied mathematics, and artificial intelligence . The ideal candidate... ...an environment of open inquiry in both teaching and research, modeling academic freedom and respectful debate. Qualifications - Ph.D....Full timeSummer work
$204k - $216k
...research role focused on the models at the foundation of... .... You study, adapt, and advance the foundational... ...experiments, evaluating models, adapting them... ...Foundational Model Research Data Scientist does that. You run the... ...models or applied NLP. Experience with...- ...guidance on program planning, execution, and evaluation for the compliance, accounts management... ...been formulated. Interprets, develops, adapts, and coordinates analytical reviews of... ...impact on operations and resources. Applies and disseminates newly developed or revised...Remote work
- Environmental Scientist, Transmission Line Routing & Siting Halff has an opening for an Environmental professional, specializing in Transmission Line Routing & Siting, to join our growing Environmental team as an Environmental Scientist. The ideal candidate will have at...Temporary workLocal areaFlexible hours
$112k
...everywhere.Machine Learning Scientists IIWithin the AI & Data... ...and operationalizes applied science driven... ...the Marketplace Health Model system to identify behaviors... ...impact.Proactively evaluate opportunities based on... ...applications or LLM-powered agents.Experience with trust...Full time$115.8k - $144.7k
POSITION SUMMARY: The Senior Design Transfer Scientist will lead and support the transition of... ...: Lead change control activities to evaluate and determine the impact of design changes... ...qualified applicants are encouraged to apply, and will be considered without regard...Work at officeImmediate startWorldwide$60 per hour
...attention to detail. Candidates will review AI evaluation tasks, identify inconsistencies, and define expected behaviors for agents. Ideal applicants have experience in policy... ...and project needs. A great opportunity to influence future AI models! #J-18808-Ljbffr Mind RiftRemote jobFreelance- ...Right Of Way Of Agents And Sr. Right Of Way Agents Coates Field Service, Inc. is seeking... ...way for private landowners, and able to adapt to tight deadlines to meet project... ...maps, electronic and paper. Ability to evaluate, interpret, and analyze engineering and right...Daily paidTemporary workWork at office
$97.3k - $125.54k
...civilian hiring freeze. As an FBI special agent, you'll directly impact national... ...brings new challenges that demand your adaptability and resilience, but you're not alone in... ...that prioritizes you. Set yourself apart. Apply today. Salary Level $97,300.00–$125,544....Work at officeLocal area- ...optical metrology applications in semiconductor manufacturing; apply robust statistical techniques to optimize applications for the... ...environment: and provide region support for new product introduction and evaluation activities (beta, head-to-head). Why Nova: Certified Best...
$30 per hour
...We are seeking a highly motivated AI Agent Intern to join Oracle's Supply Chain Applications... ...demo scripts. Research & Analysis Evaluate logistics AI use cases (forecasting,... ...Manufacturing. Engage with Oracle AI Champions, Applied Technologists, and Product Strategy....Hourly payTemporary workInternshipFlexible hours$10k
...Border Patrol Agent (BPA) - Experienced (GL-9 GS-11) - New Hire Sign-On and Retention... ...the next higher grade level (without re-applying) once you successfully complete 52 weeks... ...transcripts, etc.) to submit. You will be evaluated based on your resume, supporting...Full timeLocal areaImmediate startRelocationNight shift$10k
...Border Patrol Agent (BPA) - Experienced (GL-9 GS-11) - New Hire Sign-On and Retention... ...the next higher grade level (without re-applying) once you successfully complete 52 weeks... ...transcripts, etc.) to submit. You will be evaluated based on your resume, supporting...Full timeLocal areaImmediate startRelocationNight shift$10k
...Border Patrol Agent (BPA) – in the Federal Security and Public Safety Sector (Entry Level... ...the completion of required training and apply these skills in a law enforcement capacity... ..., etc.) to submit. You will be evaluated based on your resume, supporting documents...Full timeWork experience placementImmediate startRelocationNight shift- ...reinforcement learning environments that evaluate AI models on complex software engineering tasks... ...golden reference solutions. Evaluate AI agents' ability to reason through complex... ...Minimum weekly submission requirements may apply. Availability Selected experts...Remote jobFor contractors
- ...discretion, vigilance, and precision, our agents ensure a secure and seamless living... ...nuances of residential environments and can adapt to dynamic family routines and estate operations... ..., and discretion, we invite you to apply and become part of our elite Residential...For contractors
- ...The Role As an AI Agent Engineer, you will design... ...consistently. Apply enterprise platform guardrails... ...regulated data. Evaluate incoming integration... ...interfaces, manually routed workflows, platform-specific... ...for large language models used within enterprise...Full timeLocal areaWork from homeRelocation package
$29.89 - $37.37 per hour
...Environmental Scientist Under limited supervision and using comprehensive knowledge, conduct... ...prepare reports and memos based on data evaluation Review and interpret policies, codes,... ...initial training period. Exceptions may apply subject to the business needs of the...Remote workMonday to Friday$73 per hour
...realistic tasks that push frontier AI agents to their limits. Think... ...yourself. How to get started Apply to this post and get the... ...training prompts to refining model responses, you’ll be directly... ...understanding of how scoring or evaluation works in agent testing (precision...Permanent employmentPart timeFreelanceRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied Scientist, Agent Evaluation & Adaptive Model Routing. Be the first to apply!
- scientist ii Austin, TX
- regulatory scientist Austin, TX
- graduate scientist Austin, TX
- manufacturing scientist Austin, TX
- analytical scientist Austin, TX
- senior research scientist Austin, TX
- support scientist Austin, TX
- application scientist Austin, TX
- r&d scientist Austin, TX
- drug safety scientist Austin, TX




