Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Scientist, Agent Evaluation & Adaptive Model Routing

Full-time

Bitdeer Technologies Group

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About Bitdeer AI Lab:

Bitdeer AI Lab is a frontier AI lab under Bitdeer, a global-leading computing power solutions provider. Guided by long-termism, we are committed to exploring the frontiers of artificial intelligence with the ambition, courage, and determination to build technologies that can truly change the world.

Our mission is to turn energy into intelligence that people can actually afford to use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving decide what gets built at all. We work on this from the ground up, from the power and datacenters we own to the software that turns them into tokens — and we continue to invest in and expand the infrastructure behind it

What you will be responsible for:

  • This role builds the evaluation and decision systems that make agentic inference measurable, reliable, and adaptive. You will own and extend our LLM and agent evaluation pipeline, develop representative task suites, and measure behavior at the trajectory level across task success, tool use, quality, cost, latency, and token consumption. You will research and prototype adaptive model-routing strategies inside a shared agent harness, including rule-based baselines, cascades, stage-aware routing, uncertainty-aware selection, and escalation and recovery policies. You will take promising methods from offline evaluation through traffic replay, shadow testing, and internal pilots, while working closely with the MaaS and platform engineering teams on production integration. The role owns evaluation methodology, routing policy, and research prototypes; production gateway reliability, billing, access control, and SLA engineering are collaborative responsibilities with the MaaS and platform teams

How you will stand out:

  • Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Statistics, Electrical Engineering, or a related field, with substantial hands-on experience in LLM evaluation, agentic systems, applied machine learning, or adaptive inference
  • Strong programming ability in Python and practical experience with PyTorch and modern data and evaluation tooling; able to independently build reliable experimental pipelines and internal research prototypes
  • Demonstrated experience designing or operating evaluation pipelines for LLMs or agents, including task-level and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis
  • In addition to hands-on LLM or agent evaluation experience, candidates should have implementation-level depth in at least one core area: model selection and routing, uncertainty estimation and calibration, cascading and escalation, or stage-aware agent inference. Experience with preference modeling, contextual bandits, online learning, or broader adaptive inference methods is a plus.
  • Hands-on experience evaluating multi-turn or tool-using agents, including task completion, tool-call correctness, planning failures, recovery behavior, and cost and latency trade-offs
  • Rigorous experimental practice — controlled comparisons, honest baselines, statistical analysis, and the ability to explain clearly what an evaluation result does and does not prove
  • Ability to analyze per-request and per-trajectory quality, cost, latency, and token usage, and construct cost-quality Pareto frontiers rather than relying only on aggregate model scores
  • Experience with multi-model APIs, agent harnesses, traffic replay, shadow evaluation, A/B testing, or production model monitoring is highly preferred
  • Familiarity with model-specific differences in tool calling, context windows, prompt caching, reasoning controls, and inference systems such as vLLM or SGLang is a plus
  • Publications at top-tier ML, NLP, or systems venues, or substantial open-source contributions in evaluation, agents, routing, or inference, are welcome
  • Strong ownership and product judgment, with a track record of taking ambiguous research questions from problem definition through a working 0-to-1 prototype and measurable internal validation

What you will experience working with us:

  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

Vacancy posted 21 days ago
Similar jobs that could be interesting for youBased on the Applied Scientist, Agent Evaluation & Adaptive Model Routing in Austin, TX vacancy
  •  ...where that happens: every product built on a model is bounded by what it costs to run, so...  ...are all in scope. You will implement and adapt published methods on our models and...  ...they fall short. You will also build the evaluation discipline that makes a claim like “lossless... 
    Suggested
    Full time
    Internship

    Bitdeer Technologies Group

    Austin, TX
    a month ago
  • $142.8k - $193.2k

     ...and we're looking for exceptional scientists to help lead the way.The Last Mile Routing & Planning organization develops...  ...Learning methods- Proven experience applying these methods to large-scale,...  ...business problems- Ability to translate models into production-ready code in... 
    Suggested
    Flexible hours

    Amazon

    Austin, TX
    20 hours ago
  • $183.8k - $248.7k

    Twitch is looking for an Applied Scientist to lead computer vision and video manipulation work within...  ...(APM). You will apply large language models (LLMs), vision-language models (VLMs),...  ..., and you will have the autonomy to evaluate both first-party and third-party solutions... 
    Suggested
    Flexible hours
    Day shift
    Afternoon shift

    Amazon

    Austin, TX
    1 day ago
  • $114.6k - $234.6k

     ...Job Description The OCI AI Evaluation team builds the evidence behind model-selection, product-readiness, and...  ...retrieval-augmented generation, AI agents, NL2SQL, multimodal...  ...responsible AI. As a Senior Applied Scientist on the team, you will independently... 
    Suggested
    Temporary work
    Flexible hours

    Oracle

    Austin, TX
    2 days ago
  •  ...Applied Scientist Services at Apple help hundreds of millions of customers get the most out...  ...observational testing frameworks, counterfactual modeling, and lifetime value estimation. As a...  ...causal inference and AIML research, evaluating and integrating new frameworks where... 
    Suggested

    Apple

    Austin, TX
    3 days ago
  •  ...Manager to own AI inference and model serving for k0rdent AI, our...  ...placement, autoscaling, routing, lifecycle management, observability...  ...performance bottlenecks, evaluating system design trade-offs, and...  ...employment decision for the position applied for. You also have the right... 

    Mirantis

    Austin, TX
    3 days ago
  • We are hiring an Applied Scientist with experience in performing hyperparameter optimizations, evaluating model performance, and deploying production natural language systems. Working with an elite group of professions, you will have the opportunity to work hands on with... 

    Protopia AI, Inc.

    Austin, TX
    3 days ago
  • $272k - $431.25k

     ...new generation of interactive world-model systems. With this release, NVIDIA has...  ...need to see:Demonstrated research or applied innovation in world models, video diffusion...  ....Hands-on experience building, adapting, or deeply evaluating world models, video generation systems... 
    Full time

    Nvidia

    Austin, TX
    20 hours ago
  • $159.8k - $244.3k

     ...Artificial Intelligence Scientist to lead the...  ...track record of taking models from problem...  ...generative AI, and multi-agent solutions using complex...  ...workflow engines, evaluation frameworks, guardrails, model routing, and human-in-the-...  ...data movement.Apply strong software engineering... 
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Austin, TX
    4 days ago
  •  ...standard business hours with a hybrid work model (4 days in-office, 1 day working from...  ...external consultants on validation activities. Evaluate model performance monitoring and complete...  ...engineering. Experience developing and applying machine learning models. 5+ years of... 
    Full time
    Work at office
    Work from home
    Relocation

    The Charles Schwab Corporation

    Austin, TX
    4 days ago
  • $99.6k - $174k

     ...deployment of AI agents. We also drive the...  ...quality telemetry).Apply secure SDLC and privacy...  ...practices (threat modeling, least privilege)....  ...(RAG, retrieval, routing, tool-use, evals)...  ...including safety evaluations, iterative testing...  ...-tuning and model adaptation.Familiarity with... 
    Full time
    Work at office
    Remote work
    2 days per week

    Wolters Kluwer

    Austin, TX
    4 days ago
  •  ...needs for faculty specialized in computer science, statistics, applied mathematics, and artificial intelligence . The ideal candidate...  ...an environment of open inquiry in both teaching and research, modeling academic freedom and respectful debate. Qualifications - Ph.D.... 
    Full time
    Summer work

    UATX

    Austin, TX
    3 hours ago
  • $204k - $216k

     ...research role focused on the models at the foundation of...  .... You study, adapt, and advance the foundational...  ...experiments, evaluating models, adapting them...  ...Foundational Model Research Data Scientist does that. You run the...  ...models or applied NLP. Experience with... 

    Sapience AI Corporation

    Austin, TX
    5 days ago
  •  ...guidance on program planning, execution, and evaluation for the compliance, accounts management...  ...been formulated. Interprets, develops, adapts, and coordinates analytical reviews of...  ...impact on operations and resources. Applies and disseminates newly developed or revised... 
    Remote work

    Treasury Department

    Austin, TX
    1 day ago
  • Environmental Scientist, Transmission Line Routing & Siting Halff has an opening for an Environmental professional, specializing in Transmission Line Routing & Siting, to join our growing Environmental team as an Environmental Scientist. The ideal candidate will have at... 
    Temporary work
    Local area
    Flexible hours

    Halff Associates, Inc.

    Austin, TX
    1 day ago
  • $112k

     ...everywhere.Machine Learning Scientists IIWithin the AI & Data...  ...and operationalizes applied science driven...  ...the Marketplace Health Model system to identify behaviors...  ...impact.Proactively evaluate opportunities based on...  ...applications or LLM-powered agents.Experience with trust... 
    Full time

    Expedia

    Austin, TX
    2 days ago
  • $115.8k - $144.7k

    POSITION SUMMARY: The Senior Design Transfer Scientist will lead and support the transition of...  ...: Lead change control activities to evaluate and determine the impact of design changes...  ...qualified applicants are encouraged to apply, and will be considered without regard... 
    Work at office
    Immediate start
    Worldwide

    Natera

    Austin, TX
    20 hours ago
  • $60 per hour

     ...attention to detail. Candidates will review AI evaluation tasks, identify inconsistencies, and define expected behaviors for agents. Ideal applicants have experience in policy...  ...and project needs. A great opportunity to influence future AI models! #J-18808-Ljbffr Mind Rift
    Remote job
    Freelance

    Mind Rift

    Austin, TX
    4 days ago
  •  ...Right Of Way Of Agents And Sr. Right Of Way Agents Coates Field Service, Inc. is seeking...  ...way for private landowners, and able to adapt to tight deadlines to meet project...  ...maps, electronic and paper. Ability to evaluate, interpret, and analyze engineering and right... 
    Daily paid
    Temporary work
    Work at office

    Coates Field Service

    Austin, TX
    1 day ago
  • $97.3k - $125.54k

     ...civilian hiring freeze. As an FBI special agent, you'll directly impact national...  ...brings new challenges that demand your adaptability and resilience, but you're not alone in...  ...that prioritizes you. Set yourself apart. Apply today. Salary Level $97,300.00–$125,544.... 
    Work at office
    Local area

    Federal Bureau of Investigation (FBI)

    Austin, TX
    20 hours ago
  •  ...optical metrology applications in semiconductor manufacturing; apply robust statistical techniques to optimize applications for the...  ...environment: and provide region support for new product introduction and evaluation activities (beta, head-to-head). Why Nova: Certified Best... 

    NOVA

    Austin, TX
    1 day ago
  • $30 per hour

     ...We are seeking a highly motivated AI Agent Intern to join Oracle's Supply Chain Applications...  ...demo scripts. Research & Analysis Evaluate logistics AI use cases (forecasting,...  ...Manufacturing. Engage with Oracle AI Champions, Applied Technologists, and Product Strategy.... 
    Hourly pay
    Temporary work
    Internship
    Flexible hours

    Oracle

    Austin, TX
    4 days ago
  • $10k

     ...Border Patrol Agent (BPA) - Experienced (GL-9 GS-11) - New Hire Sign-On and Retention...  ...the next higher grade level (without re-applying) once you successfully complete 52 weeks...  ...transcripts, etc.) to submit. You will be evaluated based on your resume, supporting... 
    Full time
    Local area
    Immediate start
    Relocation
    Night shift

    Customs and Border Protection

    Austin, TX
    13 hours ago
  • $10k

     ...Border Patrol Agent (BPA) - Experienced (GL-9 GS-11) - New Hire Sign-On and Retention...  ...the next higher grade level (without re-applying) once you successfully complete 52 weeks...  ...transcripts, etc.) to submit. You will be evaluated based on your resume, supporting... 
    Full time
    Local area
    Immediate start
    Relocation
    Night shift

    Customs and Border Protection

    Austin, TX
    13 hours ago
  • $10k

     ...Border Patrol Agent (BPA) – in the Federal Security and Public Safety Sector (Entry Level...  ...the completion of required training and apply these skills in a law enforcement capacity...  ..., etc.) to submit. You will be evaluated based on your resume, supporting documents... 
    Full time
    Work experience placement
    Immediate start
    Relocation
    Night shift

    US Customs and Border Protection

    Pflugerville, TX
    3 days ago
  •  ...reinforcement learning environments that evaluate AI models on complex software engineering tasks...  ...golden reference solutions. Evaluate AI agents' ability to reason through complex...  ...Minimum weekly submission requirements may apply. Availability Selected experts... 
    Remote job
    For contractors

    YO AI Labs

    Austin, TX
    27 days ago
  •  ...discretion, vigilance, and precision, our agents ensure a secure and seamless living...  ...nuances of residential environments and can adapt to dynamic family routines and estate operations...  ..., and discretion, we invite you to apply and become part of our elite Residential... 
    For contractors

    Crisis24

    Austin, TX
    2 days ago
  •  ...The Role As an AI Agent Engineer, you will design...  ...consistently. Apply enterprise platform guardrails...  ...regulated data. Evaluate incoming integration...  ...interfaces, manually routed workflows, platform-specific...  ...for large language models used within enterprise... 
    Full time
    Local area
    Work from home
    Relocation package

    General Motors

    Austin, TX
    4 days ago
  • $29.89 - $37.37 per hour

     ...Environmental Scientist Under limited supervision and using comprehensive knowledge, conduct...  ...prepare reports and memos based on data evaluation Review and interpret policies, codes,...  ...initial training period. Exceptions may apply subject to the business needs of the... 
    Remote work
    Monday to Friday

    City of Austin, TX

    Austin, TX
    3 days ago
  • $73 per hour

     ...realistic tasks that push frontier AI agents to their limits. Think...  ...yourself. How to get started Apply to this post and get the...  ...training prompts to refining model responses, you’ll be directly...  ...understanding of how scoring or evaluation works in agent testing (precision... 
    Permanent employment
    Part time
    Freelance
    Remote work

    Mind Rift

    Austin, TX
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Scientist, Agent Evaluation & Adaptive Model Routing. Be the first to apply!