Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist (Model Evaluation)

Sanas.AI Inc.

Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross-language communication, and more. Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language. Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$120M ARR. Our team combines deep expertise in model innovation and systems engineering with a design-minded product engineering culture to build and ship cutting-edge AI models and experiences — entirely in-house. Sanas is a 130 person team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies. With Sanas, you're not just adopting a product; you're investing in the future of communication. If you’re looking to have a significant role in roadmapping and driving technical directions, if you’re looking to deploy challenging and big ideas without much overhead or slowness, if you're looking to leave your mark on an ambitious, generational mission to change how the worlds thinks about speech + AI, then Sanas is a well-suited place for you. About the Role Progress in speech AI is only as meaningful as our ability to measure it. At Sanas, model quality spans dimensions that automated metrics struggle to capture — accent naturalness, perceptual clarity, speaker identity preservation, noise suppression without speech distortion, translation fluency under real-world disfluency. We're looking for a Research Scientist who can define what "better" actually means across all of Sanas's model families, build the evaluation infrastructure to measure it rigorously, and close the loop between research progress and real-world impact. This role sits at the intersection of research, product, and infrastructure — and directly shapes how every model team at Sanas measures progress. Job Description Design and own evaluation frameworks across Sanas's full model portfolio — Accent Translation, Noise Cancellation, Speech Enhancement, and Language Translation, and more — ensuring each captures meaningful progress, not just benchmark performance. Develop novel quantitative metrics for subjective and perceptual qualities: accent similarity, naturalness, speaker identity preservation, intelligibility under noise, and translation fluency in spoken-language domains. Build evaluation systems that bridge automated metrics and human judgment — designing listening studies, MOS/MUSHRA protocols, and preference tests that are statistically rigorous and operationally scalable. Define evaluation splits, test sets, and benchmark suites that accurately reflect production conditions — diverse accents, languages, noise environments, recording devices, and telephony codecs. Evaluation infrastructure & tooling Build and maintain automated evaluation pipelines that run continuously against model checkpoints — surfacing regressions early and tracking quality trends across training runs. Develop reference-based and reference-free metrics calibrated to Sanas's specific model tasks: SI-SDR, PESQ, STOI, DNSMOS, speaker similarity, WER delta, COMET, and task-specific custom metrics where off-the-shelf measures fall short. Instrument model quality monitoring in production — detecting degradation across language pairs, accent profiles, and acoustic conditions in live customer traffic. Build tooling that allows research scientists and ML engineers to run rigorous ablations, compare model versions, and understand quality tradeoffs without needing to design the evaluation from scratch each time. Design and operate human evaluation programs — listener panels, crowdsourced annotation, and expert evaluator workflows — that produce reliable signal on dimensions automated metrics cannot capture. Conduct research into evaluation methodology itself: when do automated metrics correlate with human perception, when do they diverge, and what does that tell us about model behavior? Partner directly with research scientists across model teams to translate open-ended quality questions into concrete, measurable evaluation protocols. Cross-functional impact Work closely with ML research, product, and customer success teams to ensure evaluation reflects what customers actually experience — not just what lab conditions optimize for. Feed evaluation insights back into data acquisition and model training priorities — identifying which failure modes require more data, architectural changes, or training procedure improvements. Communicate evaluation results clearly to both technical and non-technical stakeholders, translating metric movements into product quality narratives that inform roadmap decisions. Qualifications 4+ years of research or applied research experience in speech, audio, or NLP, with a demonstrated focus on evaluation methodology and quality measurement. Deep familiarity with speech and audio quality metrics — perceptual (MOS, MUSHRA, PESQ, STOI), signal-level (SI-SDR, SNR), and task-specific (WER, speaker similarity, DNSMOS) — and an understanding of when each is and isn't the right tool. Experience designing and running human evaluation studies — listener panels, crowdsourced annotation, inter-annotator agreement analysis — with statistical rigor. Strong engineering skills: you can build production-quality evaluation pipelines, not just run scripts. Proficiency in Python and PyTorch or equivalent. Creativity in defining novel quantitative metrics for subjective or behavioral qualities — you've identified gaps in existing evaluation approaches and built something better. Ability to take open-ended research questions and translate them into concrete, measurable evaluation systems that run reliably at scale. Curiosity and rigor in equal measure — you're as motivated by discovering the right way to measure progress as by the progress itself. Bonus Experience evaluating models across multiple speech tasks — ASR, TTS, speech enhancement, speaker verification, or machine translation. Familiarity with real-time or streaming model evaluation — latency-quality tradeoffs, codec-degraded audio, telephony channel conditions. Background in psychoacoustics or perceptual audio quality — understanding of how humans perceive speech naturalness, noise, and distortion. Experience with multilingual evaluation — cross-lingual quality metrics, language-specific annotation challenges, low-resource language evaluation. Published research at INTERSPEECH, ICASSP, ACL, EMNLP, or equivalent venues on evaluation methodology, speech quality, or related topics. #J-18808-Ljbffr

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Research Scientist (Model Evaluation) in Palo Alto, CA vacancy
  •  ...Founded by a team of Stanford researchers and entrepreneurs with deep...  ...combines deep expertise in model innovation and systems engineering...  ...'re looking for a Research Scientist who can define what "better"...  ...s model families, build the evaluation infrastructure to measure it... 
    Suggested

    Sanas

    Palo Alto, CA
    2 days ago
  • Sanas in Palo Alto is seeking a Research Scientist focused on rigorous evaluation of speech AI models. You will define meaningful metrics, build scalable evaluation pipelines, and align research progress with product impact across Accent Translation, Noise Cancellation,... 
    Suggested

    Sanas

    Palo Alto, CA
    3 days ago
  • Sanas is a leading force in real-time speech AI, advancing evaluation-driven research across accent translation, noise cancellation, and language translation. We seek a Research Scientist to define meaningful progress metrics and build rigorous evaluation infrastructure... 
    Suggested

    Sanas.AI Inc.

    Palo Alto, CA
    4 days ago
  • $190k - $250k

     ...developing large-scale generative world models that learn to predict realistic,...  ...autonomous trucks. We are looking for a research scientist to lead the design and development of...  ...camera, LiDAR, and radar outputsDesign evaluation frameworks that measure world model quality... 
    Suggested
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    4 days ago
  • $185k - $400k

     ...infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-training large-scale multimodal foundation models to advance our mission of... 
    Suggested
    Remote work

    Pika

    Palo Alto, CA
    4 days ago
  • Google DeepMind seeks a Senior Product Manager embedded in Gemini research and model training. You will read evaluations, analyze model outputs, and make judgment calls on quality alongside researchers, translating user needs into product priorities and feedback loops... 

    Google DeepMind

    Mountain View, CA
    1 day ago
  •  ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using...  ...world-class researchers, data scientists, and engineers, tackling the most fundamental...  ...-training and post-training, and evaluation benchmarks. The role combines... 

    Institute of Foundation Models

    Sunnyvale, CA
    11 days ago
  •  ...Luma's industry-leading generative video models into world models: interactive,...  ...metrics that define success. It fits a researcher with deep generative-modeling or model-based...  ...downstream embodied tasks (planning, control, evaluation). Enthusiasm for open-sourcing frontier... 

    Luma

    Redwood City, CA
    4 days ago
  • Tencent’s Technology Engineering Group (TEG) seeks a research-focused engineer to advance large-scale video world models, including data set design, model pre-training, SFT, RL, and downstream applications. You will analyze R&D challenges, optimize training and inference... 

    Tencent

    Palo Alto, CA
    3 days ago
  • $195.2k - $262.2k

     ...and enterprises from data and model training through to production...  ...role Nebius Token Factory needs scientists who can turn frontier inference bottlenecks into research problems, publish credible work...  ...measurable production impact. Invent, evaluate, and productionize methods for... 
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    3 days ago
  • The Role We are looking for a Research Scientist to join the Multi-Embodiment Generalist Agent...  ...founding member. MEGA is building foundation models for general-purpose robots beyond not...  ...and training pipelines, designing evaluations and deploying policies on physical... 
    Full time
    Work from home

    Wayve

    Sunnyvale, CA
    1 day ago
  • $199.8k - $300.2k

    Research and advance red teaming methods for LLM's and diffusion models Research and develop mitigations and safeguards to ensure safe deployment of LLM's in Apple...  ...tools, metrics, and datasets for assessing and evaluating the safety of LLM's over the model deployment... 
    Relocation

    Apple Inc.

    Cupertino, CA
    2 days ago
  • A leading autonomous vehicle technology company in Mountain View is seeking a World Model Research Scientist. The successful candidate will design and train generative models, requiring expertise in AI and robotics. Responsibilities include developing techniques for realistic... 

    Kodiak

    Mountain View, CA
    5 days ago
  • $272k - $431.25k

     ...future, we’re generating it! Our world model team is pushing the boundaries of...  ...AI. We are looking for a Senior Research Manager to lead world-model evaluation and benchmarking across NVIDIA’s...  ...be doing:Lead a team of Research Scientists focused on world-model evaluation,... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $158.3k - $297k

     ...innovation.What the Role Entails1. Engage in the research and development of large-scale video world models, including the design and construction of training...  ...related to pre-training, SFT, and RL, model capability evaluation, and exploration of downstream application... 
    Full time
    Relocation package

    Tencent

    Palo Alto, CA
    3 hours ago
  • $192.2k - $260k

     ...at Amazon's Delivery Foundation Model team, where you'll work alongside world-class scientists and engineers to pioneer the...  ...technical direction for specific research initiatives, ensuring robust performance...  ...and our extensive training and evaluation infrastructure- Guide and... 
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    2 days ago
  • $165k - $185k

    Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale, California, Pittsburgh, Pennsylvania...  ..., our AI research in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI (XAI), Natural Language... 
    Work experience placement
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    3 days ago
  • NVIDIA Gruppe is seeking a Senior Research Manager to lead world-model evaluation and benchmarking efforts in Santa Clara, California. The successful candidate...  ...models for Physical AI and build a team of research scientists focused on innovative evaluation techniques. The... 

    NVIDIA Gruppe

    Santa Clara, CA
    2 days ago
  • Lightspeed Studios is seeking candidates for a role focused on the research and development of large-scale video world models. This position involves designing datasets, foundational model algorithms, and evaluating model capabilities. Ideal candidates should have a Bachelor'... 

    Lightspeed Studios

    Palo Alto, CA
    1 day ago
  • ServiceNow in Santa Clara, CA is seeking a leader for the Build Agent evaluation framework. You will own eval strategy, roadmap, telemetry, and...  ...deliver scalable, data‑driven improvements. You will benchmark models, decide on model support, and communicate results to... 

    ServiceNow

    Santa Clara, CA
    1 day ago
  •  ...based in Sunnyvale, CA, is seeking a Research Scientist to join the MEGA team as a founding member...  ...role focuses on building foundation models for general-purpose robots beyond self...  ...architectures, data pipelines, and evaluations, and deploy policies on physical robots... 

    Wayve

    Sunnyvale, CA
    1 day ago
  • $126k - $423k

     ...About the role and team We are looking for multiple passionate Research Scientists to join the Research Group at Applied Intuition. The mission...  ...: Conduct research on pretraining world-action foundation model with various world modalities including vision and physics associated... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Immediate start
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    3 days ago
  • What the role actually is Nuro is hiring an applied AI researcher focused on agent systems and evaluation. The work sits around the research and measurement...  ...approaches for autonomous driving agent behavior Work with model, autonomy, and product teams to turn research... 

    The Robot Age

    Mountain View, CA
    1 day ago
  • Nuro is hiring an applied AI researcher focused on agent systems and evaluation. The work sits around the research and measurement loop for autonomous driving systems, with emphasis on how agent behavior is tested, compared, and improved. For an impact-focused role, the... 

    The Robot Age

    Mountain View, CA
    1 day ago
  •  ...Matter Expert (SME) to support cutting-edge AI research initiatives. You will collaborate with...  ...engineering teams to provide scientific expertise, evaluate AI research outputs, and contribute to improving advanced research models. Based in or near Silicon Valley, you will... 

    Aceolution

    Mountain View, CA
    1 day ago
  • $207k - $300k

     ...large-scale data pipelines to detect model misbehavior and misuse end-to-end.Research and develop cross-context...  ...with infrastructure teams and data scientists to scale your work and regularly...  ...working on data quality, automated evaluation design and simple statistical modeling... 

    Google

    Mountain View, CA
    4 days ago
  • $147k - $211k

    SnapshotWe are seeking strong Research Scientists with expertise in AI research and experience in interdisciplinary sociotechnical modeling to join a multimodal safety research effort...  ...world social context and dynamically evaluate and evolve system behaviors over long... 
    Full time

    DeepMind

    Mountain View, CA
    3 hours ago
  •  ...At Google DeepMind, we’re a team of scientists, engineers, machine learning...  ...fairness behavior of GDM’s latest Gemini models. The role of the Research Scientist / Research Engineer will...  ...risks.Design and maintain high quality evaluation protocols to assess model behavior... 

    DeepMind

    Mountain View, CA
    4 days ago
  • $174k - $252k

     ...and pursue a long-term applied research agenda to overcome...  ...improvement.Develop high-fidelity evaluation frameworks that measure agent...  ...benchmarks for LLMs, generative code models or agentic systems.Experience...  ...types of work. As a Research Scientist, you'll setup large-scale... 

    Google

    Mountain View, CA
    3 days ago
  • $112.88k - $149.57k

     ...seeking multiple Applied and Computational Scientists to conduct research and development in the areas of mathematics of generative AI models, reinforcement learning, and algorithm...  ...research ideas, lead the experimental evaluation of these ideas, present work to government... 
    Permanent employment

    SRI International

    Menlo Park, CA
    3 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist (Model Evaluation). Be the first to apply!