Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist: Model Evaluation & Quality Metrics

Sanas

Sanas in Palo Alto is seeking a Research Scientist focused on rigorous evaluation of speech AI models. You will define meaningful metrics, build scalable evaluation pipelines, and align research progress with product impact across Accent Translation, Noise Cancellation, and Speech Enhancement. The role requires deep expertise in evaluation methodologies, strong programming in Python and PyTorch, and the ability to run human studies and ablations at scale, partnering with ML research, product, #J-18808-Ljbffr Sanas

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Research Scientist: Model Evaluation & Quality Metrics in Palo Alto, CA vacancy
  •  ...team of Stanford researchers and entrepreneurs...  ...deep expertise in model innovation and systems...  .... At Sanas, model quality spans dimensions that automated metrics struggle to...  ...looking for a Research Scientist who can define what...  ..., build the evaluation infrastructure to... 
    Quality

    Sanas

    Palo Alto, CA
    1 day ago
  • $184.7k - $324.8k

    Research Scientist / Engineer, Foundation Model Evaluation Cupertino, California, United States Software and Services We build...  ...evaluation benchmarks, metrics, and test suites that rigorously...  ...metrics predict user‑perceived quality and product outcomes. Experimental... 
    Quality
    Relocation

    Apple Inc.

    Cupertino, CA
    21 hours ago
  • $190k - $250k

     ...large-scale generative world models that learn to predict...  ...trucks. We are looking for a research scientist to lead the design and development...  ..., and radar outputsDesign evaluation frameworks that measure world model quality beyond pixel-level metrics, including scenario... 
    Quality
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    2 days ago
  • Apple Inc. is seeking a Research Scientist/Engineer to design evaluation systems for foundation models powering Apple products. You will work hands‑on across evaluation...  ...collaboration to drive model improvement and product quality. You will evaluate frontier capabilities,... 
    Quality

    Apple Inc.

    Cupertino, CA
    21 hours ago
  • $195.2k - $262.2k

     ...enterprises from data and model training through...  ...Factory needs scientists who can turn...  ...bottlenecks into research problems, publish...  ...publish or prepare high-quality technical work while...  ...impact. Invent, evaluate, and productionize...  ..., baselines, metrics, statistical reasoning... 
    Quality
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    2 days ago
  • $174k - $252k

     ...years of experience leading a research agenda. 1 year of...  ...field. Experience with AI model training, testing, evaluation, and tuning processes, as...  ...challenges and accelerate high-quality product innovation for billions...  ..., design, and evaluation metrics for research solution... 
    Quality

    Socket.dev

    Mountain View, CA
    4 days ago
  • $174k - $252k

    Senior Research Scientist, Gemini Release Evaluations, DeepMind DeepMind Mountain View, CA, USA;...  ...field. Experience with AI model training, testing,...  ...challenges and accelerate high-quality product innovation for...  ..., design, and evaluation metrics for research solution development... 
    Quality

    Google

    Mountain View, CA
    4 days ago
  •  ...development to solve global challenges and accelerate high-quality product innovation for billions of users. We...  ...priority. Responsibilities include authoring research papers, defining data structures and evaluation metrics, and driving project work to advance research... 
    Quality

    Socket.dev

    Mountain View, CA
    4 days ago
  • $272k - $431.25k

     ...generating it! Our world model team is pushing...  ...for a Senior Research Manager to lead world-model evaluation and benchmarking across...  ...team of Research Scientists focused on world-...  ...-loop benchmarks, metrics, failure taxonomy,...  ...coherence, SDG quality, and WAM usefulness... 
    Quality
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $185k - $400k

     ...intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-training large-...  ...fine-tuning.Identify, create, and leverage large, high-quality cross-modal datasets.Bring research advancements into... 
    Quality
    Remote work

    Pika

    Palo Alto, CA
    2 days ago
  • $224k - $356.5k

     ...Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a...  ...challenges and communicate effectively across research, engineering, and product teams.Ways to...  ....A strong appreciation for evaluation quality, including correctness, reproducibility... 
    Quality
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $192.2k - $260k

     ...s Delivery Foundation Model team, where you'll work...  ...alongside world-class scientists and engineers to...  ...direction for specific research initiatives, ensuring...  ...extensive training and evaluation infrastructure- Guide...  ...to improve the safety, quality, and efficiency of Amazon... 
    Quality
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    3 days ago
  • $165k - $185k

    Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale...  ...in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable...  ..., we shape the future by inventing high-quality technologies and services that spark... 
    Quality
    Work experience placement
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    17 hours ago
  • $150k

     ...You will join the Grok Voice Model team to help build the world’...  ...intensive post‑training to push quality, speed, and stability to the...  ...‑quality model training and evaluation. Work on pre‑training and...  ...framework covering objective metrics (accuracy, quality, latency,... 
    Quality
    Temporary work

    xAI

    Palo Alto, CA
    2 days ago
  • $175k - $350k

     ...with human-centered AI models that unite emotional...  ...the fun parts. Balance research curiosity with product...  ...-parameter search, evaluation, and rollout—using PyTorch...  ...to trace. Define the metrics that matter; run A/B...  ...quickly to meet aggressive quality targets. Collaborate... 
    Quality

    Inflection AI

    Palo Alto, CA
    4 days ago
  • $180k

     ...a multimodal engineer on the Imagine Model Team, you will develop cutting‑edge AI...  ...video experiences). Improve data quality through annotation, filtering, augmentation...  ...for visual and audio data. Design evaluation frameworks, metrics, benchmarks, evals, and reward models... 
    Quality
    Temporary work

    xAI

    Palo Alto, CA
    21 hours ago
  •  ...CA is seeking a leader for the Build Agent evaluation framework. You will own eval strategy, roadmap, telemetry, and cross-team quality standards, guiding an 8‑engineer team to...  ...data‑driven improvements. You will benchmark models, decide on model support, and communicate... 
    Quality

    ServiceNow

    Santa Clara, CA
    4 days ago
  •  ...industry-leading generative video models into world models: interactive,...  ...into a generated world, and own the metrics that define success. It fits a researcher with deep generative-modeling or...  ...tasks (planning, control, evaluation). Enthusiasm for open-sourcing frontier... 

    Luma

    Redwood City, CA
    2 days ago
  • $192k - $278k

    Lead model releases for Search, evaluating DeepMind release applicants against strict quality bars to determine launch readiness.Partner with cross-functional teams to resolve complex, ambiguous technical problems as Search evolves into a fully AI-enabled product.Design... 
    Quality
    Shift work

    Google

    Mountain View, CA
    2 days ago
  • Google DeepMind is seeking a senior research scientist to advance AI models, publish impactful papers, and...  ...contribute to frontier datasets and evaluation frameworks. Responsibilities include...  ...defining data structures and evaluation metrics for research solutions, driving... 

    Google DeepMind

    Mountain View, CA
    3 days ago
  •  ...Google DeepMind, we’re a team of scientists, engineers, machine learning experts...  ...behavior of GDM’s latest Gemini models. The role of the Research Scientist / Research Engineer will...  ...abuse risks.Design and maintain high quality evaluation protocols to assess model behavior... 
    Quality

    DeepMind

    Mountain View, CA
    2 days ago
  • $207k - $300k

     ...scale data pipelines to detect model misbehavior and misuse end-to-end.Research and develop cross-context...  ...infrastructure teams and data scientists to scale your work and regularly...  ...data pipelines, working on data quality, automated evaluation design and simple statistical... 
    Quality

    Google

    Mountain View, CA
    1 day ago
  • $174k - $252k

     ...pursue a long-term applied research agenda to overcome...  ...Develop high-fidelity evaluation frameworks that...  ...agent impact to software quality and developer productivity...  ...LLMs, generative code models or agentic systems....  ...of work. As a Research Scientist, you'll setup large-scale... 
    Quality

    Google

    Mountain View, CA
    17 hours ago
  • $174k - $252k

    Create comprehensive evaluation sets and benchmarks to...  ...audio-to-audio (A2A) model performance across international...  ...with cross-functional research and engineering teams...  ...work. As a Research Scientist, you'll setup large-...  ...and accelerate high-quality product innovation for... 
    Quality

    Google

    Mountain View, CA
    21 hours ago
  • $147k - $211k

    SnapshotWe are seeking strong Research Scientists with expertise in AI...  ...interdisciplinary sociotechnical modeling to join a multimodal safety...  ...context and dynamically evaluate and evolve system behaviors...  ...perspectives and harness these qualities to create extraordinary... 
    Quality
    Full time

    DeepMind

    Mountain View, CA
    3 days ago
  • $142.8k - $193.2k

     ...Relevance team works to maximize the quality and effectiveness of the...  ...machine-learned ranking models. The relevance improvements you...  ...limited to:* Analyze the data and metrics resulting from traffic into...  ...to improve search ranking.* Evaluate the proposed solutions via... 
    Quality
    Local area
    Worldwide
    Flexible hours

    Amazon

    Palo Alto, CA
    1 day ago
  •  ...hiring software engineers for the Model Behavior team to help shape...  ...strategies to deliver high-quality user experiences across...  ...with and release new models. Research & Analysis: Identify inconsistencies...  .... Experience designing evaluations or benchmarks for AI systems.... 
    Quality

    Kindredventures

    Palo Alto, CA
    3 days ago
  •  ...the training pipeline behind the models that power both Parallel’s search...  ...path from real product usage to high‑quality training data, fine‑tune and evaluate these models rigorously, and ship...  ...serve all three. You care about your research being applied to product and... 
    Quality
    Work at office
    Visa sponsorship

    Parallel Web Systems

    Palo Alto, CA
    1 day ago
  • $145k - $200k

     ...engineering team with expertise in enabling ML models in production. We deploy AI models to run...  ...ValueOwnership mindset and bias toward quality. Our software runs in environments where...  ...capabilities and the ability to quickly evaluate and integrate new models and technologies... 
    Quality
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Palo Alto, CA
    4 days ago
  •  ...ROLE:We are hiring an AI Research Scientist, Recursive Self...  ...sense: systems where models, data generators, or toolchains...  ...(e.g. synthetic data quality, targeted self-play,...  ...on counterfactual evaluation. You connect RSI concepts to concrete metrics—data efficiency, robustness... 
    Quality
    Shift work

    AMD

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist: Model Evaluation & Quality Metrics. Be the first to apply!