Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Enginer, Core Evaluations [Remote]

Full-time

Cantina

Remote
  • Remote job

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We are seeking an experienced Machine Learning Engineer (MLE) to focus on audio model evaluation, specifically for speech generation and recognition models.

This role involves designing and developing comprehensive model evaluation pipelines for both development and production environments, as well as creating automated dashboards for reporting evaluation results.

As the founding member of our evaluation team, the ideal candidate is expected to leverage their experience to lead our evaluation efforts and play a key role in the future growth of the evaluation team.

What You’ll Do:

  • Designing model evaluation pipelines for models in development and production

  • Designing user studies for subjective model evaluations.

  • Converting requirements into measurable metrics.

  • Designing and developing automated evaluation dashboard to see model performances and compare results.

  • Training new models to capture new and different evaluation metrics.

  • Communicating with the model team to help design better models based on the evaluation results.

  • Communicating with the data team to help decide the type of data necessary to improve model performance.

  • Communication with the product-manager to make sure product requirements are correctly measured.

  • Help grow the evaluation team as the founding member.

  • Lead the evaluation team in the future.

What You’ll Bring:

  • Strong experience and intuition for designing metrics that capture model performance.

  • Strong experience with designing user studies on Mechanical Turk or similar platforms. .

  • Strong experience with model training and fine-tuning for model evaluation.

  • Strong statistical knowledge and experience to statistically compare evaluation results and take decisions.

  • Very strong engineering and programming skills.

  • Experience with training ASR, TTS models.

  • Experience at ML teams working on large-scale machine learning problems. (>3B models with >1m hours of data)

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the Machine Learning Enginer, Core Evaluations [Remote] in Remote vacancy
  • $225.4k - $257.2k

    Senior Lead AI Engineer (AI Foundations, LLM Core and Agentic AI) Overview: At...  ...been an industry leader in using machine learning to create real-time, personalized...  ...similarity search, guardrails, model evaluation, experimentation, governance, and... 
    Suggested
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    Remote
    19 hours ago
  • $238k - $302k

     ...states. The Large Model Evaluation team is at the nexus of Waymo...  ...real-world driving. At its core, our progress is defined by...  ...looking for quantitatively-minded engineers to research and propose new...  ...large-scale systems. Machine learning & Quantitative Experience... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    19 hours ago
  • $213k - $263k

     ...create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How do you...  ...that a virtual world is "real"? We are seeking visionary machine learning engineers and researchers to architect the scalable deep learning systems... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    19 hours ago
  • $175k - $215k

     ...Waymo Driver. The Simulator Evaluation team faces the ultimate data...  ...are looking for a Software Engineer to build the metrics and pipelines...  ...Manager and serve as a core contributor to the simulator...  ...libraries (Pandas, NumPy), or machine learning. Background in Autonomous... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    19 hours ago
  • Role Description We are seeking an experienced Machine Learning Engineer (MLE) to focus on audio model evaluation, specifically for speech generation and recognition models. This role involves: ~Designing model evaluation pipelines for models in development and production... 
    Suggested
    Full time

    Cantina

    Remote
    6 days ago
  •  ...systems for public safety agencies, the full-time remote AI Evaluation Engineer will design and maintain automated evaluation pipelines,...  ...multiple states 3+ years of experience in software engineering, machine learning, data science, or a related technical field Experience... 
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  • $161.6k - $200k

     ...infrastructure we need for the care we deserve. To learn more, visit Hybrid 3 days (...  ...) Position Summary As an AI Evaluation Engineer at Judi Health, you will build the...  ...or Master's degree in Computer Science, Machine Learning, or a related quantitative field... 
    Local area
    Flexible hours

    Capital Rx

    Denver, CO
    19 hours ago
  • $238k - $302k

     ...Waymo AI Foundations team is to develop machine learning solutions addressing open problems in...  ..., hierarchical learning, and robust evaluation. This role follows a hybrid work schedule...  ...report to a Senior Staff Software Engineer.   You will: Work with a... 
    Full time
    Remote work

    Waymo

    Remote
    19 hours ago
  •  ...AI lab researchers to create evaluations, publish benchmarks, and...  ...institutions Work together with engineers, scientists, operators, and...  ...AI Human data is the core infrastructure to AI advancement...  ...at Handshake: Machine Learning is at the heart of Handshake... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Handshake

    San Francisco, CA
    19 hours ago
  • $204k - $259k

     ...states. The Driver Understanding and Evaluation (DUE) team at Waymo is developing rich...  ...of the Waymo Driver.  The DUE Machine Learning team will build and operate scalable machine...  ...looking for researchers and software engineers who are passionate about developing... 
    Full time

    Waymo

    Remote
    19 hours ago
  • $204k - $259k

     ...other Waymo teams that consume map data. You will: Evaluate and launch machine learning models that automatically generate map from raw data....  ...parts of the system ~ A passion for good software engineering We prefer: M.S. or Ph.D in Computer Science or... 
    Full time
    Remote work

    Waymo

    Remote
    19 hours ago
  • $170k - $216k

     ...simulation across 15+ U.S. states. The DUE Machine Learning team will build and operate scalable...  ...tools, improve and speed up the evaluation and onboard developer journeys. It will...  ...looking for researchers and software engineers who are passionate about developing machine... 
    Full time

    Waymo

    Remote
    19 hours ago
  •  ...Role Overview: Coinflow is seeking a Machine Learning Engineer to help build the intelligence layer...  ...the company Define, own, and evolve core metrics across platform health,...  ...Design how models and data systems are evaluated, improved, and trusted over time Build... 
    Full time
    Worldwide

    Coinflow

    Remote
    19 hours ago
  • $281k - $356k

     ...simulation across 15+ U.S. states. The DUE Machine Learning team will build and operate scalable...  ...tools, improve and speed up the evaluation and onboard developer journeys. It will...  ...looking for researchers and software engineers who are passionate about developing machine... 
    Full time

    Waymo

    Remote
    19 hours ago
  • $170k - $216k

     ...Waymo Driver. We use modern machine learning techniques to model the complexities...  ...weather conditions. A core challenge is ensuring our...  ...components, we create metrics and evaluation methodologies to measure...  ...role you will report to an Engineering Manager.   You will:... 
    Full time
    Remote work

    Waymo

    Remote
    19 hours ago
  • $220k

     ...This role focuses on advancing the evaluation and development of cutting-edge coding agents...  ...intersection of AI research, software engineering, and model evaluation, designing the...  ...experience in software engineering, machine learning, AI research, evaluation, or related technical... 
    Full time
    Remote work

    SaidGig

    United States
    19 hours ago
  •  ...NTC OVERVIEW: We are seeking a Machine Learning Engineer to join our team. Working at NT Concepts...  ...synthetic data generation for training and evaluation of ML Models is a plus ~ Experience...  ...problem-solving. Employees are the core of NT Concepts. We understand that... 
    Full time
    Remote work

    Nt Concepts

    Herndon, VA
    19 hours ago
  •  ...About the role We are seeking a Machine Learning Engineer to strengthen our element classification...  ...technical bottlenecks. Implement evaluation, monitoring, and regression testing frameworks...  .... An opportunity to contribute to core ML systems that power our... 
    Full time
    Work at office
    Work from home
    Flexible hours

    Teleskope

    New York, NY
    19 hours ago
  •  ...Community.   We are seeking to hire a AI/Machine Learning Engineer to our team! Role Overview: As an...  ...Hybrid Model Development: Design and evaluate machine learning models that support...  ...multi-step agent workflows. Core Development: Strong proficiency in Python... 
    Remote job
    Full time
    Work experience placement
    Work at office

    Cybermedia Technologies

    Remote
    19 hours ago
  •  .... The Role This is a founding Machine Learning Engineer role for our conversational AI coaching...  ...will directly impact our product's core and shape the future of AI-driven leadership...  ...for performance and accuracy, and evaluate them to ensure they are production-ready... 
    Full time
    Work at office
    Remote work
    Work from home

    Valence

    New York, NY
    19 hours ago
  • $135k - $145k

     ...seeking a highly skilled and motivated Machine Learning / AI Engineer to support a federal government...  ...feature engineering, model training, evaluation, and deployment. Apply deep expertise...  ...in development. ~ Understanding of core concepts such as context windows, prompt... 
    Remote job
    Permanent employment
    Full time

    Element Solutions Inc.

    United States
    19 hours ago
  •  ...philosophy and how we use AI in our recruiting process here . We are looking for a Staff Machine Learning Engineer to lead the technical vision for our Ads Conversion Core Modeling team, building the state-of-the-art systems that power our global marketplace.... 
    Full time
    Work at office
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    19 hours ago
  • $45 - $50 per hour

     ...The role of the Senior Python Engineer is to design, develop, and deploy machine learning solutions that transform data into...  ...to support model accuracy Evaluate model performance and iterate to...  ...Strong proficiency in Python and core data science libraries such as... 
    Hourly pay
    Full time
    Work at office
    Remote work
    Work visa

    Formativgroup

    Remote
    19 hours ago
  • $197.6k - $420k

     ...in the world. Our name—short for "machine learning company"—reflects our core mission: democratizing access to the...  ...including YouTube's monetization engine and key search advertising...  ...experimentation, including A/B tests, offline evaluations, and counterfactual analyses.... 
    Full time
    Temporary work
    Immediate start
    Shift work

    Moloco

    Seattle, WA
    19 hours ago
  •  ...AI Engineer Role Overview: As an AI Engineer at Particle41...  ...design, develop and deploy machine-learning and deep-learning models/solutions...  ...to model training, evaluation, deployment and monitoring....  ...About Particle41 Our core values of  Empowering, Leadership... 
    Full time
    Local area

    Particle41

    Remote
    19 hours ago
  • $180k - $350k

     ...This is not a chatbot, prompt-engineering, or RAG-wrapper opportunity....  ...building production-grade machine learning infrastructure where prediction...  ..., build, and operate the core machine learning systems powering...  ..., deployment, monitoring, evaluation, and continuous improvement.... 
    Full time
    Relocation package
    Flexible hours

    Cb Smart Recruit

    Remote
    19 hours ago
  •  ...As a Senior Machine Learning Engineer on the Economy ML team, you will build models that power ranking...  ...engineering, model design, training, evaluation, and production deployment on Roblox’s...  ...cell experiments), tying your work to core metrics such as engagement, bookings,... 
    Full time

    Roblox

    Remote
    19 hours ago
  • $150k - $210k

     ...WHOOP is hiring a Senior AI/ML Engineer to help scale the...  ...In this role, you will own core components of the AI Platform...  ...our internal AI Studio : evaluation pipelines, fine-tuning workflows...  ...years of experience in applied machine learning, AI engineering, or ML-focused... 
    Full time
    Work at office
    Relocation

    Whoop

    Boston, MA
    19 hours ago
  •  ...About the Role We're looking for founding Machine Learning Engineers (MLEs) to own and improve our core action models end-to-end - the intelligence that powers...  ...optimizations between client and server Build evaluation frameworks and data pipelines to measure and... 
    Full time
    Sleeping nights

    Composite

    San Francisco, CA
    19 hours ago
  •  ...for a pragmatic, startup-minded Senior Machine Learning Engineer or Applied Data Scientist who can take...  ..., customer, and operational data Evaluate ambiguous ideas quickly and determine...  ...alongside Go and TypeScript across our core product and backend environments. Data... 
    Full time
    Flexible hours

    Gitkraken

    Remote
    19 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Enginer, Core Evaluations [Remote]. Be the first to apply!