Evgeniya Sukhodolskaya - Data Advocate, Toloka - Data at the core of all the cool ML

28/01/2023 1h 26min Temporada 2 Episodio 4

Listen "Evgeniya Sukhodolskaya - Data Advocate, Toloka - Data at the core of all the cool ML"

Descargar episodio Ver en sitio original

Episode Synopsis

Toloka’s support for Academia: grants and educator partnershipshttps://toloka.ai/collaboration-with-educators-formhttps://toloka.ai/research-grants-formThese are pages leading to them:https://toloka.ai/academy/education-partnershipshttps://toloka.ai/grantsTopics:00:00 Intro01:25 Jenny’s path from graduating in ML to a Data Advocate role07:50 What goes into the labeling process with Toloka11:27 How to prepare data for labeling and design tasks16:01 Jenny’s take on why Relevancy needs more data in addition to clicks in Search18:23 Dmitry plays the Devil’s Advocate for a moment22:41 Implicit signals vs user behavior and offline A/B testing26:54 Dmitry goes back to advocating for good search practices27:42 Flower search as a concrete example of labeling for relevancy39:12 NDCG, ERR as ranking quality metrics44:27 Cross-annotator agreement, perfect list for NDCG and Aggregations47:17 On measuring and ensuring the quality of annotators with honeypots54:48 Deep-dive into aggregations59:55 Bias in data, SERP, labeling and A/B tests1:16:10 Is unbiased data attainable?1:23:20 AnnouncementsThis episode on YouTube: https://youtu.be/Xsw9vPFqGf4Podcast design: Saurabh Rai: https://twitter.com/srvbhr

More episodes of the podcast Vector Podcast

Economical way of serving vector search workloads with Simon Eskildsen, CEO Turbopuffer 19/09/2025

Adding ML layer to Search: Hybrid Search Optimizer with Daniel Wrigley and Eric Pugh 21/03/2025

Vector Databases: The Rise, Fall and Future - by NotebookLM 02/03/2025

Code search, Copilot, LLM prompting with empathy and Artifacts with John Berryman 10/02/2025

Debunking myths of vector search and LLMs with Leo Boytsov 17/01/2025

Berlin Buzzwords 2024 - Alessandro Benedetti - LLMs in Solr 07/11/2024

Berlin Buzzwords 2024 - Sonam Pankaj - EmbedAnything 19/09/2024

Berlin Buzzwords 2024 - Doug Turnbull - Learning in Public 18/07/2024

Eric Pugh - Measuring Search Quality with Quepid 26/06/2024

Sid Probstein, part II - Bring AI to company data with SWIRL 15/05/2024

Ver todos los episodios

ZARZA We are Zarza, the prestigious firm behind major projects in information technology.

Evgeniya Sukhodolskaya - Data Advocate, Toloka - Data at the core of all the cool ML

Listen "Evgeniya Sukhodolskaya - Data Advocate, Toloka - Data at the core of all the cool ML"

Episode Synopsis

More episodes of the podcast Vector Podcast

Information Technology (IT)

Internet Predators on the prowl

Bandwidth: Broadband or Narrowband?

Personnel recruitment via Web

Deep web or Invisible Internet

Subdomains, a glance with the experts!

Free Internet, a prediction in Nostradamus style

Educational Technology: From traditional to digital

Localhost, there’s no place like 127.0.0.1

Googling with breathtaking tricks you ignore

Gray Hat Hacking, those with ambiguous ethics…

Internet Predators on the prowl

Dot COM: The Internet’s dominant TLD