{"id":1160215,"date":"2023-07-13T09:00:12","date_gmt":"2023-07-13T13:00:12","guid":{"rendered":"https:\/\/www.prime-wow.com\/?p=1160215"},"modified":"2023-07-13T09:00:12","modified_gmt":"2023-07-13T13:00:12","slug":"metas-newest-dataset-will-train-speech-recognition-engines-on-clusters-of-speakers","status":"publish","type":"post","link":"https:\/\/www.prime-wow.com\/?p=1160215","title":{"rendered":"Meta&#039;s newest dataset will train speech recognition engines on &#039;clusters&#039; of speakers"},"content":{"rendered":"<p>It is 2023 and, sorry, Siri somehow still didn\u2019t catch that. Despite the tsunami of advancements generative AI systems have enjoyed in recent months, the synthetic assistants on our mobile devices <a data-i13n=\"cpos:1;pos:1\" href=\"https:\/\/www.engadget.com\/2020-03-24-racial-bias-speech-recognition-siri-alexa-ai.html\">remain nearly as hard of hearing<\/a> as they were in 2011. A newly developed dataset from Meta AI, however, promises to improve the performance of such automatic speech recognition (ASR) tools by clustering speech at the \u201cutterance level.\u201d<\/p>\n<p>Meta has long sought to improve its ASRs\u2019 performance, teaching them to <a data-i13n=\"cpos:2;pos:1\" href=\"https:\/\/www.engadget.com\/facebook-ai-speech-transcriptions-130005002.html\">train without the aid of transcripts<\/a>, recognize <a data-i13n=\"cpos:3;pos:1\" href=\"https:\/\/www.engadget.com\/metas-open-source-speech-ai-recognizes-over-4000-spoken-languages-161508200.html\">more than 4,000 spoken languages<\/a> and <a data-i13n=\"cpos:4;pos:1\" href=\"https:\/\/www.engadget.com\/ai-is-already-better-at-lip-reading-that-we-are-183016968.html\">even read lips at a higher proficiency than human experts<\/a>. However, many of the datasets used to train ASR models are organized by demographic \u2014 age group, gender, nationality, English accent \u2014 which limit the variation of pronunciations that models are trained on,  ultimately hindering their function in understanding a broad cross section of users.<\/p>\n<p><span id=\"end-legacy-contents\" \/><\/p>\n<p>To get around this, Meta AI has developed a dataset that instead relies on an utterance clustering method. \u201cInstead of dividing a dataset based on speakers\u2019 demographic information \u2026 our proposed algorithm clusters speech at the utterance level,\u201d the Meta AI team explained in Wednesday\u2019s blog post. \u201cA single cluster will contain similar utterances from a diverse group of speakers. We can then train our model using the various clusters and use fairness datasets to measure how the model impacts outcomes across different demographic groups.\u201d<\/p>\n<p>Meta\u2019s resulting dataset includes just over 27,000 command utterances collected from 595 paid US volunteers. Their utterances revolve around seven main themes \u2014 music, capture, utilities, notification control, messaging, calling and dictation \u2014 that other researchers can then use to train their own models and digital assistants on. Prompts included asking the speakers how they\u2019d voice search for a song or make plans with friends and deciding where to meet up.<\/p>\n<p>To evaluate this new system, Meta first trained a model on publicly-available, English-language Facebook videos. Researchers then evaluated that model using two other datasets: Casual Conversations v1, which Meta released in 2021, and a \u201cde-identified dataset collected from a data supplier for ASR,\u201d which includes 48,000 spoken utterances from 867 individuals.<\/p>\n<p>The initial results proved promising, with model performance improvements \u201con all demographic groups in our evaluation datasets, though by far the largest gains are with respect to more inclusivity of accents,\u201d per the blog. Overall, ASR performance increased by 10 percent using the clustering method, with large gains coming from the age 66-85 crowd as well, a traditionally underrepresented demographic in the voice command space.<\/p>\n<p>\u201cOur proposed algorithm is part of Meta\u2019s long-term focus on responsible AI and just one part of our holistic approach to address fairness issues,\u201d the researchers wrote. Looking ahead, the team is exploring adapting the system to other languages.<\/p>\n<p>This article originally appeared on Engadget at https:\/\/www.engadget.com\/meta-new-dataset-train-speech-recognition-engine-clusters-speaker-130012841.html?src=rss&#013;<br \/>\n&#013;<br \/>\nSource: Engadget &#8211; <a href=\"https:\/\/www.engadget.com\/meta-new-dataset-train-speech-recognition-engine-clusters-speaker-130012841.html?src=rss\" target=\"_blank\" rel=\"noopener\">Meta&#8217;s newest dataset will train speech recognition engines on &#8216;clusters&#8217; of speakers<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>It is 2023 and, sorry, Siri somehow still didn\u2019t catch that. Despite the tsunami of advancements generative AI systems have enjoyed in recent months, the synthetic assistants on our mobile devices remain nearly as hard of hearing as they were &hellip; <a href=\"https:\/\/www.prime-wow.com\/?p=1160215\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"ngg_post_thumbnail":0,"footnotes":""},"categories":[99,110],"tags":[98],"class_list":["post-1160215","post","type-post","status-publish","format-standard","hentry","category-engadget","category-unfiltered-rss","tag-engadget"],"_links":{"self":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1160215","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1160215"}],"version-history":[{"count":0,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1160215\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1160215"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1160215"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1160215"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}