{"id":1211508,"date":"2023-12-13T13:11:07","date_gmt":"2023-12-13T18:11:07","guid":{"rendered":"https:\/\/www.prime-wow.com\/?p=1211508"},"modified":"2023-12-13T13:11:07","modified_gmt":"2023-12-13T18:11:07","slug":"turing-test-on-steroids-chatbot-arena-crowdsources-ratings-for-45-ai-models","status":"publish","type":"post","link":"https:\/\/www.prime-wow.com\/?p=1211508","title":{"rendered":"Turing test on steroids: Chatbot Arena crowdsources ratings for 45 AI models"},"content":{"rendered":"<div id=\"rss-wrap\">\n<figure class=\"intro-image intro-left\">\n  <img decoding=\"async\" src=\"https:\/\/www.prime-wow.com\/wp-content\/uploads\/2023\/12\/GettyImages-152404829-800x640-1.jpg\" alt=\"A Rock'em Sock'em AI model battle.\" \/><\/p>\n<p class=\"caption\" style=\"font-size:0.8em\"><a href=\"https:\/\/cdn.arstechnica.net\/wp-content\/uploads\/2023\/12\/GettyImages-152404829-scaled.jpg\" class=\"enlarge-link\" data-height=\"2048\" data-width=\"2560\">Enlarge<\/a> <span class=\"sep\">\/<\/span> A Rock&#8217;em Sock&#8217;em AI model battle. (credit: <a rel=\"nofollow\" class=\"caption-link\" href=\"https:\/\/www.gettyimages.com\/detail\/illustration\/boxing-robots-royalty-free-illustration\/152404829\">CSA Images<\/a>)<\/p>\n<\/figure>\n<div><a name=\"page-1\" \/><\/div>\n<p>As the AI landscape has expanded to include <a href=\"https:\/\/arstechnica.com\/science\/2023\/07\/a-jargon-free-explanation-of-how-ai-large-language-models-work\/\">dozens of distinct large language models (LLMs)<\/a>, debates over which model provides the &#8220;best&#8221; answers for any given prompt have also proliferated (Ars has even <a href=\"https:\/\/arstechnica.com\/information-technology\/2023\/04\/clash-of-the-ai-titans-chatgpt-vs-bard-in-a-showdown-of-wits-and-wisdom\/\">delved into<\/a> these <a href=\"https:\/\/arstechnica.com\/ai\/2023\/12\/chatgpt-vs-google-bard-round-2-how-does-the-new-gemini-model-fare\/\">kinds of debates<\/a> a few times in recent months). For those looking for a more rigorous way of comparing various models, the folks over at the Large Model Systems Organization (LMSys) have <a href=\"https:\/\/chat.lmsys.org\/?arena\">set up Chatbot Arena<\/a>, a platform for generating Elo-style rankings for LLMs based on a crowdsourced blind-testing website.<\/p>\n<p>Chatbot Arena users can enter any prompt they can think of into the site&#8217;s form to see side-by-side responses from two randomly selected models. The identity of each model is initially hidden, and results are voided if the model reveals its identity in the response itself.<\/p>\n<p>The user then gets to pick which model provided what they judge to be the &#8220;better&#8221; result, with additional options for a &#8220;tie&#8221; or &#8220;both are bad.&#8221; Only after providing a pairwise ranking does the user get to see which models they were judging, though a separate &#8220;side-by-side&#8221; section of the site lets users pick two specific models to compare (without the ability to contribute a vote on the result).<\/p>\n<\/div>\n<p><a href=\"https:\/\/arstechnica.com\/?p=1990779#p3\">Read 10 remaining paragraphs<\/a> | <a href=\"https:\/\/arstechnica.com\/?p=1990779&amp;comments=1\">Comments<\/a><\/p>\n<p>&#013;<br \/>\n&#013;<br \/>\nSource: Ars Technica &#8211; <a href=\"https:\/\/arstechnica.com\/?p=1990779\" target=\"_blank\" rel=\"noopener\">Turing test on steroids: Chatbot Arena crowdsources ratings for 45 AI models<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Enlarge \/ A Rock&#8217;em Sock&#8217;em AI model battle. (credit: CSA Images) As the AI landscape has expanded to include dozens of distinct large language models (LLMs), debates over which model provides the &#8220;best&#8221; answers for any given prompt have also &hellip; <a href=\"https:\/\/www.prime-wow.com\/?p=1211508\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":1211509,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"ngg_post_thumbnail":0,"footnotes":""},"categories":[27,110],"tags":[73],"class_list":["post-1211508","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ars-technica","category-unfiltered-rss","tag-ars-technica"],"_links":{"self":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1211508","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1211508"}],"version-history":[{"count":0,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1211508\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/media\/1211509"}],"wp:attachment":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1211508"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1211508"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1211508"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}