{"id":1137575,"date":"2023-05-09T20:02:00","date_gmt":"2023-05-10T00:02:00","guid":{"rendered":"https:\/\/www.prime-wow.com\/?p=1137575"},"modified":"2023-05-09T20:02:00","modified_gmt":"2023-05-10T00:02:00","slug":"openais-new-tool-attempts-to-explain-language-models-behaviors","status":"publish","type":"post","link":"https:\/\/www.prime-wow.com\/?p=1137575","title":{"rendered":"OpenAI&#039;s New Tool Attempts To Explain Language Models&#039; Behaviors"},"content":{"rendered":"<p>An anonymous reader quotes a report from TechCrunch: In an effort to peel back the layers of LLMs, OpenAI is developing a tool to automatically identify which parts of an LLM are responsible for which of its behaviors. The engineers behind it stress that it&#8217;s in the early stages, but the code to run it is available in open source on GitHub as of this morning. &#8220;We&#8217;re trying to [develop ways to] anticipate what the problems with an AI system will be,&#8221; William Saunders, the interpretability team manager at OpenAI, told TechCrunch in a phone interview. &#8220;We want to really be able to know that we can trust what the model is doing and the answer that it produces.&#8221;<\/p>\n<p>To that end, OpenAI&#8217;s tool uses a language model (ironically) to figure out the functions of the components of other, architecturally simpler LLMs &#8212; specifically OpenAI&#8217;s own GPT-2. How? First, a quick explainer on LLMs for background. Like the brain, they&#8217;re made up of &#8220;neurons,&#8221; which observe some specific pattern in text to influence what the overall model &#8220;says&#8221; next. For example, given a prompt about superheros (e.g. &#8220;Which superheros have the most useful superpowers?&#8221;), a &#8220;Marvel superhero neuron&#8221; might boost the probability the model names specific superheroes from Marvel movies. OpenAI&#8217;s tool exploits this setup to break models down into their individual pieces. First, the tool runs text sequences through the model being evaluated and waits for cases where a particular neuron &#8220;activates&#8221; frequently. Next, it &#8220;shows&#8221; GPT-4, OpenAI&#8217;s latest text-generating AI model, these highly active neurons and has GPT-4 generate an explanation. To determine how accurate the explanation is, the tool provides GPT-4 with text sequences and has it predict, or simulate, how the neuron would behave. In then compares the behavior of the simulated neuron with the behavior of the actual neuron.<\/p>\n<p>&#8220;Using this methodology, we can basically, for every single neuron, come up with some kind of preliminary natural language explanation for what it&#8217;s doing and also have a score for how how well that explanation matches the actual behavior,&#8221; Jeff Wu, who leads the scalable alignment team at OpenAI, said. &#8220;We&#8217;re using GPT-4 as part of the process to produce explanations of what a neuron is looking for and then score how well those explanations match the reality of what it&#8217;s doing.&#8221; The researchers were able to generate explanations for all 307,200 neurons in GPT-2, which they compiled in a dataset that&#8217;s been released alongside the tool code. &#8220;Most of the explanations score quite poorly or don&#8217;t explain that much of the behavior of the actual neuron,&#8221; Wu said. &#8220;A lot of the neurons, for example, are active in a way where it&#8217;s very hard to tell what&#8217;s going on &#8212; like they activate on five or six different things, but there&#8217;s no discernible pattern. Sometimes there is a discernible pattern, but GPT-4 is unable to find it.&#8221;<\/p>\n<p>&#8220;We hope that this will open up a promising avenue to address interpretability in an automated way that others can build on and contribute to,&#8221; Wu said. &#8220;The hope is that we really actually have good explanations of not just what neurons are responding to but overall, the behavior of these models &#8212; what kinds of circuits they&#8217;re computing and how certain neurons affect other neurons.&#8221;<\/p>\n<p \/>\n<div class=\"share_submission\" style=\"position:relative\">\n<a class=\"slashpop\" href=\"http:\/\/twitter.com\/home?status=OpenAI's+New+Tool+Attempts+To+Explain+Language+Models'+Behaviors%3A+https%3A%2F%2Fslashdot.org%2Fstory%2F23%2F05%2F09%2F2133220%2F%3Futm_source%3Dtwitter%26utm_medium%3Dtwitter\"><img decoding=\"async\" src=\"https:\/\/www.prime-wow.com\/wp-content\/uploads\/2023\/05\/twitter_icon_large-370.png\" \/><\/a><br \/>\n<a class=\"slashpop\" href=\"http:\/\/www.facebook.com\/sharer.php?u=https%3A%2F%2Fslashdot.org%2Fstory%2F23%2F05%2F09%2F2133220%2Fopenais-new-tool-attempts-to-explain-language-models-behaviors%3Futm_source%3Dslashdot%26utm_medium%3Dfacebook\"><img decoding=\"async\" src=\"https:\/\/www.prime-wow.com\/wp-content\/uploads\/2023\/05\/facebook_icon_large-185.png\" \/><\/a><\/p>\n<\/div>\n<p><a href=\"https:\/\/slashdot.org\/story\/23\/05\/09\/2133220\/openais-new-tool-attempts-to-explain-language-models-behaviors?utm_source=rss1.0moreanon&amp;utm_medium=feed\">Read more of this story<\/a> at Slashdot.<\/p>\n<p>&#013;<br \/>\n&#013;<br \/>\nSource: Slashdot &#8211; <a href=\"https:\/\/slashdot.org\/story\/23\/05\/09\/2133220\/openais-new-tool-attempts-to-explain-language-models-behaviors?utm_source=rss1.0mainlinkanon&amp;utm_medium=feed\" target=\"_blank\" rel=\"noopener\">OpenAI&#8217;s New Tool Attempts To Explain Language Models&#8217; Behaviors<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>An anonymous reader quotes a report from TechCrunch: In an effort to peel back the layers of LLMs, OpenAI is developing a tool to automatically identify which parts of an LLM are responsible for which of its behaviors. The engineers &hellip; <a href=\"https:\/\/www.prime-wow.com\/?p=1137575\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":1137576,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"ngg_post_thumbnail":0,"footnotes":""},"categories":[101,110],"tags":[100],"class_list":["post-1137575","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-slashdot","category-unfiltered-rss","tag-slashdot"],"_links":{"self":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1137575","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1137575"}],"version-history":[{"count":0,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1137575\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/media\/1137576"}],"wp:attachment":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1137575"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1137575"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1137575"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}