{"id":1195177,"date":"2023-10-23T09:37:56","date_gmt":"2023-10-23T13:37:56","guid":{"rendered":"https:\/\/www.prime-wow.com\/?p=1195177"},"modified":"2023-10-23T09:37:56","modified_gmt":"2023-10-23T13:37:56","slug":"eureka-with-gpt-4-overseeing-training-robots-can-learn-much-faster","status":"publish","type":"post","link":"https:\/\/www.prime-wow.com\/?p=1195177","title":{"rendered":"Eureka: With GPT-4 overseeing training, robots can learn much faster"},"content":{"rendered":"<div id=\"rss-wrap\">\n<figure class=\"intro-image intro-left\">\n  <img decoding=\"async\" src=\"https:\/\/www.prime-wow.com\/wp-content\/uploads\/2023\/10\/eureka_hands-800x447-2.jpg\" alt=\"In this still captured from a video provided by Nvidia, a simulated robot hand learns pen tricks, trained by Eureka, using simultaneous trials.\" \/><\/p>\n<p class=\"caption\" style=\"font-size:0.8em\"><a href=\"https:\/\/cdn.arstechnica.net\/wp-content\/uploads\/2023\/10\/eureka_hands-scaled.jpg\" class=\"enlarge-link\" data-height=\"1432\" data-width=\"2560\">Enlarge<\/a> <span class=\"sep\">\/<\/span> In this still captured from a video provided by Nvidia, a simulated robot hand learns pen tricks, trained by Eureka, using simultaneous trials. (credit: <a rel=\"nofollow\" class=\"caption-link\" href=\"https:\/\/eureka-research.github.io\/\">Nvidia<\/a>)<\/p>\n<\/figure>\n<div><a name=\"page-1\" \/><\/div>\n<p>On Friday, researchers from Nvidia, UPenn, Caltech, and the University of Texas at Austin announced Eureka, an algorithm that uses OpenAI&#8217;s <a href=\"https:\/\/arstechnica.com\/information-technology\/2023\/03\/openai-announces-gpt-4-its-next-generation-ai-language-model\/\">GPT-4<\/a> language model for designing training goals (called &#8220;reward functions&#8221;) to enhance robot dexterity. The work aims to bridge the gap between high-level reasoning and low-level motor control, allowing robots to learn complex tasks rapidly using massively parallel simulations that run through trials simultaneously. According to the team, Eureka outperforms human-written reward functions by a substantial margin.<\/p>\n<p>Before robots can interact with the real world successfully, they need to learn how to move their robot bodies to achieve goals\u2014like picking up objects or moving. Instead of making a physical robot try and fail one task at a time to learn in a lab, researchers at Nvidia have been experimenting with using video game-like computer worlds (thanks to platforms called <a href=\"https:\/\/developer.nvidia.com\/isaac-sim\">Isaac Sim<\/a> and <a href=\"https:\/\/developer.nvidia.com\/isaac-gym\">Isaac Gym<\/a>) that simulate three-dimensional physics. These allow for massively parallel training sessions to take place in many virtual worlds at once, dramatically speeding up training time.<\/p>\n<p>&#8220;Leveraging state-of-the-art GPU-accelerated simulation in Nvidia Isaac Gym,&#8221; writes Nvidia on its <a href=\"https:\/\/eureka-research.github.io\/\">demonstration page<\/a>, &#8220;Eureka is able to quickly evaluate the quality of a large batch of reward candidates, enabling scalable search in the reward function space.&#8221; They call it &#8220;rapid reward evaluation via massively parallel <a href=\"https:\/\/en.wikipedia.org\/wiki\/Reinforcement_learning\">reinforcement learning<\/a>.&#8221;<\/p>\n<\/div>\n<p><a href=\"https:\/\/arstechnica.com\/?p=1977747#p3\">Read 6 remaining paragraphs<\/a> | <a href=\"https:\/\/arstechnica.com\/?p=1977747&amp;comments=1\">Comments<\/a><\/p>\n<p>&#013;<br \/>\n&#013;<br \/>\nSource: Ars Technica &#8211; <a href=\"https:\/\/arstechnica.com\/?p=1977747\" target=\"_blank\" rel=\"noopener\">Eureka: With GPT-4 overseeing training, robots can learn much faster<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Enlarge \/ In this still captured from a video provided by Nvidia, a simulated robot hand learns pen tricks, trained by Eureka, using simultaneous trials. (credit: Nvidia) On Friday, researchers from Nvidia, UPenn, Caltech, and the University of Texas at &hellip; <a href=\"https:\/\/www.prime-wow.com\/?p=1195177\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":1195181,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"ngg_post_thumbnail":0,"footnotes":""},"categories":[27,110],"tags":[73],"class_list":["post-1195177","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ars-technica","category-unfiltered-rss","tag-ars-technica"],"_links":{"self":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1195177","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1195177"}],"version-history":[{"count":0,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1195177\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/media\/1195181"}],"wp:attachment":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1195177"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1195177"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1195177"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}