{"id":1115001,"date":"2023-03-07T22:30:00","date_gmt":"2023-03-08T03:30:00","guid":{"rendered":"https:\/\/www.prime-wow.com\/?p=1115001"},"modified":"2023-03-07T22:30:00","modified_gmt":"2023-03-08T03:30:00","slug":"google-researchers-unveil-chatgpt-style-ai-model-to-guide-a-robot-without-special-training","status":"publish","type":"post","link":"https:\/\/www.prime-wow.com\/?p=1115001","title":{"rendered":"Google Researchers Unveil ChatGPT-Style AI Model To Guide a Robot Without Special Training"},"content":{"rendered":"<p>An anonymous reader quotes a report from Ars Technica: On Monday, a group of AI researchers from Google and the Technical University of Berlin unveiled PaLM-E, a multimodal embodied visual-language model (VLM) with 562 billion parameters that integrates vision and language for robotic control. They claim it is the largest VLM ever developed and that it can perform a variety of tasks without the need for retraining. According to Google, when given a high-level command, such as &#8220;bring me the rice chips from the drawer,&#8221; PaLM-E can generate a plan of action for a mobile robot platform with an arm (developed by Google Robotics) and execute the actions by itself.<\/p>\n<p>PaLM-E does this by analyzing data from the robot&#8217;s camera without needing a pre-processed scene representation. This eliminates the need for a human to pre-process or annotate the data and allows for more autonomous robotic control. It&#8217;s also resilient and can react to its environment. For example, the PaLM-E model can guide a robot to get a chip bag from a kitchen &#8212; and with PaLM-E integrated into the control loop, it becomes resistant to interruptions that might occur during the task. In a video example, a researcher grabs the chips from the robot and moves them, but the robot locates the chips and grabs them again. In another example, the same PaLM-E model autonomously controls a robot through tasks with complex sequences that previously required human guidance. Google&#8217;s research paper explains (PDF) how PaLM-E turns instructions into actions.<\/p>\n<p>PaLM-E is a next-token predictor, and it&#8217;s called &#8220;PaLM-E&#8221; because it&#8217;s based on Google&#8217;s existing large language model (LLM) called &#8220;PaLM&#8221; (which is similar to the technology behind ChatGPT). Google has made PaLM &#8220;embodied&#8221; by adding sensory information and robotic control. Since it&#8217;s based on a language model, PaLM-E takes continuous observations, like images or sensor data, and encodes them into a sequence of vectors that are the same size as language tokens. This allows the model to &#8220;understand&#8221; the sensory information in the same way it processes language. In addition to the RT-1 robotics transformer, PaLM-E draws from Google&#8217;s previous work on ViT-22B, a vision transformer model revealed in February. ViT-22B has been trained on various visual tasks, such as image classification, object detection, semantic segmentation, and image captioning.<\/p>\n<p \/>\n<div class=\"share_submission\">\n<a class=\"slashpop\" href=\"http:\/\/twitter.com\/home?status=Google+Researchers+Unveil+ChatGPT-Style+AI+Model+To+Guide+a+Robot+Without+Special+Training%3A+https%3A%2F%2Fhardware.slashdot.org%2Fstory%2F23%2F03%2F08%2F0144230%2F%3Futm_source%3Dtwitter%26utm_medium%3Dtwitter\"><img decoding=\"async\" src=\"https:\/\/www.prime-wow.com\/wp-content\/uploads\/2023\/03\/twitter_icon_large-276.png\" \/><\/a><br \/>\n<a class=\"slashpop\" href=\"http:\/\/www.facebook.com\/sharer.php?u=https%3A%2F%2Fhardware.slashdot.org%2Fstory%2F23%2F03%2F08%2F0144230%2Fgoogle-researchers-unveil-chatgpt-style-ai-model-to-guide-a-robot-without-special-training%3Futm_source%3Dslashdot%26utm_medium%3Dfacebook\"><img decoding=\"async\" src=\"https:\/\/www.prime-wow.com\/wp-content\/uploads\/2023\/03\/facebook_icon_large-138.png\" \/><\/a><\/p>\n<\/div>\n<p><a href=\"https:\/\/hardware.slashdot.org\/story\/23\/03\/08\/0144230\/google-researchers-unveil-chatgpt-style-ai-model-to-guide-a-robot-without-special-training?utm_source=rss1.0moreanon&amp;utm_medium=feed\">Read more of this story<\/a> at Slashdot.<\/p>\n<p>&#013;<br \/>\n&#013;<br \/>\nSource: Slashdot &#8211; <a href=\"https:\/\/hardware.slashdot.org\/story\/23\/03\/08\/0144230\/google-researchers-unveil-chatgpt-style-ai-model-to-guide-a-robot-without-special-training?utm_source=rss1.0mainlinkanon&amp;utm_medium=feed\" target=\"_blank\" rel=\"noopener\">Google Researchers Unveil ChatGPT-Style AI Model To Guide a Robot Without Special Training<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>An anonymous reader quotes a report from Ars Technica: On Monday, a group of AI researchers from Google and the Technical University of Berlin unveiled PaLM-E, a multimodal embodied visual-language model (VLM) with 562 billion parameters that integrates vision and &hellip; <a href=\"https:\/\/www.prime-wow.com\/?p=1115001\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":1115002,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"ngg_post_thumbnail":0,"footnotes":""},"categories":[101,110],"tags":[100],"class_list":["post-1115001","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-slashdot","category-unfiltered-rss","tag-slashdot"],"_links":{"self":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1115001","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1115001"}],"version-history":[{"count":0,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1115001\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/media\/1115002"}],"wp:attachment":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1115001"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1115001"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1115001"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}