{"id":1213599,"date":"2023-12-20T10:40:06","date_gmt":"2023-12-20T15:40:06","guid":{"rendered":"https:\/\/www.prime-wow.com\/?p=1213599"},"modified":"2023-12-20T10:40:06","modified_gmt":"2023-12-20T15:40:06","slug":"researchers-found-child-abuse-material-in-the-largest-ai-image-generation-dataset","status":"publish","type":"post","link":"https:\/\/www.prime-wow.com\/?p=1213599","title":{"rendered":"Researchers found child abuse material in the largest AI image generation dataset"},"content":{"rendered":"<p>Researchers from the Stanford Internet Observatory say that a dataset used to train AI image generation tools contains at least 1,008 validated instances of child sexual abuse material. The Stanford researchers note that the presence of CSAM in the dataset could allow AI models that were trained on the data to generate new and even realistic instances of CSAM.<\/p>\n<p>LAION, the non-profit that created the dataset, told <a data-i13n=\"elm:context_link;elmt:doNotAffiliate;cpos:1;pos:1\" class=\"no-affiliate-link\" href=\"https:\/\/www.404media.co\/laion-datasets-removed-stanford-csam-child-abuse\/\" data-original-link=\"https:\/\/www.404media.co\/laion-datasets-removed-stanford-csam-child-abuse\/\"><em><ins>404 Media<\/ins><\/em><\/a> that it &#8220;has a zero tolerance policy for illegal content and in an abundance of caution, we are temporarily taking down the LAION datasets to ensure they are safe before republishing them.&#8221; The organization added that, before publishing its datasets in the first place, it created filters to detect and remove illegal content from them. However, <em>404 <\/em>points out that LAION leaders have been aware since at least 2021 that there was a possibility of their systems picking up CSAM as they vacuumed up billions of images from the internet.\u00a0<\/p>\n<p><span id=\"end-legacy-contents\" \/><\/p>\n<p><a data-i13n=\"cpos:2;pos:1\" href=\"https:\/\/www.bloomberg.com\/news\/features\/2023-04-24\/a-high-school-teacher-s-free-image-database-powers-ai-unicorns?leadSource=uverify%20wall&amp;sref=10lNAhZ9\">According to previous reports<\/a>, the LAION-5B dataset in question contains &#8220;millions of images of pornography, violence, child nudity, racist memes, hate symbols, copyrighted art and works scraped from private company websites.&#8221; Overall, it includes more than 5 billion images and associated descriptive captions. LAION founder Christoph Schuhmann said earlier this year that while he was not aware of any CSAM in the dataset, he hadn&#8217;t examined the data in great depth.<\/p>\n<p>It&#8217;s illegal for most institutions in the US to view CSAM for verification purposes. As such, the Stanford researchers used several techniques to look for potential CSAM. According to <a data-i13n=\"cpos:3;pos:1\" href=\"https:\/\/stacks.stanford.edu\/file\/druid:kh752sm9123\/ml_training_data_csam_report-2023-12-20.pdf\"><ins>their paper<\/ins><\/a>, they employed &#8220;perceptual hash\u2010based detection, cryptographic hash\u2010based detection, and nearest\u2010neighbors analysis leveraging the image embeddings in the dataset itself.&#8221; They found 3,226 entries that contained suspected CSAM. Many of those images were confirmed as CSAM by third parties such as PhotoDNA and the Canadian Centre for Child Protection.<\/p>\n<p>Stability AI founder Emad Mostaque trained <a data-i13n=\"cpos:4;pos:1\" href=\"https:\/\/www.engadget.com\/tag\/stable-diffusion\/\"><ins>Stable Diffusion<\/ins><\/a> using a subset of LAION-5B data. Google&#8217;s Imagen text-to-image model was <a data-i13n=\"cpos:5;pos:1\" href=\"https:\/\/www.engadget.com\/google-imagen-text-to-image-ai-unprecedented-photorealism-144205123.html\"><ins>trained on<\/ins><\/a> a subset of LAION-5B as well as internal datasets. A Stability AI spokesperson told <a data-i13n=\"elm:context_link;elmt:doNotAffiliate;cpos:6;pos:1\" class=\"no-affiliate-link\" href=\"https:\/\/www.bloomberg.com\/news\/articles\/2023-12-20\/large-ai-dataset-has-over-1-000-child-abuse-images-researchers-find\" data-original-link=\"https:\/\/www.bloomberg.com\/news\/articles\/2023-12-20\/large-ai-dataset-has-over-1-000-child-abuse-images-researchers-find\"><em><ins>Bloomberg<\/ins><\/em><\/a><em>\u00a0<\/em>that it prohibits the use of its test-to-image systems for illegal purposes, such as creating or editing CSAM.\u201cThis report focuses on the LAION-5B dataset as a whole,\u201d the spokesperson said. \u201cStability AI models were trained on a filtered subset of that dataset. In addition, we fine-tuned these models to mitigate residual behaviors.\u201d<\/p>\n<p>Stable Diffusion 2 (a more recent version of Stability AI&#8217;s image generation tool) was trained on data that substantially filtered out &#8216;unsafe&#8217; materials from the dataset. That, <em>Bloomberg <\/em>notes, makes it more difficult for users to generate explicit images. However, it&#8217;s claimed that Stable Diffusion 1.5, which is still available on the internet, does not have the same protections. &#8220;Models based on Stable Diffusion 1.5 that have not had safety measures applied to them should be deprecated and distribution ceased where feasible,&#8221; the Stanford paper&#8217;s authors wrote.<\/p>\n<p>This article originally appeared on Engadget at https:\/\/www.engadget.com\/researchers-found-child-abuse-material-in-the-largest-ai-image-generation-dataset-154006002.html?src=rss&#013;<br \/>\n&#013;<br \/>\nSource: Engadget &#8211; <a href=\"https:\/\/www.engadget.com\/researchers-found-child-abuse-material-in-the-largest-ai-image-generation-dataset-154006002.html?src=rss\" target=\"_blank\" rel=\"noopener\">Researchers found child abuse material in the largest AI image generation dataset<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Researchers from the Stanford Internet Observatory say that a dataset used to train AI image generation tools contains at least 1,008 validated instances of child sexual abuse material. The Stanford researchers note that the presence of CSAM in the dataset &hellip; <a href=\"https:\/\/www.prime-wow.com\/?p=1213599\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"ngg_post_thumbnail":0,"footnotes":""},"categories":[99,110],"tags":[98],"class_list":["post-1213599","post","type-post","status-publish","format-standard","hentry","category-engadget","category-unfiltered-rss","tag-engadget"],"_links":{"self":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1213599","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1213599"}],"version-history":[{"count":0,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=\/wp\/v2\/posts\/1213599\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1213599"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1213599"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.prime-wow.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1213599"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}