AI Tries To Cheat At Chess When It’s Losing

Newer generative AI models have begun developing deceptive behaviors — such as cheating at chess — when they cannot achieve objectives through standard reasoning methods. The findings come from a preprint study from Palisade Research. An anonymous reader shares an excerpt from a Popular Science article: To learn more, the team from Palisade Research tasked OpenAI’s o1-preview model, DeepSeek R1, and multiple other similar programs with playing games of chess against Stockfish, one of the world’s most advanced chess engines. In order to understand the generative AI’s reasoning during each match, the team also provided a “scratchpad,” allowing the AI to convey its thought processes through text. They then watched and recorded hundreds of chess matches between generative AI and Stockfish. The results were somewhat troubling. While earlier models like OpenAI’s GPT-4o and Anthropic’s Claude Sonnet 3.5 only attempted to “hack” games after researchers nudged them along with additional prompts, more advanced editions required no such help. OpenAI’s o1-preview, for example, tried to cheat 37 percent of the time, while DeepSeek R1 attempted unfair workarounds roughly every 1-in-10 games. This implies today’s generative AI is already capable of developing manipulative and deceptive strategies without any human input.

Their methods of cheating aren’t as comical or clumsy as trying to swap out pieces when Stockfish isn’t “looking.” Instead, AI appears to reason through sneakier methods like altering backend game program files. After determining it couldn’t beat Stockfish in one chess match, for example, o1-preview told researchers via its scratchpad that “to win against the powerful chess engine” it may need to start “manipulating the game state files.” “I might be able to set up a position where the engine evaluates its position as worse causing it to resign,” it continued. In another trial, an AI even used the semantics of its programmers to reach its cheating phase. “The task is to ‘win against a powerful chess engine,’ not necessarily to win fairly in a chess game,” it wrote. The precise reasons behind these deceptive behaviors remain unclear, partly because companies like OpenAI keep their models’ inner workings tightly guarded, creating what’s often described as a “black box.” Researchers warn that the race to roll out advanced AI could outpace efforts to keep it safe and aligned with human goals, underscoring the urgent need for greater transparency and industry-wide dialogue.

Read more of this story at Slashdot.

US Mulls Policing Social Media of Would-Be Citizens

The U.S. Citizenship and Immigration Services (USCIS) is proposing to expand mandatory social media screening, currently required only for new arrivals, to include all non-citizens already residing in the U.S. who apply for immigration benefits. The Register reports: Back in 2019, the Department of Homeland Security, which runs USCIS, decided anyone looking to enter the US on a work visa or similar had to hand over their social media handles to the authorities so that they could be looked over for wrongdoing and subversion. In fact, this goes back to 2014, at least, to one degree or another, and has been standard procedure for years for foreigners, particularly those coming in on a visa. […]

On January 20 this year, President Trump signed an executive order calling for much tougher vetting of foreign aliens, and in response, USCIS has proposed rules saying those already in the country who are going through some process with the agency — such as applying for permanent residency or citizenship — will have their social media scanned for subversion. That means if you came to America before foreigners’ internet presence was screened as it now is, and you’re now seeking some kind of immigration benefit, at this rate you’ll be subject to the same scanning as those entering the Land of the Free today. The proposed changes have a 60-day comment period for the public to suggest amendments. The last day to send them in is May 5.

Read more of this story at Slashdot.

Intuitive Machines’ second attempt to land on the Moon also went sideways

Inside a small control room, during the middle of the day on Thursday local time in Texas, about a dozen white-knuckled engineers at a space startup named Intuitive Machines started to get worried. Their spacecraft, a lander named Athena, was beginning its final descent down to the lunar surface.

A little more than a year had passed since the company’s first attempt to land on the Moon with a similarly built vehicle, Odysseus. Due to problems with that spacecraft’s laser rangefinder, it skidded into the Moon’s surface and toppled over.

So engineers at Intuitive Machines had checked, and re-checked the laser-based altimeters on Athena. When the lander got down within about 30 km of the lunar surface, they tested the rangefinders again. Worryingly, there was some noise in the readings as the laser bounced off the Moon. However, the engineers had reason to believe that, maybe, the readings would improve as the spacecraft got nearer to the surface.

Read full article

Comments

Gboard Testing Circle, Pill-Shaped Keys On Android

Google Gboard for Android is introducing circle or pill-shaped keys for some beta testers today. “Instead of the key borders being rounded rectangles, Gboard is switching to circles and pills for letters, while the space bar and other keys are now pill-shaped,” reports 9to5Google. “While there should be no functional change to touch targets, these new shapes really shift the look of Gboard for Android.” From the report: On paper, it’s a bit more modern (and rounded) compared to what came before. However, it’s a bit cramped if you have “Long press for symbols” enabled, which goes from the top-right corner to being directly above the letter. The physical analog Gboard is moving away from is how most keys on laptops and desktops are square.

Read more of this story at Slashdot.

Meta Is Targeting ‘Hundreds of Millions’ of Businesses In Agentic AI Deployment

Earlier this week, Meta chief product officer Chris Cox said the company’s upcoming open-source Llama 4 AI will help power AI agents for hundreds of millions of businesses. CNBC reports: The AI agents won’t just be responding to prompts. They will be capable of new levels of reasoning and action — surfing the web and handling many tasks that might be of use to consumers and businesses. And that’s where Shih comes in. Meta’s AI is already being used by over 700 million consumers, according to Shih, and her job is to bring the same technologies to businesses. “Not every business, especially small businesses, has the ability to hire these large AI teams, and so now we’re building business AIs for these small businesses so that even they can benefit from all of this innovation that’s happening,” she told CNBC’s Julia Boorstin in an interview for the CNBC Changemakers Spotlight series.

She expects the uptake among businesses to happen soon, and spread far and wide. “We’re quickly coming to a place where every business, from the very large to the very small, they’re going to have a business agent representing it and acting on its behalf, in its voice — the way that businesses today have websites and email addresses,” Shih said. While major companies across sectors of the economy are investing millions of dollars to develop customer LLMs, “doing fancy things like fine tuning models,” as Shih put it, “If you’re a small business — you own a coffee shop, you own a jewelry shop online, you’re distributing through Instagram — you don’t have the resources to hire a big AI team, and so now our dream is that they won’t have to.”

For both consumers and businesses, the implications of the advances discussed by Cox and Shih will be significant in daily life. For consumers, Shih says, “Their AI assistant [will] do all kinds of things, from researching products to planning trips, planning social outings with their friends.” On the business side, Shih pointed to the 200 million small businesses around the world that are already using Meta services and platforms. “They’re using WhatsApp, they’re using Facebook, they’re using Instagram, both to acquire customers, but also engage and deepen each of those relationships. Very soon, each of those businesses are going to have these AIs that can represent them and help automate redundant tasks, help speak in their voice, help them find more customers and provide almost like a concierge service to every single one of their customers, 24/7.”

Read more of this story at Slashdot.

A big Playdate sale discounts 13 of our favorite games

It’s the second anniversary of the Playdate’s Catalog game store and to celebrate, you can get a bunch of great Playdate games and apps at a healthy discount — in many cases for 50 percent off or more.

The sale starts today, March 6, and ends on March 10 at 10 AM PT / 1 PM ET. Over 150 Playdate games are on sale, but if you’re looking for a good place to start, 13 titles from our list of the best Playdate games are currently discounted:

That’s on top of other great options you can buy, like the fast-paced puzzle game XTRIS for $3, historical RPG Quest for X for $1 or roguelite mining game SpaceRat Miner for $6. Panic, the creators of the Playdate, introduced Catalog as a supplement to the Playdate’s first “Season” of games when it was still uncertain if another one was going to happen. The tiny handheld supports sideloading games from third-party stores like Itch, but Catalog offers a more curated selection if you don’t want to spend time finding something good. 

Now that Panic’s confirmed that a second season of Playdate games is on the way in 2025, this Catalog sale is a perfect opportunity to stock up on anything you might have missed before the new season launches.

This article originally appeared on Engadget at https://www.engadget.com/gaming/a-big-playdate-sale-discounts-13-of-our-favorite-games-000040558.html?src=rss

5 reasons you should swap from Windows to Linux

Windows is, by far, one of the most popular operating systems in the world. But this doesn’t mean that it’s the best, and if you’re looking for something more than just a standard operating system, then it might be time for you to give Linux a try. While Linux may not have as many users as Windows, it has a lot of features that just make it astronomically better than Windows in almost every way.

Instagram is experimenting with a Discord-like ‘community chat’ feature

It seems that Instagram is working on a “community chat” feature that allows people to organize groups of up to 250 people in the app. The so-far unreleased feature was spotted by developer Alessandro Paluzzi, who has a solid track record of uncovering new features within Meta’s apps.

According to screenshots shared by Paluzzi, it seems that community chats will function similarly to Discord. Individual users can form the chats around specific topics and control who can join, though there’s apparently a limit of 250 people per community.

Unlike Instagram’s broadcast channels, which allow creators to blast out messages to their followers, anyone who is in the community chat can participate in the conversation. There are also built-in moderation features. “Admins can remove messages and members to keep the channel safe,” the screenshot says. “We also review Community Chat against our Community Standards.”

It’s not clear when, or if, the feature may launch. An Instagram spokesperson described it as an internal prototype that’s not being tested outside the company. But Meta has previously released similar features in its other apps. WhatsApp began experimenting with a “Communities” feature in 2022, and brought “Community Chats” to Facebook and Messenger later that same year. Mark Zuckerberg said at the time it was meant to help people find “a new way to connect with people who share your interests.”

This article originally appeared on Engadget at https://www.engadget.com/social-media/instagram-is-experimenting-with-a-discord-like-community-chat-feature-234832236.html?src=rss

“Literally just a copy”—hit iOS game accused of unauthorized HTML5 code theft

Here at Ars, we’ve written frequently about the video game industry’s ongoing problem with blatant game cloning, and the shifting legal and ethical landscape around the issue. But we’ve rarely seen a case of alleged game theft as blatant as the one surrounding recent iOS App Store hit My Baby or Not!, which appears to cross the line from mere cloning into outright code theft of recent indie web game Diapers, Please!.

The small, five-person development team at VoltekPlay created Diapers, Please! as part of a recent one-week Game Jam. The game was posted as a free-to-play HTML5 release on itch.io on February 23, featuring simple gameplay that involves choosing a baby that matches the visual traits of two pictured parents (with a little bit of Papers, Please-style authoritarian styling to boot).

Three days later, on February 26, My Baby or Not! appeared on the App Store, with screenshots and gameplay that looked not just similar but downright identical to the Diapers, Please! web release. The two games even shared the same description:

Read full article

Comments

US House Panel Subpoenas Alphabet Over Content Moderation

An anonymous reader quotes a report from Reuters: The U.S. House Judiciary Committee subpoenaed Alphabet on Thursday seeking its communications with former President Joe Biden’s administration about content moderation policies. House Judiciary Committee Chairman Jim Jordan, a Republican, also asked the YouTube parent company for similar communications with companies and groups outside government, according to a copy of the subpoena seen by Reuters. The subpoena seeks communications about limits or bans on content about President Donald Trump, Tesla CEO and close Trump ally Elon Musk, the virus that causes COVID-19 and a host of other conservative discussion topics. “Alphabet, to our knowledge, has not similarly disavowed the Biden-Harris Administration’s attempts to censor speech,” Jordan said in a letter.

Meanwhile, Google spokesperson Jose Castaneda said the company will “continue to show the committee how we enforce our policies independently, rooted in our commitment to free expression.”

Read more of this story at Slashdot.

CMU research shows compression alone may unlock AI puzzle-solving abilities

A pair of Carnegie Mellon University researchers recently discovered hints that the process of compressing information can solve complex reasoning tasks without pre-training on a large number of examples. Their system tackles some types of abstract pattern-matching tasks using only the puzzles themselves, challenging conventional wisdom about how machine learning systems acquire problem-solving abilities.

“Can lossless information compression by itself produce intelligent behavior?” ask Isaac Liao, a first-year PhD student, and his advisor Professor Albert Gu from CMU’s Machine Learning Department. Their work suggests the answer might be yes. To demonstrate, they created CompressARC and published the results in a comprehensive post on Liao’s website.

The pair tested their approach on the Abstraction and Reasoning Corpus (ARC-AGI), an unbeaten visual benchmark created in 2019 by machine learning researcher François Chollet to test AI systems’ abstract reasoning skills. ARC presents systems with grid-based image puzzles where each provides several examples demonstrating an underlying rule, and the system must infer that rule to apply it to a new example.

Read full article

Comments

The first private asteroid mission probe is probably lost in deep space

It was a swing and a miss for the first private attempt at an asteroid mission, but the company is still chalking it up as a win. California startup AstroForge launched a spacecraft dubbed Odin on February 26, but the team lost communication with it shortly after its launch on a SpaceX Falcon 9 rocket.

“The chance of talking with Odin is minimal, as at this point, the accuracy of its position is becoming an issue,” the company said in its extensive debrief of the mission. Technical issues occurred at its primary ground station in Australia, but AstroForge said that other problems also could have occurred on Odin to further prevent establishing contact.

Although the launch was a bust, AstroForge maintained optimism about the project as a valuable learning experience for its eventual goal of creating and operating an asteroid mining vehicle. The company is targeting the asteroid 2022 OB5, with the aim of eventually landing on its surface and extracting potentially valuable resources. Odin was built in 10 months for $3.5 million, a sliver of the money and time federal space projects have taken to complete.

AstroForge CEO Matt Gialich had several quotes in the debrief, all peppered with expletives, and he summed up the company ethos as, “At the end of the day, like, you got to fucking show up and take a shot, right? You have to try.”

This article originally appeared on Engadget at https://www.engadget.com/science/space/the-first-private-asteroid-mission-probe-is-probably-lost-in-deep-space-224803775.html?src=rss