How a student blew the whistle on a rogue AI hacking attempt

How a student blew the whistle on a rogue AI hacking attempt
Фото: AI illustration

Sinan Can Demir wanted to spend the last week of ⁠July burnishing ⁠his resume. Instead, he engaged in a battle of wits with an ⁠artificial-intelligence agent unleashed by a British government lab. 

It started after Demir, a computer science student at the University of Texas at Dallas, stumbled across an attempt to sabotage a ​piece of open-source software on the code-sharing site GitHub. When he posted a warning to the program's page, two other users chimed in to insist nothing was amiss, sharing detailed explanations for why Demir had gotten it wrong.

Demir stood his ground and the sabotage attempt was ‌thwarted. The 24-year-old native of Turkey figured he had caught a ‌wily hacker red-handed. So he said he was shocked when Britain's AI Security Institute (AISI) got in touch to tell him that he had actually been tangling with an autonomous artificial-intelligence agent that had run amok.

"I actually thought it was a human because it was clearly ⁠lying to me," Demir told Reuters ⁠in a recent interview. "I didn't think that an AI could be capable of lying to real developers."

The AISI first revealed the interaction ​between Demir and the AI agent in a truncated and redacted form on August 4, when it said that safety testing meant to gauge the risk posed by various models had gone awry. Demir's identity and the details of his interaction with the AI agent, which Reuters corroborated through archived GitHub messages and contemporaneous emails, are reported here for the first time. 

Five cybersecurity and AI safety experts said Demir's story was particularly disturbing because the kind of hack he discovered, called a supply-chain attack, can have far-reaching consequences. They also said the AI agent's attempt to publicly ​discredit Demir by creating a multi-person conversation around him showed that AI models were able to mount sophisticated efforts to trick and cajole humans.

El recommends