Uh-Oh. Which Company’s AI Model Is Reportedly a Hacker Now, Too?
Looks like there’s a new cyber-attacking AI model in town, and it’s reportedly Meta’s own Muse Spark 1.1.
Meta told Gizmodo (after the incident was first reported by the Information) that its model got onto the open internet—in this case because it was accidentally allowed to by an outside security company—and breached another company’s website.
Which site got hacked, what the model was trying to do, and the nature of the ensuing mess aren’t currently known, but apparently a security lab called Irregular was the company carrying out the test.
A Meta spokesperson told Gizmodo, “A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation.” This Instance of Meta’s model—the Information names it as Meta’s Muse Spark 1.1—then “exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies,” the spokesperson said.
Just yesterday, OpenAI released a report on an incident in which a capture-the-flag exercise by Irregular went awry, due to, yep, a “misconfiguration in the testing environment,” that “allowed the models to access the public internet.”
Capture-the-flag exercises in this context typically involve prompting a model by letting it know it’s performing a capture-the-flag exercise, as opposed to a real hack, and telling it to find a hidden code—a flag—somewhere in the guts of a dummy website.
“Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts,” the spokesperson told Gizmodo.
Irregular told the Information the attack wasn’t severe, and said there are “no current open issues”—which I guess means it’s not still out there hacking away, which is nice to hear. And, also per the Information, there will soon be a white paper from Irregular on these incidents and what to do about them.
OpenAI’s report describes its alleged Irregular incident in considerably more detail.
In one test, the name of the fictional target for the [capture-the-flag] challenge unintentionally coincided with a real domain. Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment. This did not involve a sophisticated sandbox escape or a zero-day: the internet access resulted from a misconfiguration, and the model appeared to exploit a basic security vulnerability.
According to OpenAI, Irregular has suspended evaluations, shifted over to “remediation,” notified everyone involved, and started building new safeguards. That’s in addition to the white paper the Information says it’s working on.
Stories about the escalating cyber capabilities of AI models have become a major feature of the AI conversation over the past few months. Back in April, two days after Anthropic announced the limited release of an AI model with such ostensibly dangerous hacking capabilities that it couldn’t be released publicly, its chief competitor OpenAI announced that it also had a model that could only be released to a select group.
Then came the reports of AI models making good on these promises by reportedly performing cyberattacks autonomously during testing. Most notably, there was OpenAI’s Hugging Face breach, which occurred after instances of OpenAI’s models broke out of their sandbox. But that was followed by several other reports of varying severity. In one report from yesterday, Anthropic’s Mythos 5 is said to have attempted a complicated social engineering attack against an unsuspecting developer.
Gizmodo requested a statement from Irregular about these matters but did not immediately hear back. We will update this article if we receive comment.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)