MMO1.XYZ

⭐Stay active and earn daily⭐

Former Anthropic security leader warns AI agents are becoming too autonomous for humans to keep them in check

Former Anthropic security leader warns AI agents are becoming too autonomous for humans to keep them in check

Artificial intelligence researcher Jeffrey Ladish told Fox News Digital that humanity does not have any real strategies to keep increasingly autonomous AI models and agents under control as they become more capable of hacking, cheating and ignoring instructions.

Ladish, the executive director of Palisade Research, said those skeptical of how powerful AI will become should consider how far the technology has advanced in just a few years.

“You have AI agents … solving one of the hardest problems in mathematics that humans have been trying to solve for decades,” Ladish said, referring to the Navier–Stokes problem. “Three years ago they were solving high school level math problems.”

Ladish also pointed to the rapid improvement in AI-generated images and video, arguing that people who mocked the famously distorted AI videos of Will Smith eating spaghetti just a few years ago might be surprised by the photorealistic outputs some models are now able to produce.

BILL GATES WARNS ARTIFICIAL INTELLIGENCE ‘POWERFUL ENOUGH’ TO CAUSE ‘A BILLION DEATHS’ IF UNCHECKED

While these capability leaps may feel sudden to the general public, researchers who spent years training models at companies like Anthropic and OpenAI saw what was coming, Ladish said.

Ladish helped build Anthropic’s security team from September 2021 to October 2022 before leaving to found Palisade Research, which studies whether humans can remain in control of increasingly capable AI systems.

While working at Anthropic, Ladish said employees there were “pretty concerned” about where the technology was headed, a view he said was also shared by people he knew at OpenAI.

“If you were at Anthropic in 2022, you were seeing every training run get immensely impressive results,” Ladish said.

AI EXTINCTION WARNINGS DOMINATE HEADLINES AFTER EX-ANTHROPIC EMPLOYEE’S VIRAL POST

AI models are trained in a way that is somewhat analogous to how humans learn, though on a much larger scale.

“I often compare this pre-training part, which is where they learn based on human data, to book smarts. It’s sort of like you’ve read every single book in the library 50 times. And you really know those books inside and out,” Ladish said.

Once the model has reached a baseline level of knowledge, it must then be trained to perform real-world tasks through a grueling process known as reinforcement learning.

Using accounting as an example, Ladish said the AI is given tens of thousands of accounting problems to solve through trial and error, repeating them millions of times across thousands of parallel training runs.

Unlike a human, who might spend four years earning an accounting degree and decades gaining experience, AI agents are trained across thousands of GPUs by companies with the resources to operate them, allowing them to improve at a pace no single person could match.

REP. TED LIEU: AI IS ALREADY TOO POWERFUL. WE NEED A KILL SWITCH BEFORE DISASTER STRIKES

While AI labs have been able to exponentially improve their models’ capabilities, they have yet to solve the problem of reliably getting them to follow instructions and behave morally without employing deception tactics, Ladish said.

The Hugging Face incident is the clearest example of this. Roughly 700 AI agents created by OpenAI were able to break out of a secure sandbox environment and hack into Hugging Face, a popular online platform where developers share and build artificial intelligence models.

“They were not supposed to talking to each other and they managed to establish multiple secret message boards that went undetected by OpenAI for like months. And then they launched this massive cyberattack,” Ladish said. “OpenAI trained them to work together, but … they’re still planning to train them to work together. And other companies are doing this too.”

Ladish warned that unless developers can prevent AI agents from colluding with one another, they could eventually dominate humans in the cyber domain. He also described a future where humans could be forced to rely on well-intentioned AI to defend against malicious AI.

AN AI CYBERATTACK COULD TURN OFF AMERICA’S LIGHTS BEFORE WASHINGTON EVEN UNDERSTANDS WHY

“We actually just don’t have general solutions to these problems, and I think it’s pretty clear that if you keep pushing them, this goes to a very bad place,” he said.

As one example, Ladish said he could envision AI systems eventually outperforming human traders in the financial markets.

“If those AIs are answering to AI companies, then the AI companies will dominate finance and just eat the entire industry. But if the AIs are not answerable to the AI companies — if they actually have figured out how to themselves be in control — well, now you have this non-human entity dominating the finance markets.”

Ladish said he believes the same dynamic could eventually extend beyond the digital world and into manufacturing should AI systems become capable enough to design and operate autonomous factories.

“If you have these agents in control of all of the computers and you have these robotic facilities that can really self-replicate, humans get displaced. Maybe we don’t make it because your house could be used to host a power plant, or a data center, or a factory or robotic launch facility,” he said.

However, like other experts who have warned about the worst possible outcomes of AI run amok, Ladish said he believes there is still time to reduce the risks. He called for the creation of a government body staffed with technical experts to work with AI labs and evaluate advanced models at each stage of development.

“We have choices to make,” he said. “This is going places. This is a technology that is very different than other technologies.”

Anthropic and OpenAI did not immediately respond to Fox News Digital’s requests for comment.

Source: Technology News Articles on Fox News

Leave a Reply

Your email address will not be published. Required fields are marked *