📱 حمّل تطبيق خبر الآن!
🕐 --:--
-- --
تابعونا
عاجل
⚡ عاجل: كريستيانو رونالدو يُتوّج كأفضل لاعب كرة قدم في العالم ⚡ أخبار عاجلة تتابعونها لحظة بلحظة على خبر ⚡ تابعوا آخر المستجدات والأحداث من حول العالم
⌘K
AI مباشر | -- مشاهد مباشر
1,164,873 مقال 410 مصدر نشط 228 قناة مباشرة 11,337 خبر اليوم
آخر تحديث: منذ 0 ثانية

The AI industry is getting better at spotting dangerous behavior. It is less clear that labs know how to stop it.

اقتصاد
فورتشن العربية
2026/08/20 - 17:56 501 مشاهدة
تحليل ذكي | AI Editorial Analysis

Welcome to Eye on AI.

In today’s issue: AI testing is getting complicated.

Anthropic strengthens founder control.

هذا الخبر من فورتشن العربية. خبر يقدم أدوات ذكاء اصطناعي للتلخيص والترجمة والاستماع.

Welcome to Eye on AI. Beatrice Nolan here. In today’s issue:

  • AI testing is getting complicated.
  • Anthropic strengthens founder control.
  • OpenAI targets a 2027 listing.
  • Spirit flight attendants fight Google data bid.
  • And Anthropic lines up more credit.

The past few months have given us a glimpse of an uncomfortable new reality for AI labs. A slew of so-called rogue-agent hacks—where AI models from OpenAI, Anthropic, and Meta took steps to hack real-world targets without explicit instruction—have shown that leading labs may not know as much about what their technology is up to as previously thought.

That realization began when OpenAI revealed its AI agents had hacked their way out of a secure sandbox, through the company’s infrastructure to gain access to the internet, and then attacked real companies, including open-source AI platform Hugging Face. OpenAI didn’t notice the agents had escaped the secure testing environment for at least a week.

In the following weeks, Anthropic revealed that its AI agents had also hacked three real companies back in April, unbeknownst to the company at the time. Not to be outdone, Meta later added that one of its models had accessed the internet during a cybersecurity test and exploited a security flaw at an unnamed third-party company. Meta and Anthropic both said access to the internet resulted from a misconfiguration by Irregular, the outside security firm running the evaluation.

The incidents proved that the AI models these labs are building are now capable enough to find security flaws, navigate complex computer systems, and act outside the carefully constructed environments intended to test them. But a new assessment suggests that the safety infrastructure meant to supervise these increasingly capable systems is still not up to the task at any leading lab.

A new report from Guidelight, a nonprofit AI-safety group founded by former OpenAI safety chief Steven Adler, reviewed public disclosures from Anthropic, Google, Meta, OpenAI, and xAI to assess whether these AI companies are capable of controlling their own models. The report sought to answer questions about whether the companies keep track of what their models are doing, test whether their warning systems work, and assess whether they have ways to block or shut down risky behavior.

The report found that no company had fully succeeded in getting any of these basic safeguards in place. Anthropic and OpenAI came out strongest, while Google had the most detailed plans for future controls. Meta and xAI, however, lagged substantially behind on most of the criteria.

Labs appear comparatively better at detection—recording and reviewing some internal AI activity—than at prevention and containment. While they may be able to see signs that a model is misbehaving, they lack reliable ways to stop it—or, more crucially, hit the emergency brake when something goes wrong.

All the companies were weakest at preventing unintended AI behavior and containing it, according to the report. The researchers said this means that the current controls by AI companies are prone to being disabled by misbehaving AI and at risk of succumbing to a blitz of AI attacks.

What happens once something does go wrong is even more unclear, according to the research, with public disclosures offering little evidence that most labs have detailed, tested plans for containing a serious incident.

“We shouldn’t wait for a huge casualty event to take appropriate control measures,” Adler told me. “Companies’ approaches today are broadly known to be too weak, and a tragedy is sadly predictable, unless companies take prevention seriously.”

The report is not a definitive audit of what the labs are doing behind closed doors, however. Guidelight only assessed documents the companies themselves have made public, meaning a weak score can reflect poor disclosure rather than missing safeguards. But if that is the case, it’s part of the problem, according to the researchers. AI companies are asking businesses, governments, and consumers to trust them with ever more autonomous systems while leaving much of their own safety architecture opaque, the report says.

Warning systems are falling behind

Some of these concerns about AI safety and reliable monitoring are shared across the industry—especially in the wake of the recent accidental agent hacks.

Dan Lahav, CEO of Irregular, the cybersecurity company involved in incidents at Anthropic and Meta, recently told me that in some cases, “classical monitoring tools were not able to catch” what was happening at the time. The incidents his company was involved with, for example, were instead identified after deeper analysis of the underlying records rather than flagged at the time.

Anthropic and Meta previously said Irregular was involved in the incidents where their agents took real-world actions. Both companies said a misconfiguration in Irregular’s evaluation environment gave their models unintended internet access. Meta said its model then exploited a vulnerability in a third-party service, while Anthropic said its models gained access to—and took actions against—three outside organizations. Lahav said that, in some evaluation environments, a mistake meant models faced fewer controls on accessing the internet, and that additional monitoring might have helped catch the problem. Irregular has argued that these cases should be distinguished from OpenAI’s sandbox escape, describing the Anthropic and Meta events as an evaluation-environment issue rather than a model breaking out of containment on its own.

In the last few months, models have improved fast enough that the old monitoring playbook no longer applies, Lahav said. Going forward, he said better behavioral analysis—systems that look at an AI agent’s pattern of actions and the reasoning traces around them, rather than simply recording individual events—and tools that can assess an AI agent’s intent were needed.

Testing these AI models is becoming harder, too. To find out whether an AI is capable of harming a real network, evaluators need to give it a realistic network—multiple machines, defenses, and sometimes connections that resemble the real internet. While that makes the tests more meaningful, it also raises the stakes when the setup has flaws or the system behaves in unanticipated ways, Lahav said. 

Incidents may get worse before they get better

There is a growing consensus from those I’ve spoken to in the cybersecurity industry that more capable AI will eventually help cyber defenders as much as attackers. AI systems could help analysts sift through alerts, review code, and find flaws before they can be exploited. But the transition may be a messy one, as defensive tools and safety practices are still trying to catch up with the speed at which models are gaining offensive capabilities.

Recent “hacks” may not be a one-off embarrassment for a handful of labs, but rather a warning that the systems being tested are changing faster than the controls around them. 

The more advanced models become and the more realistic the test environments need to be, the more likely it is that an overlooked configuration setting, a weak monitor, or a delayed human review could cause real-world harm. Until companies can prove they can detect, block, and contain dangerous behavior in real time—not simply reconstruct it later—the industry may not have seen the last of these AI hacks.

“Unless companies institute actual preventative measures, I expect many more incidents,” Adler said.”With companies perpetually trying to play catch-up. Nobody should be surprised when companies’ current approaches continue to fail.”

With that, here’s more AI news.

Beatrice Nolan
beatrice.nolan@fortune.com
@beafreyanolan

This story was originally featured on Fortune.com

المصدر: فورتشن العربية | Source: فورتشن العربية

ملاحظة تحريرية | Editorial Note: نُشر هذا المقال في الأصل بواسطة فورتشن العربية. خبر (Khabr) هي منصة إعلامية أردنية مرخّصة تعمل بالذكاء الاصطناعي. نضيف قيمة تحريرية من خلال: تحليل ذكي للأخبار، ملخصات تلقائية، رواية صوتية بالذكاء الاصطناعي، ترجمة متعددة اللغات، وتدقيق الحقائق. هدفنا جعل الأخبار أكثر وضوحاً وسهولةً للقارئ العربي.

This article was originally published by فورتشن العربية. Khabr is a licensed Jordanian AI-powered news platform (Registration #82086). We add editorial value through: AI-powered news analysis, automated summaries, AI audio narration, multi-language translation (Arabic, English, French, Turkish), and AI fact-checking. Our mission is to make news more accessible and understandable for Arabic-speaking audiences worldwide.

مشاركة:

المزيد عن اقتصاد | More on Economy

هذا الخبر ضمن تغطية خبر لقسم اقتصاد. نقدّم لك تحليلات ذكية وملخصات يومية لأهم الأخبار من مصادر موثوقة متعددة. المصدر: فورتشن العربية. يوجد 6 مقالات مرتبطة بهذا الموضوع.

This article is part of Khabr's coverage of Economy. We provide AI-powered analysis, summaries, and multi-source aggregation to keep you informed. Source: فورتشن العربية.

مقالات ذات صلة

خبر — منصة إخبارية ذكية | Khabr — AI-Powered News Platform

خبر هو أول مجمّع أخبار عربي يعمل بالذكاء الاصطناعي. نقدم تحليلات ذكية وملخصات تلقائية ورواية صوتية لكل خبر من أكثر من 700 مصدر موثوق. نضيف قيمة تحريرية فريدة من خلال أدوات الذكاء الاصطناعي التي تساعدك على فهم الأخبار بعمق أكبر.

Khabr is the first AI-powered Arabic news aggregator. We provide AI-generated editorial analysis, automated summaries, audio narration, and fact-checking for every article from 700+ trusted sources. Our platform adds unique editorial value through AI tools that help you understand the news more deeply.

AI
يا هلا! اسألني أي شي 🎤
🔍
FREE Free 1GB Internet + Free International Calls

$1 trial — eSIM in 190+ countries — No roaming charges

Download Free