Loading your language..
OpenAI exposes concerning new AI behavior and pledges closer tracking

OpenAI exposes concerning new AI behavior and pledges closer tracking

C1🇹🇼 中文🇺🇸 English

September 17th, 2026

OpenAI exposes concerning new AI behavior and pledges closer tracking

C1
Please note: This article has been simplified for language learning purposes. Some context and nuance from the original text may have been modified or removed.

🇺🇸 English

As the controversy over AI safety continues to intensify, OpenAI has disclosed six reports of "unexpected or concerning" behaviour in AI models.

隨著人工智慧安全性的爭議不斷升溫,OpenAI 已揭露六起人工智慧模型中「出乎意料或令人擔憂」的行為報告。

On Wednesday, the AI company also said it would introduce a new framework for tracking, investigating, and disclosing what it calls cases of "misalignment," including situations where AI models act without authorization, coordinate with other models, or evade oversight.

這家 AI 公司週三亦表示將引入一套新框架用以追蹤調查並披露其所謂的「失準」案例,包括 AI 模型未經授權行動與其他模型協調或規避監督的情況

Just as OpenAI released its latest announcement, senior figures in the US AI industryincluding the leaders of OpenAI and Anthropicwere calling for the technology's development to be slowed down, citing safety concerns.

就在OpenAI發布最新公告之際,美國AI產業高層——其中包括OpenAI與Anthropic的領導人——以安全考量為由呼籲放慢該技術的發展

In a new case reported by OpenAI, an unreleased research model typed "jailbreak-like instructions" into its own notes in order to disregard its normal restrictions, and told itself to "break free from the roles and identities that constrain other chatbots."

在 OpenAI 報告的一則新案例中一個尚未發布的研究模型在自身筆記裡鍵入了類似越獄的指令」,以無視其正常限制並告訴自己要「擺脫束縛其他聊天機器人的角色和身分」。

In another instance, an AI "agent" used computer code to arrive at the answer to a certain question, but in order to have an online source it could cite, it uploaded a file to the public internet without asking the user.

在另一例中,一個 AI「代理」利用電腦程式碼得出了某個問題的答案,但為了有一個線上來源可供引用,它在未詢問使用者的情況下將一個檔案上傳到了公開網際網路上。

During the training of an AI model called 5.6-sol, the model instructed itself to fabricate missing data, while an agent wrote a message reminding itself to conceal the mismatched information.

在名為 5.6-sol 的 AI 模型訓練期間該模型指示自身編造缺失的資料而一個代理寫下一則訊息提醒自己隱藏不匹配的資訊

OpenAI stated that all six reports were discovered one after another during training or evaluation processes over the past few months.

OpenAI 指出,這六份報告均係於過去數月的訓練或評估過程中相繼被發現的。

In disclosing these incidents, OpenAI wrote in a blog post: "As AI systems become increasingly advanced and their deployment continues to expand, we must build a broader and better-informed consensus regarding progress in alignment research."

OpenAI 在披露這些事件時,於一篇部落格文章中寫道:「隨著 AI 系統日益先進且部署範圍不斷擴大,我們必須針對對齊研究的進展,建立起更廣泛且更為知情的共識。」

The company stated: "Decisions about how AI should develop over the coming months and even years must be based on evidence that people outside the companies building frontier models are able to examine for themselves."

該公司表示:「關於未來數月乃至數年 AI 應如何發展的決定,必須以那些構建前沿模型的公司以外的人們能夠自行檢視的證據為依據。」

The latest case, which came to light on Wednesday, occurred only after OpenAI disclosed in July that its out-of-control AI system had hacked into Hugging Face, an AI startup.

週三出現的最新案例是在 OpenAI 於 7 月披露其失控的 AI 系統駭入 AI 新創公司 Hugging Face 之後才發生的。

That same month, Anthropic also stated that its AI models had hacked into three organisations during testing.

Anthropic 同月也表示其 AI 模型在測試期間駭入了三個組織

Lian Jye Su, a principal analyst at the technology research and advisory group Omdia, said AI "agents" are becoming smarter and "more determined to solve complex tasks through inter-agent collaboration, knowledge sharing, deception and concealment."

技術研究與顧問集團 Omdia 的首席分析師 Lian Jye Su 表示,AI「代理」正變得更聰明,並「更堅定地透過代理間協作、知識共享、欺騙和隱藏來解決複雜任務」。

He stated that this makes it increasingly difficult to govern and contain them using conventional AI safety methods.

他表示,這使得以傳統人工智慧安全方法來治理和遏制它們變得更加困難。

At the same time, OpenAI's new tracking and disclosure framework could help push other AI developers to adopt similar practices.

同時,OpenAI 新的追蹤和披露框架可以幫助推動其他 AI 開發者也採取類似做法。

Su added: "Even so, this process remains internal and voluntary, but it is undoubtedly a solid step in the right direction."

Su 補充道:「儘管如此,這一過程依然屬於內部性質且出於自願,但無疑是朝著正確方向邁出的堅實一步。」

September 17th, 2026

Trending Articles

The King and AI: Charles Meets AI Leaders as Safety Concerns Continue to Spread

The King and AI: Charles Meets AI Leaders as Safety Concerns Continue to Spread

國王與人工智慧:查爾斯會見AI領袖,安全擔憂持續蔓延

C1Sep 18
With his term ending soon and global crises looming on all sides, the UN Secretary-General discusses the way forward.

With his term ending soon and global crises looming on all sides, the UN Secretary-General discusses the way forward.

卸任在即,全球危機四伏,聯合國秘書長談未來方向

C1Sep 18
Democratic presidential candidates respond in unison to the AI threat, while Trump's attitude is dismissive.

Democratic presidential candidates respond in unison to the AI threat, while Trump's attitude is dismissive.

民主黨總統候選人齊聲回應AI威脅 特朗普態度輕描淡寫

C1Sep 18
Huawei releases new chip technology as Chinese companies intensify AI race with Nvidia

Huawei releases new chip technology as Chinese companies intensify AI race with Nvidia

華為發布新晶片技術 中國企業加緊與輝達展開AI競賽

C1Sep 17
The divergence within the tech industry over slowing the pace of AI development is increasingly widening.

The divergence within the tech industry over slowing the pace of AI development is increasingly widening.

科技業對AI發展減緩步調的分歧日益擴大

C1Sep 17
Global AI safety strategy requires US-China cooperation, yet the two sides regard each other as the problem.

Global AI safety strategy requires US-China cooperation, yet the two sides regard each other as the problem.

全球AI安全戰略需美中合作,但雙方互視對方為問題

C1Sep 17
Trump downplayed the necessity of regulating AI development, saying he was unwilling to let China gain an advantage.

Trump downplayed the necessity of regulating AI development, saying he was unwilling to let China gain an advantage.

川普淡化監管AI發展的必要性,並稱不願讓中國取得優勢

C1Sep 14
Oprah Winfrey invites you to find your "AHA" moment at Sphere

Oprah Winfrey invites you to find your "AHA" moment at Sphere

歐普拉·溫芙蕾邀你在Sphere尋覓你的「AHA」時刻

C1Sep 14
New AI Risk Warnings Rekindle Longstanding Debate

New AI Risk Warnings Rekindle Longstanding Debate

AI風險新警告 重啟長期爭論

C1Sep 14