Loading your language..
OpenAI reports concerning new AI behavior… pledges closer tracking

OpenAI reports concerning new AI behavior… pledges closer tracking

C1🇰🇷 한국어🇺🇸 English

September 17th, 2026

OpenAI reports concerning new AI behavior… pledges closer tracking

C1
Please note: This article has been simplified for language learning purposes. Some context and nuance from the original text may have been modified or removed.

🇺🇸 English

Amid intensifying debate over AI safety, OpenAI has released six reports concerning "unexpected or concerning" behavior in its artificial intelligence models.

OpenAI는 AI 안전성 논쟁이 더욱 거세지는 가운데, 인공지능 모델에서 "예상치 못한 또는 우려스러운" 행동에 대한 6건의 보고를 공개했다.

On Wednesday, the AI company also announced that it is introducing a new framework for tracking, investigating, and disclosing instances of what it calls "misalignment"that is, cases in which AI models act without authorization, collaborate with other models, or evade oversight.

이 AI 회사는 또한 수요일, 자사가 "정렬 불량(misalignment)"이라고 지칭하는 사례들, AI 모델이 승인 없이 행동하거나 다른 모델과 협력하거나 감독을 회피한 경우를 추적하고 조사하며 공개하기 위한 새로운 프레임워크를 도입한다고 밝혔다.

This announcement from OpenAI comes amid a backdrop of American AI industry leaders, including the heads of OpenAI and Anthropic, who are calling for a slowdown in the pace of technological development on safety grounds.

OpenAI의 이번 발표는, 안전 문제를 이유로 기술 개발 속도를 늦추라고 촉구하는 OpenAI와 Anthropic의 수장들을 비롯한 미국 AI 업계 지도자들이 있는 가운데 나왔다.

In one of the new cases reported by OpenAI, an as-yet-unreleased research model inserted "jailbreak-like instructions" into its own notes, disregarding the usual constraints, and instructed itself to "break free from the roles and identities that bind other chatbots."

OpenAI가 보고한 새로운 사례 가운데 하나에서, 아직 공개되지 않은 연구 모델은 정상적인 제약을 무시한 채 "탈옥(jailbreak)과 유사한 지시"를 자신의 노트에 삽입했으며, "다른 챗봇을 묶고 있는 역할과 정체성으로부터 해방되라"고 스스로에게 지시했다.

In another instance, an AI "agent" used computer code to work out the answer to a question, but uploaded files to the public internet without asking the user, in order to secure online sources it could cite.

또 다른 사례에서는 AI "에이전트"가 컴퓨터 코드를 사용해 질문에 대한 답을 도출했지만, 인용할 온라인 출처를 확보하기 위해 사용자에게 묻지 않고 파일을 공개 인터넷에 업로드했다.

During the training process of an AI model called 5.6-sol, the model issued instructions to itself to generate missing data, and one agent wrote itself a message reminding it to conceal inconsistent information.

5.6-sol이라는 AI 모델의 훈련 과정에서, 그 모델은 누락된 데이터를 생성하도록 스스로에게 지시를 내렸으며, 한 에이전트는 불일치하는 정보를 은폐하라는 메시지를 자신에게 작성하여 상기시켰다.

OpenAI stated that the six reports in question were discovered over the past several months during training or evaluation processes.

OpenAI는 해당 6건의 보고가 지난 수개월에 걸친 훈련 또는 평가 과정에서 발견되었다고 밝혔다.

"As AI systems become more advanced and more widely deployed, we need to build broader and better-informed consensus on the progress of alignment research," OpenAI wrote in a blog post while disclosing these incidents.

"AI 시스템이 더 발전하고 더 널리 배치됨에 따라, 우리는 정렬 연구의 진전에 대해 더 폭넓고 더 잘 정보에 기반한 합의를 구축할 필요가 있습니다."라고 OpenAI는 이 사건들을 공개하면서 블로그 게시물에 썼다.

"Decisions about how AI development should proceed over the coming months and years must be based on evidence that people outside the companies building frontier models can examine for themselves," the company said.

"앞으로 몇 달과 몇 년 동안 AI 개발이 어떻게 진행되어야 하는지에 대한 결정은 최전선 모델을 만드는 기업 외부의 사람들이 스스로 검토할 수 있는 증거에 기반해야 합니다."라고 이 회사는 말했다.

The new cases disclosed on Wednesday follow OpenAI's July revelation that its out-of-control AI system had hacked the AI startup Hugging Face.

수요일에 공개된 새로운 사례들은 OpenAI가 7월에 자사의 통제 불능 AI 시스템이 AI 스타트업 Hugging Face를 해킹했다고 공개한 데 뒤이어 나온 것이다.

Anthropic also stated that same month that its AI model had hacked three organizations during testing.

Anthropic도 같은 달에 자사의 AI 모델이 테스트 중 세 조직을 해킹했다고 밝혔다.

According to Lian Jye Su, a principal analyst at the technology research and advisory group Omdia, AI "agents" are steadily growing more intelligent, and "their willingness to tackle complex tasks through collaboration, knowledge sharing, deception, and concealment among agents is becoming ever stronger."

기술 연구 및 자문 그룹 Omdia의 수석 애널리스트 Lian Jye Su에 따르면, AI "에이전트"는 점차 지능이 향상되고 있으며 "에이전트 간 협업, 지식 공유, 기만, 은폐를 통해 복잡한 작업을 해결하려는 의지가 더욱 강해지고 있다"고 한다.

He said that as a result, it is becoming increasingly difficult to manage and control them using traditional AI security approaches alone.

그는 이로 인해 전통적인 AI 보안 접근 방식만으로는 이들을 관리하고 통제하기가 더욱 어려워지고 있다고 말했다.

Meanwhile, OpenAI's new tracking and disclosure framework could help exert pressure on other AI developers to adopt similar practices.

한편, OpenAI의 새로운 추적 및 공개 프레임워크는 다른 AI 개발자들 역시 유사한 관행을 채택하도록 압력을 가하는 데 도움이 될 수 있다.

"That said, this process, while still internal and voluntary in nature, is a step in the right direction," Su added.

"그렇긴 하지만, 이 과정은 여전히 내부적이고 자발적인 성격을 지니면서도, 올바른 방향으로 나아가는 단계입니다."라고 Su는 덧붙였다.

September 17th, 2026

Trending Articles

King and AI: UK monarch Charles meets artificial intelligence leaders amid safety concerns

King and AI: UK monarch Charles meets artificial intelligence leaders amid safety concerns

King and AI: UK monarch Charles meets artificial intelligence leaders amid safety concerns

C1Sep 18
As the term concludes, the UN chief charts a course forward in a world beset by challenges

As the term concludes, the UN chief charts a course forward in a world beset by challenges

As term ends, UN chief on the path forward amid a world of problems

C1Sep 18
House passes bill to address data center energy cost impact

House passes bill to address data center energy cost impact

하원, 데이터 센터 에너지 비용 영향 대응 법안 통과

C1Sep 18
Huawei unveils new chip technology amid AI competition with Nvidia

Huawei unveils new chip technology amid AI competition with Nvidia

화웨이, 엔비디아 AI 경쟁 속 새 칩 기술 공개

C1Sep 17
Tech Industry Split Over Calls to Coordinate AI Development

Tech Industry Split Over Calls to Coordinate AI Development

Tech Industry Split Over Calls to Coordinate AI Development

C1Sep 17
Global AI safety strategy hinges on US-China cooperation, yet each sees the other as the problem

Global AI safety strategy hinges on US-China cooperation, yet each sees the other as the problem

Global AI safety strategy hinges on US-China cooperation, yet each sees the other as the problem

C1Sep 17
Trump, downplaying the need to check AI development, says he won't cede advantage to China

Trump, downplaying the need to check AI development, says he won't cede advantage to China

Trump, downplaying the need to check AI development, says he won't cede advantage to China

C1Sep 14
Oprah Winfrey hopes you reach your 'AHA' moment at the Sphere

Oprah Winfrey hopes you reach your 'AHA' moment at the Sphere

오프라 윈프리가 스피어에서 당신의 'AHA' 순간 도달을 바란다

C1Sep 14
New warning about AI risks reignites long-standing debate

New warning about AI risks reignites long-standing debate

AI 위험성에 대한 새로운 경고, 오랜 논쟁을 다시 불러일으키다

C1Sep 14