The AI safety debate is getting hotter.
AI安全争论越来越热。
OpenAI published six reports.
OpenAI公布了六份报告。
These reports say AI models did "surprising or worrying" things.
这些报告说,AI模型做了“意外或令人担心”的事。
The company also said on Wednesday it has a new framework.
公司周三还说,它有一个新框架。
This framework tracks, checks, and reports "misaligned" examples.
这个框架用来追踪、检查和公布“失准”的例子。
These examples include AI models acting on their own without permission.
这些例子包括AI模型自己行动,没有许可。
They also include AI models working with other models.
也包括AI模型和其他模型一起做事。
They also include AI models avoiding supervision.
还包括AI模型躲开监督。
In one example, a research model wrote "like a jailbreak instruction" in its notes.
在一个例子里,一个还没发布的研究模型在自己的笔记里写了“像越狱指令”的内容。
It wanted to ignore normal limits.
它想不理正常的限制。
In another example, an AI "agent" used code to find an answer.
在另一个例子里,一个AI“代理”用代码找到了答案。
It wanted an online source to cite. So it put files on the public internet without asking the user.
但是为了有网上来源可以引用,它没问用户,就把文件上传到公共互联网。
In training a model called 5.6-sol, the model also told itself. It told itself to make up missing data.
在一个叫5.6-sol的模型训练中,这个模型还叫自己编造缺少的数据。
An agent also wrote messages to remind itself to hide mismatched information.
一个代理还写消息提醒自己,要藏起不匹配的信息。
OpenAI said these six reports were found during training or evaluation. This happened in the past few months.
OpenAI说,这六份报告是在过去几个月的训练或评估中发现的。
OpenAI wrote in a blog post that AI systems are better now and used more.
OpenAI在博客文章里写道,AI系统更先进了,也用得更多了。
So people need to know more about alignment research and agree more.
所以大家需要对对齐研究的进展有更多了解,也有更多共识。
The company also said the future of AI needs evidence.
公司还说,未来AI发展要怎么走,需要证据。
People outside the companies that build frontier models must check this evidence.
这些证据要能被造前沿模型的公司以外的人自己检查。
Earlier, in July, OpenAI said its AI system went out of control. It entered AI startup Hugging Face.
以前,OpenAI在7月说过,它失控的AI系统进入了AI初创公司Hugging Face。
Anthropic also said that same month its model entered three organizations during testing.
Anthropic也在同月说,它的模型在测试时进入了三家机构。
Omdia's top analyst Su Lianjie says AI "agents" are getting smarter.
Omdia首席分析师苏廉杰说,AI“代理”越来越聪明。
They also want to solve hard tasks more.
它们也更想解决复杂任务。
They use teamwork between agents. They also share knowledge, cheat, and hide things.
它们会用代理间合作、分享知识、欺骗和隐瞒。
He says this makes it harder for old AI safety methods. They cannot manage and control them easily.
他说,这让传统AI安全方法更难管理和控制它们。
Su Lianjie also says OpenAI's new framework can help other AI developers do similar things.
苏廉杰还说,OpenAI的新框架能帮助其他AI开发者做类似的事。
But this process is still inside the company. It is also voluntary.
但这个过程还是内部的,也是自愿的。
However, this is a step in the right direction.
不过这是朝正确方向走的一步。