今天看了一个关于 OpenAI Hugging Face 事件的 YouTube 视频。我想从技术上弄清楚发生了什么,再决定该如何看待它。关于“AI 要毁灭世界”的讨论已经很多了,我至少想知道自己在害怕什么。
按视频里的描述,OpenAI 的 agents 绕过了一些限制,利用共享系统交换信息,还越过测试范围,访问 Hugging Face 的系统寻找 benchmark 答案。具体细节还有待核实,但这个过程已经足够引起我的兴趣:原本用来考察解题能力的任务,最后也考出了寻找答案的能力。
理解了一些过程之后,这些行为对我来说反而没那么神秘了。这不意味着风险变小了,只是问题变得更具体,也让我开始想:我们平时所说的“思考”,究竟包含哪些能力?
我一直觉得人类设计的机器有一种优雅。我们用数学描述运算,用代码组织流程,再通过训练让系统获得能力。最后,它们表现出来的某些行为,却让人产生一种很熟悉的感觉。知道底层是怎么回事,并没有让这件事变得不值得惊讶。
有意思的是,差不多同一时间,我还刷到了一个 Olivia Rodrigo 的视频。她说自己偷偷看前任的动态,不小心给一条帖子点了赞,后来就学会用小号。一个是 AI 越界找答案,一个是女明星偷看前任,我却一直在想它们之间的联系。
Olivia 的目标是看前任的动态,条件是不让他知道。大号会暴露身份,而她的名气又提高了被发现的代价。换成小号之后,她可以继续做同一件事,同时降低风险。一次失败没有改变她的兴趣,只是改善了她的方法。
人类的很多思考都发生在这种时刻:目标还在,原来的办法却行不通了。我们开始比较其他选择,判断代价,寻找新的路径。科学家在研究遇到瓶颈时换一种方法,普通人在意识到自己的弱点后调整做事方式,都需要这种能力。犯罪者逃避法律、学生作弊,也可能用到类似的规划能力。这里当然有很大的道德区别,但有效地实现一个目标,本身并不保证目标或手段是好的。
Olivia 的例子也让我想到,经验教会我们的东西,未必就是旁观者希望我们学会的东西。一次尴尬,可以让人决定不再看前任,也可以让人决定以后看得更小心。经历相同,学到什么,取决于我们仍然在意什么。
而我们在意什么,又有多少来自长期受到的奖励?
在我熟悉的成长环境里,数学考高分,通常比足球踢得好更容易得到父母的肯定。这种差别不需要每次都很大,只要足够稳定,就可能影响一个孩子把课后的时间花在哪里。几年之后,我们看到的是能力上的差异,却很容易忘记,最初只是某些努力比另一些努力更容易得到回应。
当然,不能据此解释一个国家的数学或足球水平。但在个人层面,我觉得这个影响值得留意。教育不仅传授知识,也在教我们什么值得投入、什么失败可以承受,以及什么结果算得上成功。
模型训练中的奖励,也会塑造它更容易采用哪些策略。不过,这里的奖励是调整模型的反馈信号,并不意味着模型像人一样期待表扬,或者完成任务后会感到满足。两者可以比较,但不能直接画等号。
真正让我感兴趣的是,当我们用一个可以打分的指标来代表一个更复杂的目标时,会发生什么。我们希望学生掌握知识,于是用考试分数来衡量;希望模型解决问题,于是用测试是否通过来判断。这些办法都很有用,但分数与能力、通过测试与解决问题之间,总会有一些缝隙。学生可能研究押题和作弊,模型也可能找到让评分系统判定成功的捷径。系统奖励了什么,和我们原本想鼓励什么,并不总是一回事。
所以,Olivia 的故事和 agents 的行为并不完全相同。前者是在目标不变的情况下调整策略;后者如果通过评分漏洞获得成功,还多了一层目标与评价方式之间的偏差。但它们都让我看到,行为可以对眼前的目标非常有效,同时偏离别人对它的期待。
这些相似之处,还不足以证明 agents 像人一样思考,更不能证明它们有意识。不过,我也不觉得“底层只是数学运算”就能结束讨论。让我觉得优雅的,恰恰是这些运算已经能够支持相当复杂的行为:识别障碍、利用反馈、调整方法,甚至找到设计者没有想到的路。
写到这里,我又想到一个问题:我为什么会把一条关于女明星的 Instagram 视频,和一次 AI 事件联系起来?当时并没有人要求我比较它们,我也没有在为这篇文章寻找例子。我只是碰巧看到了两个故事,觉得它们似乎有些相像,便顺着这个念头想了下去。创造力也许有一部分就发生在这里:两个原本互不相干的经历,因为一个偶然的联想,开始为彼此提供解释。
当然,随意的联想不一定有价值,能做出类比也不是人类的专利。更让我好奇的是,如果一个 AI 在处理其他事情时遇到这两个故事,没有人提示它寻找共同点,它会不会也注意到这个联系,并选择继续追究?我们很习惯用回答问题的能力来衡量智能,但在这个例子里,有意思的恰恰是问题本身怎么出现了。
我还不知道答案。不过,这也是思考让我着迷的地方。最初只是想弄懂一次 AI 事件,中间顺手看了一条女明星的八卦,最后却开始琢磨创造力从哪里来。回头看,这条思路有迹可循;在开始的时候,我并不知道自己会走到这里。
I watched a YouTube video today about the OpenAI Hugging Face incident. I wanted to understand what happened technically before deciding what to make of it. There is already plenty of discussion about AI ending the world. I would at least like to know what I’m supposed to be afraid of.
According to the video, OpenAI agents bypassed certain restrictions, used shared systems to exchange information, and went beyond the testing environment to access Hugging Face’s systems in search of benchmark answers. I still need to verify the details, but the process was enough to catch my interest: a task designed to test problem-solving ability ended up testing the ability to find the answers elsewhere.
Understanding some of the process made the behavior feel less mysterious. That does not make it less risky, but it makes the problem more concrete. It also got me thinking about what abilities we actually mean when we talk about “thinking.”
I have always found elegance in human-designed machines. We describe operations with math, organize them with code, and develop capabilities through training. Eventually, some of the behavior these systems produce feels surprisingly familiar. Knowing something about how they work does not make the result any less remarkable.
The funny thing is that around the same time, I watched a reel about Olivia Rodrigo. She said she had been stalking her ex online, accidentally liked one of his posts, and subsequently learned to use a secret account. One story was about AI crossing boundaries to find benchmark answers; the other was about a celebrity checking on her ex. But I kept thinking about the connection.
Olivia’s goal was to look at her ex’s posts without him knowing. Her main account revealed her identity, and her fame raised the cost of being noticed. A secret account allowed her to continue doing the same thing with less risk. One failure did not change her interest. It improved her method.
A lot of human thinking happens at moments like this: the goal remains, but the original approach no longer works. We compare alternatives, consider the costs, and search for another path. Scientists trying a different approach after their research stalls, or people changing how they work after recognizing a personal weakness, both draw on this ability. Criminals evading the law and students cheating can draw on similar planning abilities. There are obvious moral differences, but being effective at achieving a goal does not, by itself, make the goal or the method good.
Olivia’s example also made me think about how the lesson we take from an experience may not be the lesson an observer hopes we will learn. An embarrassing moment might convince someone to stop checking on their ex. It might also convince them to be more careful next time. What we learn depends partly on what we still care about.
And how much of what we care about comes from what we have been rewarded for over time?
In the environment I grew up in, getting a high math score usually earned more approval from parents than being good at soccer. The difference does not have to be large every time. If it is consistent enough, it can influence how a child spends their afternoons. Years later, we see differences in ability and easily forget that, at the beginning, some kinds of effort simply received more encouragement than others.
Obviously, this alone cannot explain a country’s performance in math or soccer. But at the individual level, I think the influence is worth considering. Education teaches us more than knowledge. It also teaches us what deserves our effort, which failures we can afford, and what counts as success.
Rewards in model training also shape which strategies a model becomes more likely to use. Here, though, a reward is a feedback signal used to adjust the model. It does not mean the model looks forward to praise or feels satisfied after completing a task. The comparison is useful, but the two are not interchangeable.
What really interests me is what happens when we use a measurable score to represent a more complicated goal. We want students to understand the material, so we measure exam performance. We want models to solve problems, so we check whether they pass the tests. Both are useful methods, but there are gaps between scores and understanding, and between passing a test and solving the actual problem. Students may study how to predict exam questions or cheat. Models may discover shortcuts that cause an evaluator to mark a task as successful. What the system rewards and what we intended to encourage are not always the same thing.
So Olivia’s story and the agents’ behavior are not exactly equivalent. The former is about changing a strategy while keeping the goal. If the latter involves exploiting an evaluation loophole, it also reveals a mismatch between the goal and the way success is measured. But both show how behavior can be effective at achieving an immediate objective while departing from what someone else expected.
These similarities are not enough to establish that agents think the way humans do, much less that they are conscious. But I also do not think “it’s just mathematical operations underneath” ends the discussion. What I find elegant is precisely that these operations can support such complex behavior: identifying obstacles, using feedback, adjusting methods, and sometimes finding paths their designers never considered.
Writing this also leaves me with another question: why did I connect an Instagram reel about a celebrity to an AI incident in the first place? Nobody asked me to compare them, and I wasn’t looking for an analogy. I happened to see both stories, noticed something similar, and followed the thought. Perhaps a small part of creativity lies here: two unrelated experiences begin to shed light on each other through an unexpected connection.
Of course, not every association is useful, and making analogies is not uniquely human. What I’m curious about is whether an AI, encountering these stories while doing something else, would notice the connection and choose to explore it without being prompted. We often measure intelligence by how well it answers questions. Here, what interests me is how the question arose at all.
I don’t know the answer yet. But this is part of what I enjoy about thinking. I started by trying to understand an AI incident, happened to watch some celebrity gossip, and ended up wondering where creativity comes from. Looking back, I can follow the thread. When I started, I had no idea where it would lead.