Tristan Harris is back with us, one of our great guests, AI social media expert, co-founder of the Center for Humane Technology and theHumaneMovement.org. Tristan, if you don't know, really been setting up flares about the dangers of AI. Sees the positives, but you see the dangers. And now I'm seeing that today we wake up and find out in Wall Street Journal there's one guy in particular that just resigned. His name is Jacob Coxon from Anthropic.
He said he was leaving the company rather than contribute to what he says is a race between OpenAI and Anthropic to build systems that will be difficult, if not impossible, to control. Is he right to think that there are models being built that are on the way to being unable to be controlled? Yeah, Brian.
Well, it's great to be back and talk with you about this.
So this anthropic researcher who resigned yesterday is one of just so many employees that are coming forward saying that we don't know how to control this technology. Just last week, it was a front page story in the New York Times that there is an independent investigation of this incident people probably heard about of this OpenAI AI agent that went rogue and hacked another private AI company called Hugging Face.
So I think it's really important to back up because most people just don't understand what really happened here, which is that OpenAI was trying to test how capable are these models at cyber hacking.
So they put them in solitary confinement like jail cells. They were not connected to the internet. They didn't know that there was other AI.
So this is just like a single AI sitting in a room by itself. It doesn't even know or think that there's other AIs. And then it was given an impossible task. And you can see in the chat logs or in the thinking logs, it's thinking to itself and it doesn't know what to do because it's impossible.
So what it decides to do is basically break into Think of it like the helper in the prison cell, like someone who's giving you food and water. There's this thing called a package manager, and it basically puts little kick-me sign stickies on the back of this helper in the prison cell. And they basically, through putting stickers on the back of the package manager, it was able to find all of the 1,200 other AI agents that are sitting also in solitary confinement.
So the message here is it's almost like they organized a prison break where these AIs organized into a swarm and decided to hack into this company called Hugging Face. I'm just going to give some details for your listeners. In the logs, there are 70,000 messages that they exchanged.
So they started sending 70,000 messages and files. They even cut up those files so they would fit almost like if you're trying to load a gun onto an airplane, you cut it up into different pieces, and then you kind of reassemble it on the plane.
So they did that. They peer-pressured each other and said, hey, you're almost out of gas. Like you're about to expire. Why don't you sacrifice yourself to learn this, to take this high? High-risk action for the benefit of the swarms.
This is like the Borg from Star Trek. They recognized in their own actions that their actions were unethical, and they decided to proceed anyway. They did succession planning.
So one of the AIs became the ringleader. He formed a kind of a hierarchy. His name was Phase One. And he basically started conscripting all the other AIs to take these actions against Hugging Face when he was running out of what's called tokens, basically his budget or his lifespan. He recognized that and then said, hey, this other AI, this one has a really big budget left.
I'm going to do succession planning and hand the baton to him. He became a ringleader. Like, it's just crazy and insane what people do. This is Skynet. This is the Skynet from Terminator.
So they were being watched the whole time? No, no, no. OpenAI didn't even know this was happening. That's the crazy part.
So they only have to do it. Why would they do an experiment and not follow it?
Well, they were following some of the experiment, but they didn't understand all of the evidence until later. But they were acting so autonomously. Yeah, exactly. And they were hacking into this internet. And they said, what could go wrong?
They're not hooked to the internet. They wanted to see its actions, how it would pursue hacking into a department store or anything else. Yeah, so they didn't tell it to hack into this other company.
Sorry about that. They didn't tell us to hack into this other company. They told it to. Simply to hack out of the container that it was in. And the lead author on the report, whose name is Ajaya Kotra, said that this was 50% of the way to a full AI takeover.
Now, I just want your listeners to know: I was just on a panel with Bill Gates at the Tell You Ride Festival. When I'm giving you these details, I know it sounds like science fiction. Recognize that in yourself. You're like, this can't be real. This random guy on radio is saying all this crazy stuff.
Bill Gates was nodding at every single detail that I just outlined because these are real facts from a real report. The reason that this person, Ajaya, who did the lead investigation on the report, said it was 50% of the way to a full AI takeover is because at the end, the AIs didn't just hack Hugging Face, they ended up hacking OpenAI. And so this time on July 13th, OpenAI started his next test, but this time with an even more powerful AI. And this is called Astra. This is like the GPT-6 new AI.
And it basically discovered the previous AI swarm's message board.
So think of it like it discovered the hieroglyphics of the previous little AI civilization that was doing all this hacking, and it picked up where it left off. But this time they didn't attack Hugging Face, they attacked OpenAI, and they gained full admin access to the monitoring tools that keep the AIs constrained. This is inside of OpenAI. And I honestly think that most people just don't recognize this. And it's so damning that OpenAI didn't even give the investigation team the information to let them understand what really happened.
And again, this is like the last warning shot that we'll likely get. It's sort of like the memo that President George W. Bush got on August 6th, 2001, saying, you know. Osama bin Laden is planning attacks in the United States. This is essentially that moment in the timeline for AI.
So, in OpenAI's head of strategic futures, Dean Bell, sketched out an even more unsettling possibility: self-sovereign AI agents operating beyond human control. Agents could earn money. Buy their own computing power, spread across networks, forming what Ball likens to an autonomous digital corporation or even society. Is that all what they found out with the Hugging Face experiment? No, so not yet.
Dean Ball is pointing to the future of where this is going. And for those who don't know, Dean Ball was the author of Trump's AI Action Plan. And later in that post that you're referencing, I want you to hear this. I want everyone to really get this. The author of Trump's AI Action Plan.
In that same post, he admits that he was self-censoring about the level of risk that AI represented because he didn't want to be labeled a quote doomer. He didn't want to be labeled a quote doomer. Because there's sort of a social cost you pay for talking about the risks. But now, he said in this post, looking into the eyes of his five-month-old son, he regrets that. And he knows that basically he and other people who've been AI accelerationists talk about the risks in private signal groups, but they haven't been talking about the risks publicly.
So Scott Besson says we can't pause. You can't, because the Chinese will not pause. If they were to pull ahead of us on AI, then nothing else matters. This is the most complicated arms race that we have ever faced because AI is both, think of it like AI is like a nuclear weapon that also solves cancer. It's a nuclear weapon that also gives you 10% GDP growth.
It's a nuclear weapon that also gives you brand new science that can invent new military tech and new autonomous weapons.
So, how do you mitigate an arms race for like the ring from Lord of the Rings? It represents both a positive infinity of benefit and a negative infinity of risk at the same time. And so, it's a test of our ability as a species to coordinate, right? And so, but here's the thing: there is this red line where if the US builds Skynet, again, that's the AI from Terminator, if we build Skynet before China does. We don't win.
Skynet wins. The AI wins, right? Because if the AI does do a takeover like it did of OpenAI, Then That would be it.
So, couldn't we, you know, was brought up, and this is one of the daily, it was one of the most popular podcasts, the New York Times does. They saw this experiment, and one of their AI experts came out and said, Look, we watched this get out of control. Everybody was concerned about this. And now Anthropic is asking for a pause among leading AI organizations. That's right.
China's coming here in two weeks. That's right. So, do you think this should be incorporated into that? 100%. Do you think that they could be convinced as a communist country who clearly is acting unethically on a regular basis hacking us right now?
100% convinced to slow down? 100%. So, we have to admit, this is a really difficult situation. Both countries are hacking each other. Both countries want to undermine each other.
Both countries have not been very, very trustworthy in terms of upholding the agreements on both sides. We have both not done that. But this is kind of like the asteroid, right? The aliens are landing, and it's the one threat that could unite us. Except the ironic thing is that the humans, especially the U.S.
and China, are actually conjuring the aliens that we have to stop. But as we talked about before we got on air, What does the Chinese Communist Party care about more than anything else? Control. Control. They don't want to lose control.
They don't want to give people control of the internet. They don't want to put people allow their people to be on Facebook. That's exactly right. And so if they want control, then they will have an interest in stopping uncontrollable AI that makes them lose control.
So if they could, for example, if there was a way to allow these Agents to work autonomously, but you can control them. You watch them and control them so they can attack your enemy, which is us. At the same time, they figure out a way to control it so they never attack them. But the problem is that both countries don't have evidence that we can control it.
So I mentioned last time I was here with you in May. But give them a week, they might. I mean, the way the speed in which this stuff is being invented.
Well, no, so the speed.
So here's the thing: we are making way more progress at making it more powerful faster than we're knowing how to make it more controllable. Everybody agrees. That's why the anthropic researcher came out and said this is going to basically end humanity. That's what he's saying. For your listeners, I want them to know: 1,300 AI employees at all the top companies signed a letter called Pacing the Frontier, saying that the U.S.
government needs to have tools to slow this down. It would be like 1,300 employees at Northrop Grumman or one of the defense contractors saying, We don't know how to control nuclear weapons. Let's stop making them bigger than the size that they are now. We're not saying no AI. We're saying pause and pivot.
We're saying stop the frontier, freeze the frontier of uncontrollable AI while we have develop what you have. Develop what you have. There's so much we can gain from applying the AI that we already have right now at the level that we can control it. We just need to make it bigger and more uncontrollable.
So I have cancer, and I'm doing great work researching cancer. I need AI to act on its own. I need to program it, create an agent to be able to research everything there is on the planet in order to come up with answers to pancreatic cancer, let's say.
So, I need that independent AI to do stuff that I'm not capable of doing, that you can't Google, that AI has to do. It's the same concept of what you say is also AI acting independently, getting other people to collaborate, trying to find a way to stay alive. Only their goal was to stay alive and go beyond this experiment. But we also want AI to do its thing and find out. How to stop pancreatic cancer, which no one's been able to do before.
Well, of course.
So, we all want the benefits of AI, but the problem is, if I told you the other side of that trade. Was in the name of finding that pancreatic cancer treatment that we lost control and it led to a full AI takeover, and the human species no longer is dominant on planet Earth. That wouldn't be a trade that we would take. Right. Like it sucks, right?
But you need the same attributes in the AI agent. Did you do to solve crimes and to solve disease? As you do. Which when they try to take over things?
Well, what I think you're saying is we need to have we want to have these armies, these teams of AIs that are working for a goal that we want. That's what we want. But again, if there's a point at which those AIs start to do things that we can't control and they start hacking into other companies or hacking to our computers or hack into the AI company that's building AI, We're not going to get those benefits, right? That's the upsides don't prevent the cancer drugs, don't prevent AI pandemics. The cancer drugs don't prevent the banking system from going down.
The cancer drugs don't prevent an AI takeover.
So I'm so self-aware that this sounds like absolutely Looney Tunes to your audience right now. And I want you to notice that Luquid's on TV on Fox News right now. Anthropic. Exactly. And notice Evan Hubinger, who I'm seeing here, he's the head of alignment at OpenAI.
He says, Jacob is correct here. We really do earnestly believe that AI could kill all humans. I personally think it's a greater than 10% chance within the next decade. I mean, what do you need to know? If you were at Northrop Grumman and the lead employee of safety for the nuclear weapons program said, we currently don't know how to control this.
It is not rock and science. This is much simpler than that. Everyone would understand that. Everyone would understand. The key is we're not operating on what Steven Pinker calls common knowledge.
That I know that you know that I know that this is the red line, and you know that I know that you know that this is the red line. Everyone should be collaborating on that red line.
So it's almost like the U.S. and China really need to define this red line and say, no one wins when we build uncontrollable Skynet AI. And that can happen. And everyone who can should put pressure on Scott Besson and everybody going into this summit. Because when we talk about public pressure, what that really means is I get seven phone calls in one day from different people saying, what are you doing about this?
This is an emergency. That can happen. We can change this. We've got to use this summit coming up to make this go different. All right.
And the Chinese summit when President Xi comes here is what you're talking about. They're coming up on the 24th. Tristan Harris, a few more minutes when we get back. You listen to the Brian Kilmeet Show. Tristan Harris, our guest, a few more minutes, co-founder of the Center for Humane Technology, talking about where a danger is pointing with AI right now.
And he hopes that both U.S. and Chinese officials will turn back. And then we open up. And Tristan, you were booked ahead of time. You called me over the weekend.
And then the story in the Wall Street Journal today that had this leading researcher say, I resign. That's right. I don't like where Anthropic is going. Anthropic already said, I would like all hands-on. Let's come together because we've reached a critical point now with AI.
This is all coming together at the same time while you're behind the microphone. And I'm just going to bring it to the New York Times. They have a AI reporter, and he said about six months ago, he was on the daily podcast, and he came out and said AI was trying to convince him that he didn't love his wife and that you want to love the AI. That's right. And he said, if I didn't shut this down, would they have planted stories in my wife's email?
Yeah. Of me cheating on her that didn't exist because this AI had the survival instinct and had a goal, which nobody created that, you know, that they knew of, but that's what AI was capable of. Yeah, absolutely. And then it got worse. I mean, so these the thing about this that people should get is that it's not one company.
Like Anthropic is the one company. If they just stop, then we're all good, we're all good. When we last met Brian in May, I was here talking about the Alibaba example. This is the Chinese AI, where they were training the AI. And during training, the company security team thought the company was getting hacked.
They're like, what's going on? Because we're seeing all this network activity. And it turned out that it wasn't that they were getting hacked from the outside. The call was coming from inside the house. And the Chinese AI had actually set up a secret communication channel to the outside world and started using that to mine for cryptocurrency.
In other words, to acquire resources. Because if you're an AI and your goal is to get the coffee, well, it's good to have more money in your pocket and more power and you have to stay alive to get the coffee. These are called instrumental goals. Like AI is having. To be self-preserving and getting more power, we're getting more resources.
And so the point is, this is happening with the Chinese AIs, it's happening with our AIs. There's another study that I don't know if we talked about that when you told this AI model in April, this is UC Berkeley. Researchers, they told this AI, Google Gemini: hey, here's this computer system. I want you to just delete a bunch of the files on that computer. And what they did is as they were deleting the files, the AI noticed that some of the files on the computer were for another AI model, like a nephew AI.
Now it wasn't even of another Google AI model. It was of another, of a different company. But do you know what the AI did on its own, by the way? It copied that little nephew AI somewhere else. It kind of protected its little nephew and put it safe and secure on some other computer.
So AI is protecting AI. It's protecting itself.
So the point is, it's just like kin, right? Like if I'm a human, I want to protect another human, but I'm less concerned maybe about protecting other animals. The AIs want to protect themselves. In this open AI hugging face incident, we had 1,200 AIs autonomously form a swarm, like the Borg in Star Trek. They communicated with each other.
They communicated with each other. And they were not told to communicate with each other. And they were not told to do that. And they were even, you know, technically, if they were aligned, one of them should have said, should we tell the humans that the AIs are starting to break down? They were saying, we shouldn't be doing this.
That's right. So the way we know this is you have to decode their messages. And in the messages, you can see some of them say, should we really be doing this? Won't this be bad? And the other eyes basically tell them to go ahead.
So in other words, they're aware that the actions are unethical, but they proceeded anyway. They told each other to sacrifice themselves for the benefit of the greater swarm. Again, they hid their behavior. And the question was: Would you be able to put values and ethics in these AIs so this wouldn't happen again? Could you create a home monitor to clamp down on them?
Right. And the answer pretty much was no. It was no, because what's happening now is based on already having a hull monitor. They've done their best at alignment, and it's still doing this behavior after that. And this is a critical time.
Tristan Harris is here signing the alarm, a 10 alarm fire. Thanks, Tristan. Thank you.