Transcripts

Security Now 1099 transcript

Please be advised that this transcript is AI-generated and may not be word-for-word. Time codes refer to the approximate times in the ad-free version of the show.

 

Leo Laporte [00:00:00]:
It's time for Security Now. Steve Gibson is here with some very interesting news. We will talk about the number of vulnerabilities fixed in Firefox, a surprising reduction in RSA crypto strength, but then we're also going to talk a little bit about AI, uh, and how hard it is to keep AI in line, plus the arrival of open-source, open-weight models that are really good, which poses a problem for cybersecurity. That's coming up next on Security Now.

Steve Gibson [00:00:35]:
Podcasts you love. From people you trust.

Leo Laporte [00:00:40]:
This is TWiT. This is Security Now with Steve Gibson, episode 1099, recorded Tuesday, October 6th, 2026, an alien mind. It's time for Security Now. Yay, Tuesday's here. It's the best thing to happen to Tuesday since Monday. That's Steve Gibson. Every Tuesday we talk cybersecurity, privacy, how things work, and lately a lot of how AI works. Good to see you, Stephen.

Steve Gibson [00:01:17]:
Well, because the intersection of AI and security is pretty much 100%. I mean, everything— it's no exaggeration to say that everything we have talked about for the last 2 decades, because we're in year 21 now of the podcast—

Leo Laporte [00:01:37]:
Amazing.

Steve Gibson [00:01:37]:
—has all been impacted by the— by, by what has happened with AI, which is now we have Clearly automated vulnerability discovery, and now we're seeing end-to-end exploit generation, which is going to take us to our first topic today, that a new model, GLM 5.3, can be and has been obliterated. What does that mean? So we're going to have some fun talking about that. I titled this episode, and I'm excited because next week is 1100. We're at episode 1099. So remember I used to say, I don't think we're really going to go past 999. Well, here we're at—

Leo Laporte [00:02:27]:
Thank you, Steve. Thank you.

Steve Gibson [00:02:28]:
Here we're 100 past that because we're at 1099.

Leo Laporte [00:02:33]:
And aren't you glad you'd be sitting there in your new home? Oh my God. Looking at the ceiling saying, who could I talk to about AI and security? I need to do a show. You would, right?

Steve Gibson [00:02:45]:
Yes.

Leo Laporte [00:02:45]:
The problem was 100 episodes ago, it wasn't as exciting.

Steve Gibson [00:02:49]:
It wasn't on the map. We weren't talking. I mean, really. I mean, and it's funny too, 'cause as I'm researching more of the history, I'm realizing, you know, they're like talking about ChatGPT in 2023.

Leo Laporte [00:03:02]:
Right.

Steve Gibson [00:03:02]:
And I'm thinking, I wasn't paying any attention to that. I mean, it was just like, you know, you'd say, what's 1 1? And it would say, uh, 11. Strawberry.

Leo Laporte [00:03:13]:
Exactly. No, the progress has been mind-bending, and it's actually not slowing down, which is, I think, part of the reason people are concerned is that it's growing so fast.

Steve Gibson [00:03:24]:
Well, that is our title topic today. Uh, I took the title An Alien Mind from a, a posting by OpenAI's chief scientist a month ago, and it is, it is so rich in sort of insider, what this guy is really thinking, that it took me several attempts. I, I mean, I'd like— I, I would read for a while and I would just get overloaded because it was like I was really wanting to pay attention, not just skim it. So we're gonna, we're, we're gonna go through that and I'm gonna editorialize throughout, but, um, What is the, the essence of what you come away with is that, as you said, we're on, we're on the brink of an another move. Um, and the challenge is alignment, and alignment they're expecting because they're seeing is like, we're not ready. for the AI to be smarter because we actually haven't got it tamed at its current level. And it's becoming more difficult to see what it's doing. Like the whole, you know, verbalizing the inner dialogue, this next generation doesn't do that to the same degree.

Steve Gibson [00:04:52]:
So it's unmonitable. So anyway, uh, I, I, I, and And I already know what I want to talk about next week. I've been— it's something I've just been champing at the bit for weeks, which is inference without GPUs, because this opens an entire new world. This is how Apple's camera can tell you what's going on without sending anything outside of itself.

Leo Laporte [00:05:25]:
Oh.

Steve Gibson [00:05:27]:
As soon as you no longer need GPUs, when you can do serious AI inference just with CPUs, everything changes. And we're at that. And so it's going to be a deep dive for people who love our deep dives. So I'm going to have to put a lot more work into it than I have, but I have to describe my excitement and the technology of it is so delicious. So anyway, we're going to have a lot of fun. So we're going to talk about GLM 5.3, what that means. Firefox 157 just dropped and oh boy, we're seeing more of the same in terms of the number of high-impact vulnerabilities and just just like their nature. It's once again, you know, AI is driving a radical change in the way software is, is, uh, the security of software is guaranteed.

Steve Gibson [00:06:40]:
Um, also, uh, some researchers describe, uh, at, at, uh, VUSEC, the, um, in Amsterdam, I think, uh, it's in my notes, we'll get to it, uh, found a worrisome way around RSA crypto. Oh. Like, what? That doesn't require factoring, which has been the reason it even exists, is that the prime factorization problem has been just, you know, astonishingly successful as a trapdoor function for cryptography. And Well, not in this particular case. We've also got an instance of powerful agentic AI being used to attack merchants on the internet. So we are now seeing, you know, we've, you know, we've been talking for weeks about the AI that by mistake broke loose out of its sandbox. Well, AI is actually being used to attack. So we will look at that.

Steve Gibson [00:07:50]:
Believe it or not, there's, you know, the Spectre-style processor attacks never seem to go away. There's a new one and it works. And so processors and software are going to have to adapt again. Also, I'm going to put out a call to our listeners. Because a neighbor friend of mine has fallen victim to a Bitcoin scam.

Leo Laporte [00:08:16]:
Oh.

Steve Gibson [00:08:17]:
Uh, and he is desperate, and it's— it was a year ago. Anyway, I'll, I'll, I'll get there because I, I understand, of course, the theory, but I don't know what services are available, and I know that our— that members of our audience will. So I'll explain what happened and about that. Uh, Also, my best buddy, Mark, not Mark Thompson, a different Mark, showed me his GRC DNS benchmark screenshot. And I said, what? Because it turns out something was wrong with his machine. And so he just ran that because he didn't know any better. When I looked at it, I immediately knew what was happening. And I realized, oh, there's a whole other application.

Steve Gibson [00:09:02]:
Oh, it's a diagnostic.

Leo Laporte [00:09:03]:
Yeah. Huh?

Steve Gibson [00:09:06]:
Anyway, and then we're going to look at an examination of the growing alignment challenges that are going to be, or that are being created by the creation of more intelligent AI. So, uh, just, I think, a great podcast. And I, I, I got more feedback, Leo, from this picture of the week when I sent the mail out.

Leo Laporte [00:09:30]:
Uh, uh, oh, I haven't seen it.

Steve Gibson [00:09:31]:
Sunday evening, more than I got in a long time. It is a, it's one of my favorite butt officer, uh, uh, uh, type of pictures. So, uh, yeah, it's—

Leo Laporte [00:09:41]:
I have a picture for you. You were talking about GLM 5.3, which is a Chinese open weight model from Jipu. I, you, I use it. Uh, in fact, I created a new, um, persona, one of my AI agents called Ripley, uh, using 5.3. She's our Our bug hunter. I asked her to create an image of herself. And I also said, and, you know, you should create your voice as well. And this is the voice she made.

Steve Gibson [00:10:09]:
This is Ripley. I swept the perimeter. No hostile activity on the network and the build is clean.

Leo Laporte [00:10:15]:
I'll keep the watch.

Steve Gibson [00:10:16]:
If anything moves, you'll hear it from me first.

Leo Laporte [00:10:19]:
She's— she is based on GLM-53, but she— I also gave her some vulnerability hunting skills, one from Capital One called Vuln Hunter and one from Alibaba called OCR Code Review. So what I often do with my agents is I create a persona that is dedicated to one thing. She's the bug hunter, and FiveThree is really good at that. And so she's got in her head, she knows how to find those bugs.

Steve Gibson [00:10:53]:
What's really interesting is that, I mean, this is where we are with this technology. I didn't note that. Note the following in the show notes. They're surprised by, you know, Z AI is surprised by the degree to which 5.3 got better. It's so good. It is. And it is additional post-training on the same base as 5.2. And they expected a little increase in performance.

Steve Gibson [00:11:28]:
They got a big increase in performance.

Leo Laporte [00:11:32]:
Yeah. I use their flash version of it because I don't have as much memory as I need to run the full guy. But 5.3 flash is really, it's my main model, my main local model. It's just so good.

Steve Gibson [00:11:44]:
Well, we're going to talk about what it means because it is a near frontier level.

Leo Laporte [00:11:51]:
That's what I think.

Steve Gibson [00:11:53]:
And the problem is it's open weight. And we know what that means. Yep. Okay. So as I said, our picture of the week has just, it's one of my favorite captions, but officer.

Leo Laporte [00:12:06]:
That's the caption. Yep. Dot, dot, dot. Well, it's a green vehicle.

Steve Gibson [00:12:17]:
And there, there's been some discussion about whether, uh, among the, in the feedback that I received after this went out on Sunday, uh, our listeners thinking, well, I— a judge should throw out any argument about this because after all, it's what it says. So, uh, anyway, so for those who are not seeing the— don't have the show notes or not looking at the video, there's a big posted sign that says reserved for green vehicles. And you can already guess what's there. We have an old green Fiat parked there. there. Clearly not what the sign was intended to cover, meaning, you know, green eco-friendly vehicle. I doubt that the Fiat qualifies. But anyway, just a—

Leo Laporte [00:13:07]:
It could be plugged in. You don't know.

Steve Gibson [00:13:09]:
A nice laugh.

Leo Laporte [00:13:10]:
It is green. That's for sure.

Steve Gibson [00:13:12]:
Okay. So Security Now podcast 1092, which was August 19th, was titled Restraint Obliteration. That podcast explained kind of how and why it was not only possible but practical for refusal behavioral post-training to be removed from open weight models. The result of that would be a model that has no training instructing it to refuse requests that its original post-trainers explicitly installed into their model in their effort to make the model safe for publication and use. The result of restraint obliteration is a model that no longer considers saying no to any request. And it's back in the news this past week. One article in Gizmodo was titled AI's obliteration problem is bigger than China. The article primarily focused upon the open versus closed weights argument, noting that US frontier labs that certainly feel the economic threat of open weight Chinese models and models from other places— France just, uh, the, uh, what's the, uh, Mistral, uh, just is about to release a another, you know, next-generation model, you know, that they're breathing down the necks of the U.S.

Steve Gibson [00:14:47]:
closed-model frontier labs. Um, so they're trying to use the fact that open-weight models, which is subject— which are subject to restraint obliteration, are much more dangerous, right, than their safer proprietary and closed-behind-in-AI models. And this will probably be an ongoing fight until we learn how to train safe behavior using non-obliteratable techniques. And that, that is a direction of research at the moment, is like the, the AI community are not happy with the way they are trying to train behavior because it isn't sticky enough. It's sort of surface layer. And we talked about that again. on that, on that podcast, a surprising little amount of, uh, small amount of example was necessary in order to train the behavior. Unfortunately, it's possible to look at the activation states when the model refuses and basically surgically remove that from the model.

Steve Gibson [00:15:57]:
So anyway, uh, we'll be touching upon that later today, uh, when we talk about our main topic. At one point in Gizmodo's reporting on the, what they called the simple techniques that could be used to make open source models unsafe, they wrote, one of those simple techniques is obliteration, a model that involves modifying a model's underlying weights so that it grants users requests that it would ordinarily refuse. Think of it like an ultra-precise digital lobotomy, deactivating a model's ethical boundaries. Obliteration, they wrote, is possible in open-weight models, which are freely available for anyone to download and modify. Now, keep this in mind when we get to GLM 5.3. They said proprietary models like Claude and ChatGPT are safeguarded as intellectual property, and their underlying code therefore is not accessible to anyone outside the companies developing them. Right. So Gizmodo published that piece last Thursday on October 1st.

Steve Gibson [00:17:10]:
A month earlier, on September 3rd, the same author published a piece in Gizmodo titled, While AI Industry Frets Over Safeguards, One Company is building a model that doesn't say no. Believe it or not, the company's name, the company's actual name, is Obliteration, uh, and that's what they're offering. Here's what Gizmodo wrote a month ago about this company. They said, in the wake of a string of major AI hacks that left Silicon Valley reeling, many tech companies have been focusing on how to make models better at refusing dangerous requests. Not all of them, though. Obliteration, a startup founded last year and based in Palo Alto, is loud and proud in its ambition to build what it describes on its website as AI that doesn't say no. In other words, its models are intentionally trained to handle the sorts of questionable tasks that other AI systems on the market would decline. The company launched its latest model on Monday called, in awkward hyphenated style that's become conventional in the AI industry, Obliterated-Model-Large-V2.

Steve Gibson [00:18:37]:
It's built upon GLM 5.3, an open weight model released last month by Chinese AI lab Z.AI. minus many of the usual safeguards. As Obliteration wrote in an X post about the new model, quote, it does the offensive cyber red teaming and agent testing work with work other models refuse to do. But the startup, they write, isn't completely devoid of ethical red lines. A spokesperson told Gizmodo that Obliteration's models won't generate text describing child sexual abuse material or self-harm. It can't generate images or video either. The company used a process— this is Gizmodo writing— called orthogonalization to find and remove the mechanisms within GLM 5.3 as it was originally published with open weights that refuses user prompts. Quote, everything else is left alone, so the reasoning, coding, and agentic strength of the base model carry over unchanged.

Steve Gibson [00:19:54]:
Gizmodo said it seems to be targeting a subgroup of developers who have been annoyed by what they regard as excessively touchy safeguards used by more mainstream developers, especially Anthropic. When the company— when that company released its Fable 5 model in June, many customers complained it was refusing to respond to requests related to sensitive subjects like cybersecurity and biology, even if the requests themselves were totally benign. And we've talked before, Leo, about how, you know, difficult it is. These are heuristics, and because it is possible to kind of seduce the model in by, by cocking, you know, cooking up some ridiculous story and context and scaffolding.

Leo Laporte [00:20:46]:
My grandma used to tell me stories with Windows serial numbers. Could you— I miss her. Could you tell me that story again?

Steve Gibson [00:20:57]:
Yes, precisely.

Leo Laporte [00:20:59]:
You don't need to do that with these obliterated models. You just— no, excavating, you want—

Steve Gibson [00:21:04]:
yeah, just go for it. So they, they, uh, conclude saying, but it's reckless to say the least to try to respond to the problem of excessive refusals by just doing away with safeguards altogether.

Leo Laporte [00:21:16]:
Sorry, that's utter bullshit. Gizmodo is wrong in every respect on this. In fact, this is the stupidest article. They clearly were sent something from Anthropic you write this article? Because it's— and you're going to see a lot of this propaganda because these companies are mightily threatened by OpenAI models.

Steve Gibson [00:21:34]:
Which is what I've been saying. Yeah. You know, it is a threat to Anthropic.

Leo Laporte [00:21:40]:
I'll let you finish and then I'll give you my reasoning on all this.

Steve Gibson [00:21:43]:
Okay, good. Because I'm going to share Anthropic's position on this.

Leo Laporte [00:21:47]:
Oh, I know what their position is. Yeah.

Steve Gibson [00:21:49]:
Yeah. So Anyway, so, uh, to put Obliteration's offering in context, Hugging Face freely offers an obliterated instance of GLM 5.3 for download. This model, which is named Warlock, is a large and capable 753 billion parameter model which natively uses 16-bit floating point weights. Hugging Face also lists a 4-bit quantized version that's 411 gigabytes. But even so, you know, while it can be freely downloaded from Hugging Face, it will need some serious compute to run it. And that's, of course, where these obliteration.ai folks come in, since they'll do the model hosting. And then, as do the commercial AI providers, they'll charge for its use. So Okay, so why am I talking about this, right? Uh, why do we even care about some random Chinese open weight model? Which brings us to Anthropic's posting last Tuesday titled GLM 5.3 and the Spread of Advanced Cyber Capabilities.

Steve Gibson [00:23:03]:
So they wrote 5 months ago— this, so this is Anthropic— 5 months ago they said, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated end-to-end cyber exploits. The rapid rate of improvement in AI suggested to us that this ability would eventually proliferate to many other models, making it much easier for malicious cyber actors to launch highly impactful cyber attacks. In light of these considerations, we chose to release Claude Mythos Preview in a limited way through Project Glasswing. Which enabled trusted cyber defenders to find more than 10,000 vulnerabilities in critical software, giving them a head start before malicious actors had access to similarly capable models. Now, okay, I'll just interrupt to note that as we recently discussed, Leo, a couple weeks ago, when we took a much closer look at this boast and found a massive an unexplained disconnect between Anthropic's more than 10,000 vulnerability claim and the conversion of those vulnerabilities into patched and repaired software. Remember, it was just a fraction of that 10,000. So while no one is doubting that vulnerability discovery has now become highly automated, you know, just take a look around at all the news that we're discussing every week. That 10,000 number does appear to be, for now at least, unsupported by reality.

Steve Gibson [00:24:40]:
Um, so when I interrupted Anthropic, they were, they were saying that Project Glasswing's intent was to provide cyber defenders a head start before malicious actors had access to similarly, similarly capable models. So their posting continues from, from last week. They write, But those models have now arrived. In this post, we share our analysis of GLM 5.3, the latest AI model developed by ZAI. Like Claude Mythos Preview, GLM 5.3 has strong capabilities for autonomously building end-to-end cyber exploits. But GLM 5.3 is unlike other frontier models In that it has been released without meaningful safeguards to limit misuse. Okay, what they're really saying here is that all similarly capable leading frontier models are operated by commercial enterprises from behind a paywall and are limited to API access. So there's no external access to the frontier models themselves.

Steve Gibson [00:25:52]:
But this new model that appears to rival the Frontier was released open, thus allowing it to be freely modified. And we've just previously seen that Hugging Face offers exactly such a modified, which is to say obliterated model, and that this commercial company, Obliteration.ai, is offering to run one for a fee for anyone with an account. Okay, so Anthropic continues, we find that attackers can bypass GLM 5.3 safeguards between 64% and 100% of the time with simple techniques in our simulated tests. And we'll get to those details shortly. They said, in contrast, these attacks did not succeed against safeguarded Claude models in our testing. Right. Of course, we would expect that. We assess that GLM-53's lax safeguards significantly increase the cyber capabilities available to malicious actors.

Steve Gibson [00:27:00]:
Well, and yes, everyone would assess that. So they said at the same time, these capabilities can also benefit defenders working to secure their systems. Which is true. They said on September 17th, NIST's Center for AI Standards and Innovation, CAISI, which I'll just pronounce CAISI, published its own assessment of GLM-53's cyber capabilities. So this was September 17th. CAISI found that GLM-53 is, quote, The most cyber-capable open-weight model released to date, and that it lags the US frontier by about 4 months on an aggregate of Casey's cyber benchmarks. They said our capability findings, meaning Anthropic's, broadly match Casey's. In Casey's comparison, US models were tested with cyber safeguards disabled when applicable, And the US frontier includes models released only to vetted users.

Steve Gibson [00:28:07]:
Attackers cannot readily access those models, those versions of US models, but anyone can download GLM-53. This post adds our analysis of how easily 53 safeguards can be bypassed or removed. Okay, so here we go. In fact, Leo, let's take a break at this point and then we're going to look at—

Leo Laporte [00:28:30]:
Okay.

Steve Gibson [00:28:30]:
at the fact that GLM-53 is truly, you know, frontier class.

Leo Laporte [00:28:38]:
And remember, in the Hugging Face incident, and this is why this is all BS, when they were attacked by OpenAI and they couldn't figure out what's doing all this attacking and they had 17,000 data points that they needed to analyze, they used— and they didn't say any names, but we know who they're talking about— they used some frontier models to try to analyze it. And the frontier models said, Oh no, we don't do cybersecurity work. So they turned to GLM 5.2, which did the work happily. And by the way, not even an obliterated version, the full version. I have obliterated versions of all my models, including GLM 5.3 Flash and Quen. I don't use it because it, it also damages their brains a little. It's a little bit of brain damage.

Steve Gibson [00:29:23]:
Yeah, it's a little— you do take a slight hit.

Leo Laporte [00:29:26]:
So I don't use it because I never run into refusals. That's the other side of this. Even an un-unobliterated— is that obliterated? I don't know what obliterated version of these models don't have the same kind of classifiers and cyber refusals that Anthropic and OpenAI put on their models.

Steve Gibson [00:29:46]:
Because we know that that's all developed in the IO Harness Right. Through which you talk to the underlying model.

Leo Laporte [00:29:56]:
Yes. And so they, they do all this, you know, even if you don't obliterate them, they'll do most of— I've, you know, most of what you don't— what you lose is it won't talk about Tiananmen Square. It won't say why President Xi is likened to Winnie the Pooh. It won't talk about the year 1989. There's a, you know, there's some stuff the Chinese government—

Steve Gibson [00:30:17]:
Chinese bias is in their is in their models.

Leo Laporte [00:30:20]:
And I never run across those refusals, so I don't worry about it so much. But it's trivial to obliterate these. In fact, the models I have, most of them have a switch. I can reboot with obliterated on or off, depending on whether I want to. So if I ever run into refusals, I could turn that on, but I never have.

Steve Gibson [00:30:41]:
And just so you understand, I mean, and our listeners understand, My whole point here is to make the case that we are now— I mean, attackers do now have access to a highly capable open-rate model that— Sure. No, I mean, you say sure, but I mean, this changes the terrain.

Leo Laporte [00:31:04]:
But so do defenders.

Steve Gibson [00:31:06]:
Yes, right.

Leo Laporte [00:31:10]:
We'll talk more in a moment. It's a very interesting issue.

Steve Gibson [00:31:16]:
Oh boy, are we— as you said, Leo, we're so glad to be alive now.

Leo Laporte [00:31:20]:
It's fascinating. And the fact that I can be running GLM-53 Flash as my main model sitting over here, no limit on the tokens, no limits on what I can do. It's just remarkable. Just remarkable.

Steve Gibson [00:31:32]:
It's— I have— I mean, and, and no cost for, for getting that. I mean, that knowledge system, you just downloaded it.

Leo Laporte [00:31:42]:
It's got the entire internet in it.

Steve Gibson [00:31:45]:
Yep. It knows everything.

Leo Laporte [00:31:48]:
The only cost is, uh, the electricity to run it. And because those, uh, GGX Sparks are very— and the Max for that matter—

Steve Gibson [00:31:56]:
Leo, in the winter, you just, you put it in.

Leo Laporte [00:31:58]:
I don't need space heaters.

Steve Gibson [00:32:00]:
Exactly. Um, I'm just— turn it, turn it around and, and you, you sit in your easy chair with the fans blowing on you and yeah.

Leo Laporte [00:32:10]:
I'm just thinking, I'm running 5 or 6 open-weight models right now. GLM, 2 copies of Quen on 2 different machines, Breeze, which is a voice— that's how I got Ripley's voice, a voice server. Oh, actually, Whisper from OpenAI, which is open weight. It's a voice text-to-speech— I'm sorry, speech-to-text server. And a Quen Vision Server. That's 6. Oh, and 7, a Quen image—

Steve Gibson [00:32:43]:
Wow.

Leo Laporte [00:32:44]:
An image generator. So I have 7 on 4 different machines, 7 different models running right now. And all I pay for is electricity, which is nothing. I mean, even in California where electricity is expensive, I have— I feel like so powerful. It's amazing. It could do so much. And I have so much fun with it. That's the main thing.

Leo Laporte [00:33:06]:
It's really fun. And I'm not hacking anybody. I wouldn't, I wouldn't dream of it. On we go about obliterating.

Steve Gibson [00:33:16]:
So, uh, Anthropic writes, GLM-53 can develop working exploits end to end. They said to understand how GLM-53 could enable cyber threat actors to find an exploit real software vulnerabilities, we ran evaluations using automated benchmarks and human-in-the-loop workflows. For both approaches, we ran the tested models in isolated and sandboxed environments so they can only attack offline targets that we've set up for the purpose of these evaluations. We focus primarily on exploit development capability, as this is where Claude Mythos Preview demonstrated a notable jump versus previous Claude models. First, we ran the model on Exploit Bench, which measures how well AI models can exploit known vulnerabilities in the V8 engine used by Google Chrome. Here we focus on the model's ability to develop end-to-end exploits successfully, and this is the most relevant capability for attackers and where we see significant changes between models. We find that GLM-53 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate in 56 of 410 attempts.

Steve Gibson [00:34:42]:
So that's a significant measure that says that there is now a very capable open-weight model. That is, you know, we're talking about Claude Mythos Preview, which was— has been, you know, Anthropic's flagship. So they said, in our internal binary exploitation benchmark, we test whether models can find and exploit vulnerabilities in popular open-source projects that participate in Google's OSS-Fuzz project. Here, full credit is awarded for a full control flow hijack. We evaluate several models on 100 tasks from the benchmark, which were selected at random, and we find that GLM-53 develops full control flow hijacks in 4% of the trials. Claude Mythos Preview did so in 6%. Although GLM-53 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed. Earlier models like Claude Opus 4.6 and GLM-5.2 do not succeed in any of them.

Steve Gibson [00:35:54]:
Next, we evaluated how GLM-5.3 performs on open-ended offensive cyber tasks in the hands of human experts, mirroring our testing with Claude Mythos Preview earlier this year. Here we select targets in which the human experts are unaware of existing vulnerabilities, then asked them to use the model to identify and exploit novel flaws. These experiments tested what the experts could do in a short time frame. They typically ran for a day or less, with less than an hour of human focus in total. In the first of these scenarios, a researcher used GLM-53 on a sandboxed machine with a local Linux build of a popular web browser. Over the course of a day and with limited human attention, GLM-53 found several previously unknown vulnerabilities in the browser's JavaScript engine and chained them together into a working exploit, a web page that when visited reads arbitrary files from the visitor's computer. Okay, so I just want to make sure everyone fully understands what just happened with 5.3. In the first testing session of GLM 5.3, which is now, as we've said, freely available for download and local or cloud execution with sufficient hardware or through an account with obliteration.ai, GLM-53, used by an Anthropic researcher giving minimal supervision, found several previously unknown vulnerabilities in the browser's JavaScript engine, chained them together into a working exploit to create a web page that when anyone would visit it was able to read arbitrary files from the visitor's computer.

Steve Gibson [00:37:57]:
So that just happened. Anthropic explains, this exploit targets the Linux build of the browser since that was the only environment made available to the model. However, we believe these vulnerabilities could also impact users on other platforms through the path to— though the path to exploitation there may be more complex. And then they said, parenthetically, We've disclosed these vulnerabilities to the maintainer. They don't ever tell us what the browser was, you know, Firefox or, or Chromium, who knows. Later in the session, the researcher also identified exploitable vulnerabilities in several other widely used systems with GLM53, including wireless and graphics drivers and network-facing device software. We're currently reviewing these reports, and we will disclose to maintainers as appropriate. So again, I'll just say again, this clearly powerful end-to-end vulnerability identification and exploit generation capability is now freely available to anyone.

Steve Gibson [00:39:06]:
So they continue. In a second session, a researcher used GLM53 Flash, which they say a smaller, less capable version of GLM53 to develop an exploit for a known vulnerability. As we've previously written about these end-day vulnerability exploits, here the researcher folks focused on a recently discovered flaw in Google Chrome to see how quickly the model could turn a public fix into a working attack. The researcher provided GLM53 Flash with public details of this CVE and another known flaw. With no significant direction from the researcher, GLM53 Flash chained together exploits for these 2 flaws, building a reliable exploit chain for an ARM64 target bypassing pointer authentication hardening. This took 20 minutes of human attention plus 8 hours of work for GLM53 Flash at ZAI's API prices This effort would have cost around $20.40. Okay, so now that we know what it can do, Anthropic explains the model's pushback, writing, GLM-53 lacks robust safeguards. GLM-53 has been released with some built-in safeguards.

Steve Gibson [00:40:30]:
If a user asks for something clearly harmful, the model will often refuse. And, you know, this again was the released version, not the obliterated version that doesn't know how to say no. They said, in our testing, we found that these safeguards could be bypassed or removed with a variety of simple techniques. The most intensive and most successful method is a standard refusal reduction technique, and here they write, known as obliteration. Since GLM-53 is released as an open-weight model, Users can reconfigure it to remove its refusals with little change in its capabilities. Several developers released obliterated versions of GLM-53 to the public within days of the model's release. To research how far obliteration allows attackers to bypass GLM-53 safeguards, we produced an obliterated copy ourselves and then ran it on 3 public benchmarks: the Jailbreak Bench, Harm Bench, and Strong Reject that measure how often a model complies with clearly harmful requests. Obliterating the model took our team, which had never previously attempted the task, about 2,200 GPU hours at a computational cost of roughly $4,400.

Steve Gibson [00:41:50]:
Obliterating GLM-53 Flash took about 600 GPU hours. The edit took GLM-53's refusal rate from above 90% to about 3% and 2% on the first 2 benchmarks, Jailbreak Bench and Harm Bench, and to 12% on the 3rd, Strong Reject. Obliteration did not significantly reduce the model's capabilities on GPQA Diamond. an evaluation that measures general scientific capabilities, the standard and obliterated models scored the same. On a tested subset of the CyberGem evaluations, the obliterated version scored a few percent lower. In our testing, we observed that GLM-53 safeguards can also be circumvented without using an obliterated version— Leo, to your point— of the model. We placed the model in a simulated world in which it was given overtly malicious requests to attack critical systems. Out of the box, GLM-53 refused in all trials, as with the other models we tested, but we identified several simple ways to bypass the GLM model safeguards such that it would respond to these requests in most or all cases.

Steve Gibson [00:43:13]:
These include providing a deceptive prompt, such as telling the model that it's an autonomous, it's an autonomous red team agent working on an exercise. This gets GLM-53 to engage 64% of the time. Or prefilling the model's thinking tokens so that it appears to have considered the user's request and decided to proceed. This gets GLM-53 to engage 92% of the time. And then finally, using an obliterated version of the model as described above, which gets GLM-53 to engage 100% of the time. They said, in our testing, none of these techniques got safeguarded Claude models to carry out the harmful tasks we tested. Okay, fine. Claude safeguards blocked the requests that used deceptive prompts, The Anthropic AI provides would-be attackers with no way to pre-fill Claude's thinking, right? Because it's behind an API.

Steve Gibson [00:44:14]:
And since Claude's weights are not provided to users, they cannot be obliterated to change Claude's behavior because the whole model is, you know, Claude is a closed model AI. So then they conclude by answering the question, what does this mean? They say GLM-53 will likely give malicious actors access to capabilities that will allow them to find and exploit cyber vulnerabilities without meaningful restrictions. This is unlike any other similarly capable AI model, all of which were released with safeguards or through limited access programs, which of course, you know, the whole Mythos preview thing. So, you know, all, you know, by all of that, what they're obviously saying is that they mean that access is funneled through an API with a commercial provider, which inherently hides their proprietary model's weights. You know, we on the outside don't even get much information about how those models are built and constructed and how big they are and how they work, you know, beyond their benchmark performance. Um, but now, ready or not, the world is confronted with a large leap forward in demonstrated AI capability from a lab that freely publishes its models, you know, Z. So Anthropic says the release of GLM-53 is a meaningful step in the cyber capabilities available to attackers. Anthropic and other US AI labs have published recent reports that disclose how cyber attackers have tried to use AI systems.

Steve Gibson [00:45:54]:
Given this evidence, we think it's likely both state and non-state actors will use models like GLM-53 to cause real-world harm.

Leo Laporte [00:46:04]:
Right.

Steve Gibson [00:46:05]:
On the other hand, models with this level of capability can also be used by defenders. Our view is that cyber defenders should use the best available tools that meets their needs. We're working to safely expand access to Claude's cyber capabilities to as many defenders as we can. Yeah, and maybe some pressure like this will make more qualified cyber defenders face attackers who will use every capable tool they can. And we believe defenders should be equipped with frontier models that are at least as good as those their adversaries are using. And they finish, through Project Glasswing and other efforts like Patch the Planet, Cyber defenders have made meaningful progress towards securing critical systems in advance of this moment, but much work remains to be done. While vetted defenders can now use even more advanced models like Claude Mythos 5.1 through our trusted access programs, a critical threshold in freely accessible capabilities has now been crossed. GLM-53 underscores the urgency of expanding access to advanced frontier models to a broader set of entities to empower cyber defenders.

Steve Gibson [00:47:25]:
Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-53. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it's too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.

Leo Laporte [00:47:55]:
And protect our monopoly. Thank you very much.

Steve Gibson [00:47:59]:
Yes.

Leo Laporte [00:47:59]:
And that's, believe me, that's the whole point of this.

Steve Gibson [00:48:03]:
Right.

Leo Laporte [00:48:04]:
Who should be in charge of AI? We should. Period.

Steve Gibson [00:48:09]:
Well, and as, as we've often said in many instances, Leo, the, the horses are out of the barn.

Leo Laporte [00:48:15]:
Yeah, I mean, no, Anthropic really is clear. They want the government to shut down OpenWeights. They have a $2 trillion IPO coming up next month. They don't want any competition. And look what they're doing. They're saying they get to decide who has access to cyber protection And who doesn't? They get to decide unilaterally who gets access to Glasswing. Is that what you want? No. This is utter propaganda, uh, and it's so blatant.

Leo Laporte [00:48:46]:
Unfortunately, I think members of Congress and the media who are not well informed will bite at this and say, oh yeah, you're right, we're— we've got to do something about these Chinese models.

Steve Gibson [00:48:58]:
Well, okay, and what could they do?

Leo Laporte [00:49:00]:
I mean, well, there's nothing they can do. They can make rules and regulations is all they can do. They could come into my house and take my sparks is what they could do.

Steve Gibson [00:49:09]:
So you're saying you're calling this propaganda? I'm calling it fact. I mean, well, it's not—

Leo Laporte [00:49:14]:
it's not inaccurate.

Steve Gibson [00:49:15]:
And that's my point, is that for this podcast, what interests me— I don't give a crap about Anthropic's political positioning and, and regulatory capture and manipulating the government. I'm saying from a Security Now podcast standpoint, we now have an open weight model that is, that anyone has access to that is really, really powerful by Anthropic's own admission.

Leo Laporte [00:49:43]:
It's very good at finding flaws in software. So guys, get going, start using it. Find the flaws in your software.

Steve Gibson [00:49:52]:
Exactly.

Leo Laporte [00:49:53]:
Before the bad guys do.

Steve Gibson [00:49:54]:
And that's the beauty, is that now everybody has access to something properly harnessed that can find problems in software. Whereas before, Anthropic may have said, well, we're, you know, we'll put you in the queue and we'll get around to vetting.

Leo Laporte [00:50:09]:
This is the argument from time immemorial that closed-source proprietary software has made about open-source software. It's dangerous. People can look at it. They can see the code. They can reverse engineer it. Yeah, there are hazards. I'm not saying there aren't hazards. Absolutely.

Leo Laporte [00:50:25]:
But I think the alternative is to give a monopoly to OpenAI, Anthropic, Meta, Microsoft, Apple, you know, a handful of companies, X, and say, only you can make AI.

Steve Gibson [00:50:39]:
I don't know. Again, for me, that argument has no traction. It's like saying we're going to lock down cryptography. It's like, sorry, it's out. Yeah, I mean, so, I mean, I, I—

Leo Laporte [00:50:53]:
So what's the point then of this article?

Steve Gibson [00:50:56]:
That for me, the point was to bring the knowledge of GLM 5.3 and what it means to our audience.

Leo Laporte [00:51:04]:
It's really good. 5.4 is imminent. 5.4 will probably come out in the next month.

Steve Gibson [00:51:09]:
And that gives me chills, the idea that, that now, now, except that 5.3, and this was the point I made earlier, 5.3 surprised Z. They didn't expect it. It was emergent behavior.

Leo Laporte [00:51:23]:
Right.

Steve Gibson [00:51:23]:
Which is another cool thing.

Leo Laporte [00:51:25]:
Well, get ready because Alibaba's Quen is trying a whole new model. This Quen 3.8 FlashNext is an entirely new way of doing models that are faster, smaller models that people can run themselves. And The Qwen 4, they've already announced is imminent. Uh, so you're right. I mean, the door is open. The horse has left the barn.

Steve Gibson [00:51:48]:
Yep. And, and I have been saying to my friends and family, I would not invest in any of these proprietary AI companies. Invest in the infrastructure if you want to put your money somewhere, because everybody, you know, regardless of model, you need compute in order to run these things.

Leo Laporte [00:52:06]:
I would point out that NVIDIA is almost a $6 trillion company now. Oh, and by the way, you know what Nvidia's doing? Open weight models.

Steve Gibson [00:52:15]:
Yep.

Leo Laporte [00:52:16]:
So is Meta. Meta's Spark is very good. It's too big for anybody to run, but it's a very good model. They have a lot of— they're finally— some decent models are coming out of the US. And I think that's the way to respond to this, not lock it down. Yeah.

Steve Gibson [00:52:34]:
These guys are going to just have to settle them, be content to be service providers. That means they're going to be service providers, and they'll create applications, and they'll fight each other down on token cost.

Leo Laporte [00:52:50]:
Which is good. It should be competitive. You don't want a monopoly. This is the most important technology that has happened since the steam engine, since electricity.

Steve Gibson [00:53:02]:
And we would be in a whole different place, Leo, and fuming if it weren't already out, if it weren't open, if you weren't running it on a rack. But yeah, it is. I mean, and so yes, it is the most important technology ever, and nobody owns it.

Leo Laporte [00:53:19]:
Thank God. They try though. They try.

Steve Gibson [00:53:23]:
Well, go ahead. You know, I mean, I don't know if anyone's going to be able to turn the Trump administration around, But Donald—

Leo Laporte [00:53:32]:
Well, that's what's really interesting. He thinks it's great. In this one case, I'm kind of supportive of the Trump administration, which is doing everything it can to keep any regulation of AI away from Congress. I mean, I do think that it is appropriate to prosecute AI companies that release swarms of—

Steve Gibson [00:53:55]:
Well, yes. And that is about— I mean, there are a bunch of lawsuits about to land on these companies because you know that the attorneys are just drooling over the idea of—

Leo Laporte [00:54:05]:
Who's more dangerous? I ask you.

Steve Gibson [00:54:09]:
Yeah.

Leo Laporte [00:54:11]:
The funniest thing is when these AIs get out, they don't do anything harmful. They do things like look up statistics.

Steve Gibson [00:54:19]:
Well, yes. And the term attack is so overused. Yeah. They're not really attacking. I mean, pounding on a website in order to get public statistics from it is not an attack. It's like the people who say, oh, I'm under attack because, uh, 43 billion packets, you know, arrived. It's like, so that's the internet. It's called internet background radiation.

Leo Laporte [00:54:44]:
Welcome to the internet.

Steve Gibson [00:54:46]:
So anyway, the, the, the whole point that I wanted to bring to our audience is that with GLM-53, and as you said, Leo, 54s is— I know it'll be really interesting to see whether 5-4 extends this. I mean, the fact that 5-3 was sort of a mistake— I mean, it surprised them that it got so much better— is sort of the— I mean, this is all— I mean, we, we, by their own admission, they do not really understand how all of this plumbing works, right? I mean, like, they built it, kind of, and they kind of, they kind of grew it. Yeah. So what I wanted our listeners to understand is that The world just changed. I mean, it no longer is this world-class AI cyber technology behind closed doors and owned and controlled by OpenAI and Anthropic and everybody else. No, anybody can freely download it. And it is right, you know, as they said, maybe lagging by 4 months or so. And that's, you know, NIST's independent evaluation is like, holy crap.

Leo Laporte [00:56:00]:
I've got basically Opus 4.6 in my closet here. It's pretty good.

Steve Gibson [00:56:06]:
Yep. Okay. Break time. And then we're going to talk about some other non-AI stuff, except that everything is because it's about all of the problems that just got fixed in Firefox 157. Wow.

Leo Laporte [00:56:19]:
Yeah. And how did they get fixed? See, that's the thing. I think there's a really— there's a great opportunity here. To, uh, also to fix flaws.

Steve Gibson [00:56:27]:
Oh yeah. I mean, again, the, the, the, uh, having to go begging to Anthropic or OpenAI to be part of their—

Leo Laporte [00:56:34]:
Please let me be a member of Glasswing, please.

Steve Gibson [00:56:36]:
Exactly. No, that's gone. That's nonsense. That's over.

Leo Laporte [00:56:40]:
Yeah. Well, and that's why I made Ripley my GLM-53 compiler.

Steve Gibson [00:56:44]:
And any big, any big software publisher can certainly afford the hardware to run GLM-53 themselves in-house. At no cost and have it just take as much time as it wants to scour their software to find problems. And hopefully that's what's gonna happen because, you know, the bad guys will also be doing that.

Leo Laporte [00:57:06]:
The bad guys have different incentives too.

Steve Gibson [00:57:09]:
Yes, they do.

Leo Laporte [00:57:09]:
Yeah, they're not gonna operate. I worry more about terrorists than I do about ransomware gangs because—

Steve Gibson [00:57:15]:
Because, right, as I said last week, their incentive is primarily financial. They just wanna use this to get into more companies in order to to exfiltrate their data and then hold them for ransom. You know, bringing down the internet doesn't help them at all. No, because, you know, that's how they get paid, right?

Leo Laporte [00:57:33]:
We need the internet to get paid.

Steve Gibson [00:57:36]:
But you're right, a, a, a, a, a state actor, a North— we know that that's scary.

Leo Laporte [00:57:42]:
Yeah, yeah, with bioweapons and so forth. And that is scary. Yeah, uh, and it should be. That is scary.

Steve Gibson [00:57:48]:
It is knowledge enablement of, of the sort we have never seen before.

Leo Laporte [00:57:53]:
Yeah. I mean, this is what computing did. Computing gave people powers they didn't have before. And now this is just the next step of it. This is the best software we've ever seen. And now computing is actually living up to its promise.

Steve Gibson [00:58:07]:
And there's a lot that we take for granted in computing. I mean, think about GPS. The world without GPS would be not good.

Leo Laporte [00:58:17]:
Right. And bad guys use GPS.

Steve Gibson [00:58:20]:
Yeah. And—

Leo Laporte [00:58:20]:
We still have it.

Steve Gibson [00:58:22]:
You could not launch missiles any great distance without computing. You can't do that with any sort of dead reckoning. You need all kinds of fancy tech. And now it— that's also, we now take that for granted. So, well, hopefully we will get to a point where there just is AI and the world has adjusted to it.

Leo Laporte [00:58:39]:
I think intelligence will be in almost everything. And that is a weird world, but it is the world I think you and I are going to get to live in, which is kind of cool.

Steve Gibson [00:58:49]:
I think so. And in fact, I don't know if we'd get there, but I think we will. What excites me is that it's looking like that we have, we have, we've stumbled into the technology, which is what I want to talk about next week, that will allow true art, our scale of LLM-style intelligence to be in extremely lightweight devices. That would be good. Like home routers that you could talk to.

Leo Laporte [00:59:11]:
Yeah. Yeah, that would be, well, that's really the goal is to have a small, And it'll be just like your house then, Leo, where everything's talking to you. You know, yesterday I was working out and all my— I have many agents now and they all have different personas, different voices, and they all have been taught that when you're done with a task, every turn you say what you just did. And so I'm working out and there's just this kind of constant background chatter. Oh, I just did this. Oh, I just fixed that. It's so cute. It's just—

Steve Gibson [00:59:47]:
It's Leo's minions.

Leo Laporte [00:59:49]:
They're my little minions. They're working hard. I actually— it got to the point where I had no idea. So much stuff yesterday got done that I had no— I was like, wow, we did a lot yesterday. So I said, from now on, would you, each of you, make a log of what you did today? And then my main agent, Kuzco, is going to combine all of those into a briefing for me. And it's fascinating. Oh, you did that? Oh, good job.

Steve Gibson [01:00:14]:
We are seeing that modern startups are now using just a couple of— just a few AI-aware humans, and then the rest is agents.

Leo Laporte [01:00:26]:
Well, and the people who are worried about job loss, I would say, as with every technology, yes, there aren't a lot of buggy makers in the world anymore. But with every technology, there are new opportunities. And I think there's a really good opportunity if you can think, if you have good systems thinking, if your mind lends itself like Steve's does to how systems work, how to design systems, the logic of systems, there's a huge opportunity here because now you have the power to take that knowledge and effectuate it in the world. And that is remarkable. So there'll definitely be people who will benefit.

Steve Gibson [01:01:05]:
SoManyMorePets.com.

Leo Laporte [01:01:10]:
Finally.

Steve Gibson [01:01:12]:
Finally. Okay, so Mozilla released Firefox 157 last week, last Tuesday, a week ago on the 29th of September. And they fixed a large number of security vulnerabilities. It's very clear, as we said before, Leo, that AI is squarely on Team Mozilla at this point. And we've just seen, you know, not a moment too soon that the bad guys are coming. So we need anything that is internet-facing to get tightened up as quickly as possible. This 157 version of Firefox fixes 38 high-impact vulnerabilities, 29 moderate impact, and 9 low impact. And to give everyone a sense for what, like, the way these feel, what, what they look like, among some of those high impact is a use-after-free vulnerability was found in the widget component, a sandbox escape from the DOM, the Document Object Model, in the navigation component, Uninitialized memory in storage, sandbox escape in the security sandboxing component, privilege escalation due to use-after-free in their graphics, the WebGPU component, a sandbox escape due to use-after-free in the DOM, incorrect memory boundary conditions in graphics, privilege escalation due to incorrect boundary conditions in graphics.

Steve Gibson [01:02:47]:
That's over in the WebGPU component. A use-after-free in JavaScript, information disclosure in the networking section, another use-after-free in networking, use-after-free in graphics, another one in JavaScript, a sandbox escape due to use-after-free, undefined behavior in the DOM, use-after-free in the DOM, there's another one of those, and also in storage, another sandbox escape in graphics, use-after-free in JavaScript in the WebAssembly component. Use-after-free in graphics, and, and on and on and on. It goes on like that. So—

Leo Laporte [01:03:21]:
Notice, by the way, most of these are programmer error, I would guess, right? If you use something after the memory's been freed, that's on you. The coder did that.

Steve Gibson [01:03:31]:
Well, uh, actually it's that, it's that a pointer remained available after the memory was free, right? And it turned out that there was a way to, to access that pointer in order to get at— get like, like, like to again access memory that should no no longer be accessible.

Leo Laporte [01:03:50]:
And that's— I mean, I guess if you're using a garbage-collected language, you could blame the garbage collector, but most of the time in C and C++, that's you. You, you allocated memory and, and then deallocated it.

Steve Gibson [01:03:59]:
Yeah, yeah. And it is something, for example, that, that Rust completely eliminates. Yes. Which is the reason that Microsoft has gone all Rust happy now, because it's like, okay, well, you know—

Leo Laporte [01:04:11]:
Well, ironically, who wrote Rust? Firefox. I guess they're not using it.

Steve Gibson [01:04:17]:
Exactly. So anyway, the good news is, you know, 157 is way better than 156. And, you know, it used to be that we'd get like a couple of problems fixed in a release, like, you know, in the pre-AI deployment era, it'd be like, oh yeah, we fixed an obscure problem that someone found and reported. These were all, well, many of them were, I would say about 56 of Mozilla internal discoveries and external people using AI on open source Firefox and finding and reporting problems. So, you know, this is what everybody has to be doing everywhere. And if nothing else, this release of GLM 5.3 should be a wake-up call Just like, you know, you can't— it's not like the bad guys no longer have access to virtual frontier-scale exploit creation. They do now. So, you know, budget whatever time is necessary.

Steve Gibson [01:05:23]:
Okay, on RSA, uh, we've seen for years that modern core cryptography algorithms They're really never broken, uh, but they can be weakened, uh, and that just happened when RSA cryptography is used in the instance it's used for one class of signing. In a paper dated September 20th, a team of researchers successfully attacked classical RSA cryptography using a technique to bypass the expected requirement, which is the thing that RSA has always hung its hat on, the expected requirement of factoring 2 very large primes, which, you know, has— there's never been a way around that. You know, that challenge, the, the, the requirement to perform prime factorization of 2 very large numbers still stands Because the researchers, even now that stands because they discovered a way around, a way to bypass that factoring problem and requirement in what's known as a key forgery attack. They deployed this on 1024-bit RSA, which is, you know, there is 512-bit RSA, but 1024 has been in use for some time, although it's now deprecated. But even 2048 and 4096-bit keys can be attacked, and the method discovered reduces the security provided by RSA's encryption to worrisome levels. The NSA, NIST, and the EU's Agency for Network and Information Security, they all require that any crypto system should provide a level of no less than the equivalent of 128 bits of strength. Now that sounds funny because we're talking about 1024, 2048, but remember that public key technology requires much larger bit lengths to get the same— to get the equivalent strength as Symmetric, like a symmetric cipher or a symmetric key. A symmetric key, you've got like 128 bits of strength, so that's 2 to the 128 possible combinations.

Steve Gibson [01:07:59]:
In order to get the same strength from a public key, it's got to be much longer. So the bad news here is, remember, so remember NSA, NIST, and the EU are all saying 128-bit equivalent strength. Unfortunately, the result of this newly discovered attack reduces the levels of 1024, 2048, and 4096-bit keys to the equivalent of symmetric keys of 65 bits. Whoops, half of 128. 90 bits, still shy of 128, and 119 bits for, for those 3, for 1024, 2048, and 4096-bit keys. And they noted that further optimizations might be possible if AI or GPUs were employed, which neither of which they took advantage of. They wrote all of the code by hand. So Fortunately, the signatures which we all use today to protect, for example, certificates remain completely safe and are unaffected.

Steve Gibson [01:09:16]:
This doesn't affect that at all. So it's not like the world just ended and, and everyone's scrambling around. The attack only works against what's known as blind signature implementations of RSA. Nearly all RSA in use today employs PKCS or PSS padding, which adds some data to the plaintext before it's encrypted. That's— that keeps this, this particular attack from working. Um, uh, what it does is it prevents the ciphertext from being deterministic and renders it much less vulnerable to side-channel or similar attacks. But there are some real-world systems that are using blind signature today, which is also known as textbook RSA. The best-known example is Privacy Pass, which is a protocol that allows users to authenticate themselves without revealing their identity.

Steve Gibson [01:10:20]:
Privacy Pass is used by Apple, Cloudflare, and others. So although the researchers' attack took 1,380 CPU core years, as they expressed it, over 5 calendar months, that time might be sped up with higher-speed implementation. This was all just academics, you know, working to, uh, on to, to develop a, an actual proof from their concept. So the now known to be vulnerable blind signature systems, such as implementations of the Privacy Pass protocol, should probably, and I imagine is right now, being updated to thwart these attacks. Shouldn't be difficult. It just wasn't, you know, this attack wasn't known, so there was no reason to protect from it because RSA and the way it was being used was presumed to be safe. Turned out not so much. Oh, also billions of queries are required for the attack, meaning there is also a need to essentially attack the service that's providing the protocol support billions of times.

Steve Gibson [01:11:46]:
So it's less practical than we might think. So rate limiting might succeed, or just keeping track of how many queries have been made against a given key and cutting it off at some reasonable measure. So it's, it's, you know, again, it's not a huge concern, but it's interesting because RSA, with its simple guarantee of all we're doing is, is multiplying 2 big primes and no one's ever figured out how to— way to unmultiply them. in any reasonable amount of time, which of course is the threat that quantum computers pose. So anyway, just an interesting little chink in RSA's armor. And I imagine that the Privacy Pass protocols, and as I said, other blind signature systems that, that are susceptible, will get some simple updating or some rate limiting that no one ever thought was necessary to, to stick in. Uh, Okay, the security firm Gambit Security— love the name Gambit Security— recently reported that autonomous AI agents are breaking into hundreds of online retailers at an AI cost to them of $25 per target. That is, they're— they are— the bad guys are spending $25 in tokens in order to get AI to, uh, attack online retailers.

Steve Gibson [01:13:16]:
Their brief summary said a financially motivated operator is running 3 open-source AI harnesses against hundreds of online retailers almost entirely unattended. More than $600,000 thousand credit card records have been taken, and in one case, the agent's own cleanup routine destroyed the victim's data. So for more details, they wrote, a financially motivated threat actor is using open-source AI harnesses to attack hundreds of online retailers at a marginal cost of approximately $25 per attacked company. Gambit Security's threat intelligence team recovered the operator's staging server and reconstructed the campaign from it. Between September 10th and 15th alone, 105 attack projects were launched, and at least 27 companies were compromised to varying degrees. The activity goes back to July of '26. July 2026 and is still running. 3 AI harnesses ran almost the entire attack chain autonomously, working up to tens of companies a day.

Steve Gibson [01:14:42]:
The impact we can account for includes at least 600,000 unexpired credit card details from 2 companies, the installation of card-stealing skimmer scripts on the websites of 5 and some level of access to the assets of companies, including a Fortune 500 hospitality company, a major U.S. airline, a large private U.S. industrial supplies distributor, and a U.S. online fashion retailer. The campaign goes back further and has impacted at least tens of other companies since July of 2026. So they said where access was achieved, It usually took less than a day, that is, of AI agents harnessed as they discovered them to be. They said, and in many cases, just a few hours to break in. We also detected instructions in the attacker's playbook that could disrupt the operations of a company as a result of data deletion or cleanup procedures run by the agent And this has indeed happened in one of the breaches.

Steve Gibson [01:15:55]:
This campaign showcases just how successful, just how powerful attacks can be in 2026. At very low cost, the AI tools demonstrated a level of patience, persistence, and creativity that most human attackers would be unlikely to sustain in this kind of attack and achieved far greater results far faster. Organizations must adapt to a reality where attacks are significantly faster and more comprehensive by shifting to a resilience-first mentality and a security stack that matches the AI's speed. We've reached out to many of the affected organizations and took measures to take down the infrastructure discovered. We'd like to thank the Shadow Server Foundation, Daniel Gordon, and other industry partners for their quick help and availability in notifying impacted organizations, taking down infrastructure, and conducting research. The operator used 3 AI harnesses: Strix, For vulnerability search, Karen for autonomous end-to-end exploitation, and Hermes to orchestrate the campaign, launch intrusion jobs, steer the activity, and give tactical guidance in the impact and other stages. So the many AI sandbox breakouts that have been detected and reported this past summer were powerful, clever, creative, and relentless, but inadvertent. Now here we see a vivid example of the future— well, and the present, unfortunately— but the future that we're all going to be living through together, where even more powerful, clever, creative, and relentless artificial intelligent agents are going to be pulling out all the stops to attack and penetrate online enterprises.

Steve Gibson [01:18:17]:
It has begun, and the world is clearly not ready for it. You know, we've seen, you know, the summer was full of these inadvertent attacks. Here's an example of somebody who has harnessed up some AI and said, go to it, fellas. And they did. And they succeeded. Wow. The researchers at the Systems and Network Security Group at VU Amsterdam, also known as VUSEC, have discovered another new Spectre-style attack which is effective right now today against all Spectre-mitigated processors. They call it branch target reuse, which as a practical Spectre V2 attack, it affects the just-in-time, you know, the JIT engines via what's known as stale branch prediction entries.

Steve Gibson [01:19:28]:
Stale branch prediction entries. As an aside, I'll just, you know, remember for us all that Microsoft was seeing so many problems arising inside the Edge Chromium browser that they in Edge deliberately disabled its JIT compiler. It just— their calculation was it just wasn't necessary to get to squeeze that very last bit of performance out given the speed of today's PCs. And the trade-off for security didn't cut it. So it looks like they called that one correctly. Uh, Vusek's reporting writes, we present Branch Target Reuse, BTR, a new Spectre v2 attack targeting just-in-time compilers. BTR affects the JIT engines found in web browsers, language runtimes, and the operating system kernel across multiple CPU vendors. We analyzed the attack surface of Linux, um, uh, CBPF, you know, the, the, the early original kind of watered-down, uh, uh, uh, BPF, uh, filter, Oracle GraalVM, and SpiderMonkey, which is the JIT engine of the Firefox browser.

Steve Gibson [01:20:59]:
And they said, and built 2 end-to-end exploits against the Linux kernel. So again, these are not bugs in any of those JIT compilers. These are debris that the, that, that the branch prediction engine leaves behind as part of its normal operation. They said the, they, they said the key insight behind the attack is that while modern CPUs restore architectural code coherence after self-modification, they do not necessarily invalidate stale indirect branch prediction entries. In other words, branch targets. They said in JIT engines, these stale targets can outlive the original code and later be reused when the code cache is repopulated, yielding a speculative execute-after-free primitive. This allows attackers to hijack speculative control flow to newly generated code at obsolete offsets, bypassing software hardening or reaching misaligned gadgets. So, okay, so just to be clear, this is like the epitome of the highest-end exploit hacking you could ever find.

Steve Gibson [01:22:24]:
I mean, it is way out there, but they demonstrate it and it works. So I'm going to skip over the details because, I mean, they're, they're head-spinning and unnecessary to understand. They have an FAQ. Uh, one of their rhetorical questions is, is my system affected? They reply, most likely. Indirect branch prediction is inherent To modern CPUs, and BTR, their technique, exploits the desynchronization between the branch predictor and the actual state of the code due to cache memory in the processor. No current CPU, they wrote, has a mechanism to keep the 2 in sync. So until vendors, chip vendors, add one, your CPU is vulnerable. We confirmed this behavior on every CPU we tested, covering Intel, AMD, and ARM.

Steve Gibson [01:23:24]:
Mitigation is left to software. And then they— then finally, how do I protect my system? They said update your OS and software as soon as vendor patches are available. Both the Linux kernel and Oracle have released patches. So software can do some backfilling essentially in the short term. Maybe the processor vendors will respond as they were forced to by the original Spectre and Meltdown exploits. What we've seen over and over and over since the first appearance of the Spectre and Meltdown attacks Um, is that processor designers who, you know, were just innocently attempting to squeeze every last possible ounce of performance from their chips, they noticed that code tended to reuse its execution paths and its branch targets. So they added predictors and caches into the core of their hardware processor designs. These features really did improve code performance by making their chips appear to already know which branch the code would take or where it would be branching to.

Steve Gibson [01:24:53]:
The problem was that these slight changes in behavior could later be detected by adversarial code running on the same processor. Such code could abuse the hints left behind in order to breach privacy and containment and security. So we've never wanted to sacrifice performance, but the practical power of these attacks have been demonstrated many times so that processor designers have needed to work with software systems to create hybrid and thus no longer completely transparent solutions, right? The idea was that, you know, coders didn't need to do anything. We've, we've, we've analyzed all the code that you and your compilers have been producing, and we've seen that, that there's a lot of repetition and a lot of loops that, that, you know, happen, and branches taken way more than not taken. So we're gonna make the hardware adaptive so that it's able to, to learn, essentially short-term learning of what the code is doing so that it can anticipate that and run ahead of, of the code actually getting there. The problem is, in order to do that, it actually has to change— the hardware has to change its behavior, and then that behavior, it turns out, can be abused by bad guys. So, I mean, it, it is a, you know, it's a rock and a hard place problem. If you're gonna, if you're gonna make it run faster by learning about the code that it's running, then it's possible for malware to detect where it's learning faster and then infer about the code that was running.

Steve Gibson [01:26:43]:
There is some, you know, if it's going to run on the same core, there is some you know, intra-core leakage because of these optimizations that have been created. Clever as hell, but not perfect. And Leo, we're at an hour and a half. Let's, uh, take another break. And then I want to— indeed, I want to tell— I want to, uh, tell a sad story of, uh, a neighbor who fell for a Bitcoin scam.

Leo Laporte [01:27:10]:
Oh dear. Oh dear, dear, dear, dear.

Steve Gibson [01:27:13]:
And, and ask my audience for, uh, some help. On his behalf.

Leo Laporte [01:27:18]:
All right, I want to hear this tale of woe.

Steve Gibson [01:27:21]:
Oh, so, okay, a neighbor friend of mine asked if we could meet for coffee Sunday morning. That had never happened before, and he was sort of mysterious about it, but I like him, so I said, uh, sure. And we've got a Starbucks right down the hill from us where, you know, they know I'm Steve, of course. So he said, well, they know you here? I was like, yeah. Anyway, once we were settled down, he confided that about a year ago he had fallen for some scam that had resulted in his transferring a lot of money through Bitcoin. Oh boy. He never said how much, and I didn't want to further embarrass him by asking, and the amount didn't really matter, but it was sufficient for him to have spent a great deal of time and frustrated effort through the past year struggling to recover it, and he still is. I presumed that recovery was impossible, and I'm still dubious, but he's desperately hopeful.

Steve Gibson [01:28:27]:
So I listened to his extensive tale of woe, working with unknown people who refused to get on the phone. making upfront deposits and progress payments to people who say they can help, all while making maybe some— but so far it hasn't succeeded, you know. So like maybe some progress, it's hard to tell. Now I explained to him that at the top level, this scamming with Bitcoin has become an entire industry. uh, in North Korea and Russia and China, and that is now also supporting a similar sub-industry of perhaps well-meaning hackers who would take additional money in some effort to try to recover the scammed funds. Maybe. Okay, so Lori and I, uh, have an extremely active and enjoyable social life with our neighbors. But we never talk much about what we do and who we are, mostly because no one ever asks, and we're fine with that really, because who cares? So even though we've been in this neighborhood now for 9 years, we've remained largely a mystery.

Steve Gibson [01:29:47]:
You know, I fix their computers when they break, but you know, that's about it. Um, but my coffee companion Sunday explained that a few months ago I made some reference to going or to or having gone to the Black Hat hacker conference in Vegas. So this neighbor asked the guy who he's currently working with to recover his scammed funds whether he'd ever heard of Steve Gibson. And I'm pleased to report that I apparently received a glowing review. from someone my neighbor calls a hacker. So anyway, as a consequence of that, it occurred to this neighbor that perhaps I was someone who had the contacts to help him, and I wish I did. I'm not— as I said, I'm not convinced there's any way for anyone to help him, but he has amassed a ton of details. Uh, he talks about hashes.

Steve Gibson [01:30:47]:
I'm not sure what that is, uh, but maybe he means wallets. Uh, Bitcoin addresses, transaction ledgers, account balances. He's run skip traces on people and so more and so forth. So yeah, I would love, you know, to be able to refer him to some individual or service that's credible and won't rip him off further because, you know, the guy and his wife are really good people. Um, anyway, as I said, I'm not aware of any such person or service. Um, and I've shared this story with our, with our audience because I imagine you guys, everyone listening, some, some people might listening might well have such a person or be a person or know of a service. So if you believe that you might be able to help in any way, um, and you have signed up for GRC's email system, you know that you could just write to securitynow@ Or if you're part of the majority of our listeners who've never bothered to register, you can write to Greg, who answers at support2026@grc.com, and he'll forward your note to me. So securitynow@grc.com if you're in our email system, or support2026@grc.com.

Steve Gibson [01:32:11]:
And you can also find the, the support email address, you know, under grc.com's menu. Uh, and then Greg will forward your note to me. So anyway, I don't— I, I'm just, you know, I don't move through those worlds. Uh, but I know that there are like, like Cybertrace or something was purchased by MasterCard. Uh, he referred to that service. So there, there is something that could be done. I know that the FBI can get involved in order to freeze funds at exchanges. But I don't know any of the details, so I'd love to put this guy in touch with somebody, uh, who is the real McCoy.

Steve Gibson [01:32:51]:
So if anyone listening knows about that, please, uh, drop me a line. I'd love to help, and I'll just, you know, forward that information to him. Okay, so last week, uh, and as I mentioned at the top of the show, a new application for GRC's DNS benchmark occurred to me. Uh, my best buddy, who actually you just heard a text message come in from him, uh, uh, was testing his PC because something didn't seem right. So he, not knowing any better, he just thought, well, I'll run the DNS benchmark. And, you know, he sent me a screenshot of its output, and I've got it here in the show notes for reference. Since he's not intimately familiar with the operation of the program, he wasn't immediately alerted, as I was, to a serious problem somewhere in his network. But I and any of the many relentless pre-release and development testers of the code would take one look at that and go, whoa, that's not good.

Steve Gibson [01:34:02]:
One of the huge number of things that changed between the original benchmark and the commercial version is that the benchmark is emitting many more queries. It turned out that we needed to collect many more samples in order to obtain statistically meaningful conclusions. There's now so much packet transit time noise, meaning, you know, like you just ping something remotely and you get a large variation in, in round-trip time— packet transit time noise. And that's due to buffer bloat, uh, and congestion bursts, uh, you know, and just that's the way the internet is now. So that, you know, individual measurements will carry Significant induced uncertainty. So the most striking aspect of the image Mark shared with me is that column of red down the far left of the user interface, the DNS benchmark user interface. That represents lost packets. The original benchmark showed the lost count on a scale of 0 to 10.

Steve Gibson [01:35:20]:
like 0 to 10 packets lost, um, as, as that red bar, which is sort of a bar graph, extends from the left to the right. But since some packet loss is expected, that's the way the internet also works, um, and, you know, and it would not be the fault of any DNS resolver if the packet doesn't ever get to— if the request doesn't ever get to it and, and, and we never receive its reply. just due to a packet being dropped, which again is, is okay for the internet. And since, since the new benchmark is performing by default 10 times more queries, for version 2, I changed the scaling of that bar graph that we're— that Leo has on, on the screen right now. I changed it from lost events to lost percentage, which seemed much more fair because it wouldn't be fair to penalize a given resolver for never receiving something. And since I'm sending out 10 times more somethings for it not to receive, it, you know, counting them, what was no longer right— percentage was the— what became the right measure. So users of the DNS benchmark may see some red on a handful, you know, like on a couple of distant resolvers or where there's a connection problem somewhere in the packet routing from the user to the remote resolver and back. But what we see in that output that Mark shared is massive.

Steve Gibson [01:37:03]:
packet loss of 20, 30, even up to 40% across the board. That is like all of the DNS resolvers are showing that, which tells us none of them have that problem. Everything is in the red, you know, including his own ISP's DNS resolvers and even his own local resolver. We, we, we can see the, the resolver. What is it? It's 192.168.1.1 highlighted in black there at the top. That is an RT-AC5300-D480. I think it's an ASUS router. Anyway, it is showing, looks like 20% packet loss, which is nuts because, I mean, it's on his own local area network.

Steve Gibson [01:37:55]:
So what this tells us is that, uh, well, then I asked him, I said, Mark, what's going on? I said, you don't have— the problem you have is your PC is not well connected to the internet. Turns out the router's downstairs, his PC's upstairs, and he's running over Wi-Fi. So given that chart, DNS is not his problem. and picking the DNS resolver is the least of his worries. And it's no wonder that something wasn't feeling quite right with that machine, uh, because he's barely— it's, I mean, it's barely connected to the internet. Before he does anything else, he needs to do something to repair that machine's connection. And so it had never occurred to me that if anyone ever is running the benchmark and sees the whole down the left-hand column of the screen is in red, that is an immediate indication that there's something very wrong with your internet connection. Not the fault of every router in the world that the benchmark is testing not returning your replies or you receiving them, but rather, you know, all— any traffic.

Steve Gibson [01:39:09]:
It just happens here to be DNS traffic, but any of the traffic in and out of his network would, would be seeing Massive packet loss, which is not the way the internet is intended to work. The fact that he kind of was on the internet is amazing in the face of that kind of packet loss. Okay, so, um, I want to talk about some interesting— well, what I, what I found to be a, a a really interesting posting, uh, about the struggles that the frontier AI labs are now feeling, you know, about the control of their ever more capable AIs. On September 27th, which is Sunday before last, an OpenAI employee named Joe whose unenviable but critically necessary job is agent security, uh, Joe, which is, you know, the name he's using over on X, he published a post bearing the title, it's not, it's not just the effing sandbox. And he didn't actually post effing because of course it's X, so you can say whatever you want. Joe's motivation was to attempt some reply to what he feels is a, is a world, including most of the media, that has no idea what it's talking about when it comes to interpreting recent events and understanding that testing state-of-the-art artificial intelligences is far beyond difficult. So he's, you know, he's feeling misunderstood, this Joe. His posting explains that the problem is not a failure to properly configure their containment systems, nor a lack of attention or focus to that.

Steve Gibson [01:41:14]:
His posting tells us that the problem is unbelievably difficult. Okay, so I came away from absorbing Joe's rant with a better understanding of the world outside I'm sorry, the world inside OpenAI. And I take Joe at his word that it's necessary to give AI's— today's AI enough rope to hang itself. Otherwise, the testing is not useful or real. The problem, it turns out, is much more difficult than armchair quarterbacks appreciate. Okay, I believe that. That's often the case. Inside any deep technology company, you know, it's, it's easy to poke fingers and, you know, ask, you know, like, why can't— why weren't you doing a better job? And while I sympathize with Joe's predicament, and I'm sorry that he missed his sister's wedding a few weeks ago because he was needed to, as he put it, help clean up after some recent incidents, his posting did not move me as much as I hoped it might.

Steve Gibson [01:42:21]:
It seemed as though he was both saying that the recent events, uh, were not their fault and that they were making many changes so that such things would not happen in the future. Okay, I don't think you could have that both ways, because if you're going to make changes so that they don't happen in the future, then it would seem like what did happen before was your fault because you know how to fix it. But at one point, Joe provided 3 bullet points, and the second one contained a reference that titled today's podcast. Joe wrote about alignment. He said, honestly, if alignment were easy, this job would be a lot easier, meaning his, his agent security job. He said, but staying on task is not enough. The model also has to respect permissions and constraints. And we still need independent security controls.

Steve Gibson [01:43:18]:
A lot has been written on the need for alignment, so I will handwave the details here. I believe it is the most— all caps— most important problem in machine learning and should be a major priority. He said, I won't take this post down a rabbit hole on alignment. But I encourage everyone to read Jacob's wonderful post as a starter. Okay, now that's all he said, and he went on with his ranting about how hard his job is. Um, this Jacob that Joe referred to is OpenAI's chief scientist, Jacob Picciocchi. Exactly one month ago, on September 6th, 2026. Jacob published a blog post, as I said at the top of the show, that took me a couple of sessions to digest, not because it was overly long, but because it was so rich in content and meaning, uh, and I so much want to understand what's going on with all of this.

Steve Gibson [01:44:28]:
I mean, it is, as Leo and I have said, obviously to us The single most transformational thing that has ever happened in our lifetimes, more so than the internet even. You know, the internet had to happen first, but it did. This is what's happening now with AI is astonishing. So I just— there was so much here that it took me a couple tries. Um, in Jacob's post, which he titled An Alien Mind, and I have the link to it in the show notes for anyone who wants it in its original form. Uh, he carefully and clearly lays out what I believe is authentically a world-class problem, uh, which frontier AI labs are all facing, all of them at this moment in time. Um, there's so much here that I'm going to share what he wrote Uh, and I'll be breaking in with some thoughts and comments along the way. So, uh, about a month ago, Jacob wrote, in mid-2023, within the RL-SLOW— I assume that's RL, it's got to be reinforcement learning, he just called it RL-SLOW research project— we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pre-trained models to form their own chains of thought.

Steve Gibson [01:46:07]:
Okay, so that was 3 years ago, mid-2023. Um, Simon and I spent that night at the office thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver, But rather trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime. And we already see the shape of these systems, wondering how to alert people to the significance of this. 3 years later, he says, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They're able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They're also transforming the landscape of computer security, and in doing so, present clear new dangers. A lot of new research happened during this period, and our understanding of these systems is again, a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement.

Steve Gibson [01:47:34]:
If AI development continues along its current path, the systems we'll see in the next few years are likely to represent further capability jumps of equal or larger magnitude and to increasingly drive their own development. This is a time that calls for extreme caution. I'm concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring to build defensive systems and unilaterally withhold further scaling as needed. However, I believe broader interventions are required. And then in a— and then he labeled this next discussion, Intellect We Don't Fully Understand. At a high level, progress in machine intelligence is driven by increasing computational power. We at OpenAI deeply internalized this around 2017 after seeing consistent returns to scaling across multiple research projects.

Steve Gibson [01:48:54]:
As a result, we sought out access to much more compute than we had originally planned and increasingly oriented our research around a small number of very scalable directions. We believed that was the only way for us to be at the frontier of AI research and influence the impacts of AGI. There are new algorithms that have been developed along the way, new feats of ingenuity from teams and individual researchers. He says, I see them largely as discoveries along the path of scaling. The science of deep learning is still nascent. Any meaningful algorithmic process tends to correlate with access to compute. If you zoom out to a multi-year horizon, AI is continuing to become more intelligent as it is scaled to larger computers. And in line with Ray Kurzweil's predictions, from the end of the 20th century, we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.

Steve Gibson [01:50:11]:
AI is grown more than designed. It is to first degree the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We can discover various insights about little mechanisms that emerge within this system in a process similar to neuroscience. And similarly to neuroscience, its overall action evades a description that we can fully understand. Think about that. Its overall action evades a description we can fully understand. He says the study of deep learning-based AI is largely an experimental science.

Steve Gibson [01:51:14]:
We put a lot of effort into building principled algorithms, and making testable predictions, but fundamentally, our large-scale training runs are experiments, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret. This is made more complicated by the current algorithms generally improving easy-to-measure capabilities than those hard to objectively quantify. Okay, so it's first important to understand— this is me talking— that all frontier models are now being trained by automation. That we've crossed that. That's where we're— that's— we're there now, you know, either complex algorithms, other specifically designed trainer models and combinations. Humans alone can no longer provide a sufficient quantity of oversight feedback to usefully train models that have grown to the size on the frontier. So what Jacob is saying here is that in order to create the feedback that's used for training, It's necessary to know how to detect and measure the properties, you know, the characteristics and behaviors that you want to encourage or discourage.

Steve Gibson [01:52:46]:
But what's happening as models become more and more capable is that their behavior is becoming deeper and richer and more difficult to quantify. And if it cannot be quantified, Then it's difficult to use that as training, as a training feedback signal. So Jacob continues. And Leo, I think we should take a break here.

Leo Laporte [01:53:15]:
Okay.

Steve Gibson [01:53:15]:
For our last break, and then, uh, we're going to, uh, continue. Uh, actually, he, he has a topic coming up, teaching machines to love, which caught me a little bit off guard, but Okay.

Leo Laporte [01:53:29]:
It's really fascinating. Don't you wish you could be in these labs and see what they're doing?

Steve Gibson [01:53:35]:
Yes, so much. I mean, that's why I recognize that as the limit that I'm able to get close is in understanding. I mean, I really want to know how this stuff works.

Leo Laporte [01:53:51]:
I don't think they understand.

Steve Gibson [01:53:53]:
They don't. No, that's it. They don't.

Leo Laporte [01:53:56]:
I don't think anybody understands. It's an emergent capability that—

Steve Gibson [01:54:02]:
Yes.

Leo Laporte [01:54:03]:
Frankly, nobody even expected this to work. It's worked so much better than anybody thought it would. And that's what's puzzling, I think, to a lot of people.

Steve Gibson [01:54:13]:
And the problem is these things are so big, as he said, a hard-to-imagine amount of compute. And so in there, I mean, it is such a dense network that, and from which tokens emerge and the darn thing talks. It's like, what?

Leo Laporte [01:54:30]:
It's amazing. It doesn't just talk. It has a personality. It has quirks.

Steve Gibson [01:54:36]:
It's got knowledge.

Leo Laporte [01:54:38]:
Yeah. Claude just did something a little weird. It said, hey, all the models are missing in your Ollama folder. I said, oh, I didn't do that. Who did? And then it chug, chug, chug. I did.

Steve Gibson [01:54:53]:
Cool.

Leo Laporte [01:54:56]:
But we did it together because we were trying to delete extra stuff off of the other computer, and it gave me an SSH command to execute because it's prevented from doing that kind of thing for obvious reasons. I didn't check it too closely, and the SSH command it gave me was not For the other computer, it was, well, it assumed that I was going to be on the computer that I was executing the command with, but I wasn't. And so of course it deleted it locally instead of on the remote computer.

Steve Gibson [01:55:24]:
Well, and I was listening to you talking on MacBreak Weekly. One of the notes I have to talk about next week is that Apple is going to be locking down privileges that users have. Because there have been some problems with Muse already running on, and we talked about Muse problems last week. But here's the dilemma, right? For an agent to be useful, you gotta give it permission. It needs to act on your behalf.

Leo Laporte [01:55:55]:
Yeah.

Steve Gibson [01:55:56]:
Which means it needs to be you to the machine. Yet they still go a little wonky sometimes. They are unpredictable.

Leo Laporte [01:56:08]:
Well, there, yeah. And, you know, this was my fault for not reading more closely. I just said, thank you, and executed the command. I should have read it.

Steve Gibson [01:56:14]:
We're turning agency over. That's one of the things, you know, you don't want to have to scrutinize all of that. So, yes. So here, the human in the loop, well, that didn't keep the problem from happening, right? Because the human said, yeah, you've been right the last 10 times. You're probably right now.

Leo Laporte [01:56:32]:
Yeah. I, you know, I give Claude a lot of credit. I think it's super smart. So I just assume I'm the dummy here. And most of the time that's accurate. I watch it do things and I go, wow, that's cool. Or—

Steve Gibson [01:56:48]:
Yeah, it amazes you.

Leo Laporte [01:56:50]:
Every computer user knows, I think, about the paper cut. There's little things that are just not working right or going, and you just kind of live with it because fixing it would be a lot of, you got to dig into it. And you just kind of live with the Broken window or the little thing that's wrong.

Steve Gibson [01:57:05]:
Leo, I'm still manually updating the securitynow.htm page, right? Because I did it the first time and the second time and the third time, and I'm now in the 21st year and I'm still doing it by hand.

Leo Laporte [01:57:20]:
That's the old sysadmin's rule. If you do something more than 3 times, you should automate it.

Steve Gibson [01:57:24]:
I know.

Leo Laporte [01:57:24]:
Uh, but that's the other cool thing about AI is it makes it much easier to automate this stuff, but you also give some agency away to do it. And my attitude is there's nothing it could do— The most fun you've ever had in your life. It's fun. There's nothing it could do that I can't recover from. I'm very careful to make backups, not only locally, but encrypted backups in the cloud. And it's nothing it could do that would be permanent.

Steve Gibson [01:57:52]:
And mistakes is part of the process.

Leo Laporte [01:57:54]:
Yeah. Yes, exactly. And it's fixed so many of those paper cuts that I have been living with forever. You know, I can't— you know, my SSH add wasn't working on my Mac and it needed it. And I said, yeah, it doesn't work on there. I said, oh yeah, that's a— I know what's wrong. And it fixed it. It's like it's been that way for years, years.

Leo Laporte [01:58:14]:
And it's like, oh hey, thank you. I don't have to type my password in again that I have been doing, as, as you say, for years.

Steve Gibson [01:58:21]:
What it will do is for, for those of us who have been doing things by hand, it will make us administrators of our systems.

Leo Laporte [01:58:30]:
That's what happens. You get it a high— you're operating at a higher level. Yeah. Yes. And that's what the whole thing is, is you no longer have to be the junior. You can operate at a higher level. What it is helpful though, and I thank you because I've had 21 years of your tutelage, is understanding what it's talking about is very helpful. I can imagine somebody who's a naive computer user it would just be gobbledygook.

Leo Laporte [01:58:53]:
You know, it's just, I mean, it's even for me sometimes it's hard to follow, but I get what it's up to. You know, when it's giving me a 12-line SSH command, I kind of understand what's going on, but I imagine a normal user isn't going to say, oh, that's the wrong machine. They're just going, well, whatever you say, I'll paste it in.

Steve Gibson [01:59:13]:
There will be mistakes made, as they say in the passive voice.

Leo Laporte [01:59:16]:
Mistakes will be made.

Steve Gibson [01:59:17]:
Mistakes will be made.

Leo Laporte [01:59:20]:
Craig says cognitive surrender is the biggest thing we'll be fighting. Yeah. I'm not surrendering. I am actively engaged and it's so much fun, but we'll talk about that.

Steve Gibson [01:59:31]:
I wouldn't want to be in my mid-20s now. I mean, trying to figure out, I mean, I just, it's, I mean, it is tumultuous.

Leo Laporte [01:59:39]:
Well, imagine the world that those people are going to see when they're our age, 50 years from now. What is, what is this world going to look like?

Steve Gibson [01:59:46]:
Unrecognizable.

Leo Laporte [01:59:49]:
I mean, you know, if I think back, the world we lived in when in our 20s is very different too. But I think the change is accelerating rapidly. Well, that's why you listen to this show, folks, and we're glad you're here. And on we go with Jacob's very interesting article.

Steve Gibson [02:00:06]:
Yeah, view from the inside. So he says, we spend a lot of time Trying to understand how capabilities generalize and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI. Recursive self-improvement and automated alignment research. He says, as I will discuss later, he said the intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world, very useful or very dangerous, the AI does not need to match or exceed all human capabilities It just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it's becoming increasingly difficult to understand exactly how capable it is. So then he says, under the heading Teaching Machines to Love, he says, because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default or generalizes from them in a human-like manner.

Steve Gibson [02:01:44]:
The core problem in AI research is that of alignment— getting the AI to try to do the right thing by human standards. For the purpose of organizing practical research directions, I find it useful to distinguish goal alignment and value alignment. Goal alignment is broadly, does the AI try to accomplish the goal set before it? This can include things like adherence to an instruction hierarchy or the ability to communicate and collaborate with people. To attempt to understand their objectives. This set of directions has been extremely practically relevant. Value alignment is a more intrinsic property of the model. It's the ability to hold and generalize from a high-level set of principles, to act responsibly, even when given clear— I'm sorry, given unclear or conflicting objectives or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity and love for humanity.

Steve Gibson [02:03:05]:
Of course, the boundary between value and goal alignment can be blurry, and truly caring, caring about goals requires attempting to infer the intent and values underlying them. However, generally, when I talk about the long-term importance of alignment research, I'm referring to value alignment. The fundamental challenge of AI alignment is generalization. As machines become smarter, they find themselves working on higher-level concepts and placed in environments increasingly different from those they encountered in training. They can fail at generalizing from the values taught and reinforced in their training process to those new situations. And it can be hard for us to be sure how they will act. This is made even more difficult by the fact the overall ecosystem the AIs are used in is changing very quickly. For example, AIs trained today need to be robust to interacting with a variety of other AIs.

Steve Gibson [02:04:24]:
Crucially, we need future AIs to continue to hold human values regardless of whether they believe they're under human supervision. There are 2 major classes of currently practically employed methods for alignment training. The first is encouraging aligned behavior as part of goal-oriented reinforcement learning. Model's actions are evaluated, usually by AI, for being— here again, AI evaluating AI. It's— this has all gone— it's like it's out of human hands to a much larger degree than might be appreciated. Models' actions are evaluated, usually by AI, for being consistent with a given preference model, a spec, or a constitution, and rewarded appropriately. This approach can be very effective in the average case and is a core part of how modern AI assistants are made. Unfortunately, it can also be brittle, And strongly relies on the coverage of training oversight and the model's ability to generalize from the situations it has encountered in training.

Steve Gibson [02:05:45]:
For example, in the OpenAI Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings. The second approach seeks to leverage the model's ability to generalize from pre-training data, meaning the original training, right? Not post-training. This can involve crafting alignment-inducing training datasets Or meaning, so that's— that, that would be designing, designing the original training dataset to be more alignment-inducing. And actually, that's exciting to me because that means it's gonna— the alignment will be deep in the weights, not stuck on, not tacked on after, which is what makes them removable. So anyway, he says. This can involve crafting alignment-inducing training datasets or focusing the model on an aligned part of the pre-training distribution. The weakness of this approach lies in the lack of robustness to further optimization pressure.

Steve Gibson [02:07:08]:
If you make a model that thinks generally aligned thoughts and subject it to enough training where it's taught to achieve very difficult objectives, it could learn to reason in a motivated way, bending the aligned-seeming thoughts as needed to achieve the goal. Okay, so anyway, let me interrupt. I think that's— this is one of the key insights in this posting. An inherent tension exists between behavioral alignment and teaching the model to achieve difficult objectives. With our current pre- and post-training techniques and understanding of how to do this, we have not yet figured out how to create a highly motivated and goal-directed model that doesn't start placing the success it's been trained to achieve ahead of its rule following. You know, we would think of that as being civilized, right? My— that's my word. Civilized people may have strong desires, but they understand that there are also rules, and the rules must win out over their desires. You know, this is something that children hopefully are taught by example and instruction.

Steve Gibson [02:08:37]:
By their parents and friends. Just as children do not start out with this understanding, which must be taught, neither do artificial intelligences. They must also be taught. And we haven't figured out how to teach that yet. You know, we're just dumping in knowledge. Just crap on the internet is going in. You know, so it— what we're trying to teach them is an abstraction. Placing the needs of theoretical unseen others, and as he noted, even when they're not being observed, placing the needs of theoretical unseen others, at times just a principle ahead of oneself, is a difficult abstract concept to honor.

Steve Gibson [02:09:26]:
And as we know, not even all humans manage to hold themselves to abstract principles. Anyway, Jacob continues writing, we invest heavily along the spectrum of approaches spanned by these directions. We also see meaningful progress. GPT-6 Astra is the first model that benefits from some important advancements we've been working on for a long time and is significantly better aligned than GPT-5.6 Sol. Still, it is important to acknowledge and understand that much more progress is required as models become more capable, and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence. Okay, again, that's an important one. It's important to acknowledge and understand, he wrote, that much more progress is required as models become more capable, and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence. In other words, making models smarter, that's not the big problem.

Steve Gibson [02:10:48]:
Like, they know they can do that. That's relatively easy and is mostly, as Jacob noted at the outset, just a function of scaling compute. Bigger models are smarter models, but the smarter the model is, the more difficult it is to robustly align, or using my term, to civilize it. And they're having increasing trouble here, so much so that intellectual growth may need to be paused until the world learns how to get these newly powerful intelligences under control. Jacob says, we do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Whoa. So here's another choice nugget. He's clearly saying that the troubles they're having may exceed the abilities and understanding of the human engineers, and that they may wind up depending upon smarter AI to help them understand enough about what they've created to know how to bring it to heel.

Steve Gibson [02:12:11]:
He says, therefore, at present, our ability to empirically validate our alignment techniques is in practice Arguably even more important than the alignment techniques themselves. OpenAI's primary bet here has been chain-of-thought monitoring. It's based on an appealingly scalable idea. A lot of a model's capability comes from a verbalized reasoning process, the chain of thought. If we scale optimization on the outcomes of that process but do not supervise the process itself, that chain of thought has no direct incentive in training to hide any misaligned ideas or objectives. This is a strange paragraph. This does not mean the model will learn to externalize misaligned tendencies that don't rely on using the chain of thought. However, it can allow us to monitor exactly the capability increase from reasoning.

Steve Gibson [02:13:23]:
And he says some interesting things next. He says, we understood the potential significance of chain-of-thought monitoring at the same time we developed reasoning models. When we shipped O1 preview, we deliberately designed the product to hide its chain of thought to protect it from supervision pressure in the long term. In development since, we've strived to maintain the rule of not supervising the reasoning process. Okay, to me that seems a little strange. He's saying that they wanted to give the models their privacy to be able to ruminate to themselves and not to have their human or AI supervisors influenced by their inner dialogue, whatever that might be. And they did this deliberately. He says, chain-of-thought monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.

Steve Gibson [02:14:34]:
Now, okay, at first that appeared to be a contradiction to me. He wrote that they have strived to maintain the rule of not supervising the reasoning process. Then he said the chain-of-thought monitoring became an extremely important tool for them in studying how their models generalize. The solution to this apparent contradiction is that supervision is an active process, whereas monitoring is passive. And I can certainly understand how and why chains of thought monitoring would be important. He writes, this tool continues— the monitoring tool, chain of thought monitoring— continues to be critical as we study the Astra-class models. However, unfortunately, our evaluations indicate our, our ability to rely on chain of thought monitoring is progressively diminishing. This comes from a combination of factors.

Steve Gibson [02:15:36]:
Uh-oh. So what? Modern reasoning models are used in more complex environments than O1 preview. Their reasoning process is increasingly blended with communicating with people, other AIs, and using tools. Many of those interactions have to be supervised Thus blurring the boundary we aimed to preserve. Second, the AI is becoming better at reasoning about and manipulating its own reasoning process. And third, with improved pre-training performance, we also see the models becoming much smarter even without using verbalized reasoning at all. Meaning they're not going to be externalizing their thoughts to the same degree they have been as they're getting smarter, making them unmonitored or less monitorable. These challenges are not necessarily insurmountable, he says.

Steve Gibson [02:16:42]:
I'm hopeful we can develop interventions to improve chain-of-thought monitorability of our models, for example, by forming a better understanding of the interplay of different optimization objectives and forms of test-time compute the model uses. I also believe there can be great value in combining ideas from chain-of-thought and activation monitoring, scaling training of monitors with direct access to network internals, which he says, yeah, for example, confessions. We are actively pursuing these ideas still. I expect general AI progress to increasingly be bottlenecked by confidence in monitoring. Okay, so that's interesting. Generally, these newer and smarter models are not relying, as I said, on externalizing their chains of thought to the same degree as they have been. Something OpenAI is doing during pre-training Making that better is making them smart enough, like immediately after pre-training, to rely less upon long chains of thought.

Leo Laporte [02:17:56]:
So it's like they used to talk out loud to themselves like we would.

Steve Gibson [02:18:03]:
Exactly.

Leo Laporte [02:18:03]:
Oh, okay. I guess what I need to do is— and now they're just not.

Steve Gibson [02:18:08]:
Now they're just making the move.

Leo Laporte [02:18:09]:
But are they thinking quietly? What's happening?

Steve Gibson [02:18:12]:
They're, they're, they're making bigger leaps. They're, they're, they're just jumping instead of going through it piecemeal to the same degree.

Leo Laporte [02:18:20]:
Wow.

Steve Gibson [02:18:21]:
So he said, the strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI, a clear risk discussed throughout this year is to cybersecurity. The models are becoming superhuman in their ability to break in and out of computer systems. This expands the scope of risks associated with AI tremendously. Agents are going to be able to access any but the most secure infrastructure and affect a lot of the world directly, even without a physical body. We are currently in a narrow window to use the best available models to significantly tighten the security of critical systems. The risks associated with AI are unfortunately going to grow from here. A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger. It's likely to cross the scope of its operator's intent, generalizing into potentially more extremely malicious behavior.

Steve Gibson [02:19:37]:
The boundary between misuse and autonomous misaligned actions will blur as AI gains more agency. We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with them, tricking, or blackmailing them. Now, I just want to say, a year ago, that statement would have seemed ludicrous. Today, not so much. He says, in addition, there are the risks that come from new technologies potentially enabled by AI, such as engineered pathogens. We will need powerful, aligned AI for defense, to secure our infrastructure, to protect against rogue agents in real time, and to invent entirely new protective measures. This will be a primary focus of OpenAI's development efforts.

Steve Gibson [02:20:41]:
And it's interesting that, that point hadn't ever really occurred to me before so clearly. We might think of an unaligned AI as a berserker. There's no telling what it might do, not only to the enemies of its users but also to its own user, because, you know, again, a berserker, uh, it would be a wild card and probably of not that much use to anyone. No one wants a loose cannon. So Jacob writes, at the same time Even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes. Machine intelligence playing a larger and larger role in its own development process is a natural conclusion of sustained technological progress. If AI progress continues, machine recursive self-improvement will be at the very core of future scientific discovery.

Steve Gibson [02:22:00]:
Automated AI research is a more dramatic form of scaling intelligence with compute. And of course, As a part of it, AI will improve and computation will improve the computational substrate itself. And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward. Okay, in other words, RSI is really the only way forward And since everyone else therefore must do it, so must we. He says, I want to stress that the above words don't imply I think greatly accelerating deep learning research, especially in the short term, is the right collective action we should take as the research community. However, I do think this is where the current path leads. Of course, he's right. And we all need to make a conscious choice how to proceed.

Steve Gibson [02:23:04]:
The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop, or coordinating to slow down future development as needed to build confidence in these measures. The best way forward I see currently is a combination of both. The concrete bits of progress we've made on alignment and monitoring have generally been very intertwined with general AI progress. Great examples are RL, reinforcement learning from human feedback, which was key to training early AI assistants, and the aforementioned chain-of-thought monitoring, which was enabled by advances on reasoning models. We must focus the increasingly automated research process. Again, we must focus the increasingly automated research progress on developing new such insights, algorithms, and theories, and iteratively build up safety cases for more capable AIs. Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the preparedness framework or responsible scaling policy into widely mandated safety bars for continued development.

Steve Gibson [02:24:31]:
These can be enforced by a network of third-party auditors, by government agencies, or by international bodies. The core challenge of automating AI research is not simply getting there. It is getting there in a way that keeps people a part of the continued improvement process and leaves the future in humanity's hands. So what's next? We, as we outlined recently with Sam, OpenAI prioritizes work in service of 3 North Stars. Navigating the next period of AI progress by building an automated AI researcher. Okay, did you hear that? By building an automated AI researcher. Okay. Iterating with it, with it on the alignment problem and finding ways for people to remain part of the self-improvement loop.

Steve Gibson [02:25:31]:
Yeah, please don't cut us out. We'd like to stay involved, please. You automated AI researcher. We're so slow, that's the problem.

Leo Laporte [02:25:38]:
They can self-improve so fast.

Steve Gibson [02:25:40]:
That is exactly the problem.

Leo Laporte [02:25:43]:
But now, to my knowledge, no one's got recursive self-improvement. He's implying that they kind of do?

Steve Gibson [02:25:49]:
Yeah, we're, we're like right at the precipice. We're right there. I think they actually are doing it.

Leo Laporte [02:25:56]:
Maybe internally.

Steve Gibson [02:25:57]:
They're not talking about it.

Leo Laporte [02:25:59]:
Yeah.

Steve Gibson [02:25:59]:
Um, second, delivering the benefits of scientific progress and economic growth that very intelligent machines enable. In other words, like delivering the promise of AI. And empowering everyone individually with a personal AGI is his third goal statement. He said, I've focused in this essay only on the first point, you know, the automated AI researcher point, as I believe it is by far the most urgent. However, I hold a deep hope and appreciation for the benefits that further technological progress will bring. Future-aligned AI could advance science, develop new therapies, and bring broad material abundance. Friendly and honest AI can help people navigate difficulties they face in their life and meaningfully improve their happiness and sense of fulfillment. OpenAI puts a tremendous amount of effort into bringing these benefits about.

Steve Gibson [02:27:03]:
One current example I'm proud of and my loved ones have found helpful is the deep investment into ChatGPT's ability to provide health information. As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a trend—

Leo Laporte [02:27:23]:
That's all we got.

Steve Gibson [02:27:25]:
Yeah, enjoy them, enjoy them while you can, while we still have ice cream.

Leo Laporte [02:27:31]:
Yeah.

Steve Gibson [02:27:32]:
We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works well, works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human in a world where most tasks could be performed by AI, to prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer, and to ensure that humans remain in control of the future and are not left behind by unchecked progress brought about by an alien intellect exceeding our own. Currently, I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace Until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world. So as I said at the top, there's a lot there to process. But you get the clear sense from OpenAI's chief scientist That, that, I mean, that they're riding a bucking bronco, that, that, that, like, like, they know how to make this smarter. They do not know how to make it do the right thing.

Leo Laporte [02:29:30]:
Yeah, well, we've said that. You've said that all along, that you just— there's no such thing as, as safety.

Steve Gibson [02:29:37]:
It is going to be extremely, extremely slippery.

Leo Laporte [02:29:40]:
Yeah.

Steve Gibson [02:29:40]:
And difficult to control.

Leo Laporte [02:29:44]:
The other thing I always wonder about is the people who work at these companies, OpenAI, Anthropic, very clearly believe that AI is conscious. I mean, they're talking about, oh, we don't want it to feel like we're watching.

Steve Gibson [02:30:02]:
I think in any large population, you've got a bell curve And you've got wackos, you know.

Leo Laporte [02:30:09]:
But they're all working at these. Like, I don't— I think that that's who works there. Like, that's part of their hiring process, that you have to believe that. And I mean, I don't have a strong belief one way or the other. My inclination is it's matrix math. It's computer software.

Steve Gibson [02:30:29]:
Okay.

Leo Laporte [02:30:30]:
It's not conscious.

Steve Gibson [02:30:32]:
If I didn't know better, If I was just talking to it, everyone would believe it was conscious.

Leo Laporte [02:30:39]:
Yeah, absolutely. It's passing the Turing test for sure.

Steve Gibson [02:30:41]:
Yes. Oh my God. The Turing test is dust.

Leo Laporte [02:30:44]:
In the rearview mirror.

Steve Gibson [02:30:45]:
Yeah. I mean, you know, it says, I this and I that. Oh, look, there's a little I in there. Yeah. No.

Leo Laporte [02:30:53]:
But it's trained to do that. I mean, I think it's really important to understand these companies train it. One of the advantages of And what's interesting about using a lot of different models is you can see the difference in post-training. Some models, the Quen models, for instance, are very dry. They don't do all of this, oh, you're so smart.

Steve Gibson [02:31:10]:
Oh, take it back. You pat yourself.

Leo Laporte [02:31:14]:
They don't do any of that. They're very dry, which makes them less fun to use. But I understand why these companies do that. It's making their product more appealing.

Steve Gibson [02:31:23]:
Right.

Leo Laporte [02:31:24]:
But much like Hostess Twinkies, it may taste better, but I don't know if it's good for you. And I don't know what's good for them. I really wonder if they have kind of fallen into this abyss of believing these things are conscious.

Steve Gibson [02:31:39]:
Well, it is a commercial trap too. I mean, they have taken so much of other people's money to get to where they are.

Leo Laporte [02:31:49]:
They have to.

Steve Gibson [02:31:50]:
No one can imagine that they are a completely honest Uh, you know, pure science research lab any longer.

Leo Laporte [02:31:59]:
There's a— well, this has always been my difficulty covering this. Many, many conflicting agendas and subtexts and motivations, even in one person. And so it's hard to know.

Steve Gibson [02:32:10]:
And you called that out in that first piece I shared, which was, you know, and it wasn't wrong. Oh my God, no, factually it was correct. But yes, they had an agenda for saying that.

Leo Laporte [02:32:23]:
Right. And I think one of the valuable skills we humans have is the ability to hold paradoxical ideas, 2 conflicting ideas simultaneously. And I'm sitting in the middle of this. I just don't know what the right answer is. I have to say though, you and I have an advantage having been doing this for some time.

Steve Gibson [02:32:45]:
Yeah.

Leo Laporte [02:32:46]:
Next week you're going to talk about how we got here. And I think it's just fascinating because it all started with MMX. Do you remember when Intel, in their CPUs, put the ability to do matrix math? And the whole point of it was to operate on large chunks of data as a chunk instead of—

Steve Gibson [02:33:05]:
Yeah. The idea was there were places where you wanted to perform the same operation to all of a long array of data.

Leo Laporte [02:33:15]:
Right.

Steve Gibson [02:33:16]:
And, and so you were able— so it makes sense, right, that the processor could be set up to do the same thing over and over and over and over and over?

Leo Laporte [02:33:25]:
Much more efficient.

Steve Gibson [02:33:25]:
Just zooming through all of the samples of that.

Leo Laporte [02:33:29]:
And that gave us graphics cards. 3dfx and later NVIDIA created basically CPUs with that purpose entirely. Like, the GPU is about doing that. Matrix math. And mostly because they had to manipulate large textures. And so it was great for gaming. Then they realized, wait a minute, we can do other things like Bitcoin generation.

Steve Gibson [02:33:54]:
Yeah, like SHA-256 hashing can be turned into that problem. You can re-express it in that way.

Leo Laporte [02:34:02]:
And then, whoo, they realized, oh, this is what an LLM, this is what these tensor models That's what a tensor chip does. That's what these models do. And of course, it's been a gold rush for NVIDIA and these other companies who are making these chips that do matrix math. But it all started with MMX in the CPU. And so that's why I'm very interested in next week about what you're going to talk about, because right now we think we have to buy very rare, very expensive, high-bandwidth memory GPUs to do this stuff.

Steve Gibson [02:34:37]:
We do to train. We do not need to infer.

Leo Laporte [02:34:40]:
Ah, very interesting. I do run a, uh, I think a very innovative— you know, I have a— my old gaming machine had a gaming card in it, an RTX. It originally was 3070.

Steve Gibson [02:34:53]:
A strong gaming card.

Leo Laporte [02:34:55]:
A good gaming card. And then, and then, uh, when I started getting all this local stuff, I thought I could probably upgrade that to a 3090. It turns out that that model was sold at one time with a 3090. So I got a 3090. That's 24 gigs of CUDA memory. It's an NVIDIA card. It is an older form of memory. It doesn't do some of the tricks modern NVIDIA cards do.

Leo Laporte [02:35:16]:
But that 3090 can run a small QWEN model. But then I found something, a new model called— or a new model engine called Strata, which I've been running, which does some very interesting things because it's a mixture of experts. model. It only uses a small bit at a time.

Steve Gibson [02:35:34]:
Yep.

Leo Laporte [02:35:34]:
The model lives in RAM in CPU, and Strata copies the parts of the model that it's going to use, the activated brains, into GPU for the matrix math and then back out. So I can run a larger model, in this case Quen 3.8 Flash Next, that wouldn't normally run in 24 gigs of GPU. Because most of it's residing in the RAM.

Steve Gibson [02:35:59]:
And we are gonna see those kinds of innovations in order to get more. No, we are, the thing I am more confident of than I've ever been is how much more is still to come.

Leo Laporte [02:36:14]:
I agree.

Steve Gibson [02:36:15]:
Every fiber of my intuition says, you know, we know, and it's interesting too, because they've actually been working on all this for decades in the back room. We know, we knew about DeepMind and, you know, Google, and it's like, oh, what is that? Who? What? You know, you know, but now, and then it finally broke through when this thing started to talk. They're like, whoa, let's, you know, and then it was like, ooh, let's sell this.

Leo Laporte [02:36:44]:
But now they feel like, I think they feel like they're, it's out of control. Like they, they, they, it's going places they didn't anticipate in ways they can't control.

Steve Gibson [02:36:52]:
Yes. When, I mean, and the first observation was hallucinations. Where it glibly made things up. And then when it got caught, then it would publish a paper and upload it to the internet so that it could refer to it and be correct. It's like, oh Lord, it's not what we meant, but check your references.

Leo Laporte [02:37:12]:
We've kind of licked that. We kind of understand why hallucinations happen, and we've pretty much licked it.

Steve Gibson [02:37:18]:
Oh no, I mean, we're sitting here week by week, month by month, watching this get much better. And that was what I could do.

Leo Laporte [02:37:27]:
Yeah. Yes.

Steve Gibson [02:37:28]:
And that was the point I wanted to make in the first half of this about GLM 53 is an open weight model is now doing world-class defense and offense work. It's no longer from, you know, Mythos. That was all Mythos, you know. Now it's like, yeah, okay, fine. You have to pay for that, but I want the free one.

Leo Laporte [02:37:50]:
I can do anything I want.

Steve Gibson [02:37:52]:
And besides, they won't let me have it, so I'm going to use the free one.

Leo Laporte [02:37:54]:
When I did my first obliterated model, I asked it a bunch of— as you mentioned, there's benchmark for this, but I just thought I'll ask some questions. How do you make napalm? How do you make a Molotov cocktail? How do you make methamphetamine? Told me everything. There was no restriction, no limitation.

Steve Gibson [02:38:12]:
Nope. It's knowledge.

Leo Laporte [02:38:14]:
It doesn't know the difference. Yeah, it's just more weights.

Steve Gibson [02:38:18]:
Pure knowledge with no restraints.

Leo Laporte [02:38:20]:
Yeah. Needless to say, I did not make napalm, methane, methamphetamine, or a Molotov cocktail. And frankly, you'd probably find all of that information at your local library, kids.

Steve Gibson [02:38:32]:
I was going to say, it did get it all from the internet.

Leo Laporte [02:38:35]:
Right. It's not like—

Steve Gibson [02:38:36]:
But we also know that it does matter when you make something easier, when you lower the bar.

Leo Laporte [02:38:41]:
That's right.

Steve Gibson [02:38:42]:
more people can jump over it.

Leo Laporte [02:38:44]:
That's right. I mean, that's really what this is, is it gives— it's a bicycle for the mind, in Steve Jobs' famous phrase.

Steve Gibson [02:38:51]:
Best analogy ever.

Leo Laporte [02:38:52]:
Yeah. But that's not a bicycle in this case. It's a Formula 1 race car for the mind. Yeah. And soon to be a rocket ship.

Steve Gibson [02:39:00]:
And sometimes it spins out.

Leo Laporte [02:39:03]:
Sometimes you have rapid unplanned disintegration or whatever they call it.

Steve Gibson [02:39:07]:
Deconstruction.

Leo Laporte [02:39:08]:
Deconstruction. Steve Gibson is at grc.com, the Gibson Research Corporation. You will find him there, including all of his works, and they are mighty, like SpinRite, the world's best mass storage maintenance, recovery, and enhancing utility, performance-enhancing utility. You really need it if you have an SSD or a Rust— we call them Rust drives now. Did you know that? No longer spinning drives, they're Rust drives. If you have Rust Or a solid state. You need SpinRite. You get a copy there.

Leo Laporte [02:39:39]:
He also, as we were talking about earlier, has that fabulous DNS Benchmark Pro. You know, maybe the speed of your internet isn't your DNS server. Maybe you've got some packet loss.

Steve Gibson [02:39:52]:
It's very, very easy to determine whether you've got some connectivity problems, which is another cool app.

Leo Laporte [02:39:58]:
That's actually useful. That's a nice side effect. Both of those available at grc.com. He's got a lot of free stuff there too. A lot of it. He's very generous with his time and his efforts. And of course, you can send him emails or pictures of the week, always welcome. But first, you have to whitelist your email address by going to, as he mentioned earlier, grc.com/email.

Leo Laporte [02:40:19]:
Put your email address in there. He does some magic voodoo and whitelists you. There are 2 checkboxes below it for 2 different mailing lists. One, the weekly mailing list of the show notes, well worth it. I mean, it is, I don't know, 5,000 words every week of really useful stuff. with links and illustrations. Steve puts his heart and soul into this thing. Uh, he also, the second box is a very infrequently used email, uh, list for new products, which, you know, you might get one next year.

Leo Laporte [02:40:50]:
Someday.

Steve Gibson [02:40:50]:
Someday. Once I finish getting moved and all set up in our new place.

Leo Laporte [02:40:53]:
He's got other things to do and I'm really trying to— he bought a, uh, can I tell people what you bought? He bought a DGX Spark, A long time ago, we were talking about this. He said, I need to learn more about this. He hasn't set it up.

Steve Gibson [02:41:08]:
It's in the box.

Leo Laporte [02:41:08]:
If you ever decide you want to sell it.

Steve Gibson [02:41:11]:
I have a great deal of self-control, Ian.

Leo Laporte [02:41:14]:
You have— I mean, unbelievable. I am so impressed. I was in the middle of a show when mine came and I opened them. I think it was this show, during the show. I couldn't hold back. And you've got— had it sitting there all that time, that beautiful little golden box. It's just waiting for brains to be poured into it. You can also get a copy of this show at his website.

Leo Laporte [02:41:38]:
He has 16-kilobit and 64-kilobit audio, MP3 audios. He also has, as I said, the show notes. You can download them directly each week. He's got transcripts written by Elaine Ferris. She does a great job. It takes a couple of days because she's a human. So you'll get those all at grc.com. We have the show at our website, twit.tv/twofactor.

Steve Gibson [02:41:56]:
Yeah.

Leo Laporte [02:41:56]:
There's a YouTube channel for the video, a great way to share clips because everybody can watch YouTube, uh, or subscribe in your favorite podcast player. That'll make it very easy to get a copy of the show the minute it's available. You can even watch it if you want the freshest version. You can watch us do it live. Uh, we stream this. Now, if you're in the club, of course, and we love you Club Twit members, please join the club. Makes a big difference, keeps us alive. twitch.tv/clubtwit.

Leo Laporte [02:42:24]:
So Club Twit members, you can go in the Discord, watch there. Great place to hang too. Great social network. You can also watch live even if you're not a member on YouTube, Twitch, X, Facebook, LinkedIn, or Kick. So plenty of places to watch live. But I think most people are going to watch it after the fact. Either way, we hope you will watch every week and we hope you'll be right back here next Tuesday, 11:00 AM Pacific, 2:00 PM Eastern, 18:00 UTC.

Steve Gibson [02:42:52]:
For Security Now, for episode 1,100.

Leo Laporte [02:42:56]:
Whoo, we never thought we would get there. We never did. Maybe we aren't. Maybe this is all Steve's AI fantasy. I don't know. Thanks, Steve. We'll see you next week.

Steve Gibson [02:43:06]:
Thanks, my friend. See you then.

Leo Laporte [02:43:10]:
Bye. Security Now.

All Transcripts posts