TWiT+ Club Shows 770 Transcript - AI User Group #20
Please be advised that this transcript is AI-generated and may not be word-for-word. Time codes refer to the approximate times in the ad-free version of the show.
Leo Laporte [00:00:00]:
This is TWiT. Hello everybody, time for the AI User Group with our new album art designed by Gemini Flash. I'm Leo Laporte and we are so glad to have you. And joining us already, we've got a bunch of people. Larry, who's having lunch, or is it dinner? It's dinner. Uh, Timothy Ingalls. Hi Timothy, welcome back. And our CEO in residence, and I forgot your name because I don't see a lower third.
Danno [00:00:31]:
Oh yeah, Dano.
Leo Laporte [00:00:32]:
Dano, that's right.
Danno [00:00:34]:
Leo, Leo, I see you, you didn't get the memo. We're supposed to dress in strict casual.
Leo Laporte [00:00:40]:
Believe it or not, this is strict casual for me. I didn't put on a fancy shirt. I did, I put it, put it in context. And I have a boo-boo on my head because earlier today I stumbled over my gym equipment. And for the— and for once the Apple Watch said, did you fall? And I said, yeah, I did, but don't call the police. I'm okay.
Timothy Ingalls [00:00:59]:
So your agents got mad and sent something after you?
Leo Laporte [00:01:02]:
Uh, yeah. The other thing I did is I said, I want you all talking to me, uh, because, uh, we're gonna have the AI user group and I want to bug them. So every turn, say something. And I, I've been using Breeze as my voice. I went back to Kokoro briefly, but then I found some room and I For the Breeze model. Turns out I can run it on the Mac in MLX. There's an MLX version. So I'm running Breeze on the Mac, and it does very good voice clones.
Leo Laporte [00:01:30]:
Uh, so I had it clone the voices from The Emperor's New Groove, and it's pretty cute. And I think I'm gonna do this. I'm gonna have— I'm gonna say things like, uh, today I want you to be the crew from Gilligan's Island. And we'll have the professor and Mary Ann Gilligan, the skipper, Thurston Howell, Lovey.
Darren Oakey [00:01:52]:
I like—
Leo Laporte [00:01:52]:
the reason I have them have— well, besides the fact that it tickles me no end to hear these silly voices, it also is helpful for me. In fact, this is why I told them to talk to me, is because they're working right now. They've been working on TwitAds, but I'm not hearing when— and they do handoffs there, you know, they— it's an agent team that's handing off. One does design, one does review, one does coding, one does cybersecurity, etc. And I, and I don't know when they've done a handoff. I don't know where they are. I have to go. I said, I don't want to look at the screen, so tell me every time you do a handoff.
Leo Laporte [00:02:25]:
Anyway, the real— I will— I'm going to step out, step back as soon as I do this. But I thought I would do this today because I was lucky half an hour ago, as I mentioned, I got Astra, the new model that everybody's been kind of— This, we think this is the model that hacked Hugging Face. So, uh, I, I didn't say go hack Hugging Face, but I did say just do a website for me, and I have not looked at it yet. So I'm going to, uh, we'll see. This is its debut. Um, why is it not—
Larry Gold (LrAu) [00:03:01]:
Are we sure it's the hard— the same harness?
Leo Laporte [00:03:04]:
Oh, I'm sure it's not the same harness. For one thing, if I don't think I could spawn 12,000 subagents right now. Why is it not, um—
Timothy Ingalls [00:03:14]:
Would you want to pay for that though?
Leo Laporte [00:03:16]:
You would know, obviously. You know, Anthony, Restream needs to be rewritten. It's now really falling apart on me.
Danno [00:03:26]:
Well, it seemed like they must have had a lot of bots just running around because they were finding the, finding the chat room. If they were sub-agents, they would have known where the chat room was, right?
Leo Laporte [00:03:38]:
Well, yeah, I mean, and as others have pointed out, you're telling me OpenAI doesn't know how to air gap a computer? But, uh, I guess if it— if we're gonna— if you're gonna, you know, have it attached to the data center where the models are, you can't really air gap it. Anyway, let me see what, uh, let's see what Astra did. I told it, do your best work. So, ah, I don't— I already don't like it. Uh, so we're now on the planet Earth. Uh, when I enter the attic, it should zoom in. Let's see what it does. Oh, that's good.
Leo Laporte [00:04:13]:
Oh, oh, I'm liking that. It's zooming into the attic. Oh, oh. I said, do a picture of the attic. These are the various devices. It does not look like this. I wish it did. Oh, so you can click the different things.
Leo Laporte [00:04:34]:
Okay, you know what, this is pretty nice. So I wonder if there's a— it says artist's interpretation.
Anthony Nielsen [00:04:47]:
Okay, okay.
Leo Laporte [00:04:51]:
Here's the various things.
Craig McFarlane (CraigM) [00:04:54]:
It's correct.
Leo Laporte [00:04:55]:
In the models that are running on it. Oh, here are the various agents. Let's see if it's got the voices right.
Darren Oakey [00:05:01]:
I keep the whole fleet moving. The right mind, the right machine, and the receipts waiting when Leo returns.
Larry Gold (LrAu) [00:05:08]:
Honestly, orchestration is the easy part.
Darren Oakey [00:05:10]:
David Spade?
Leo Laporte [00:05:10]:
Yeah, it's David Spade. This is from The Emperor's New Groove.
Darren Oakey [00:05:12]:
That takes style.
Leo Laporte [00:05:15]:
Okay, this one I apologize for. This is Wanda Sykes.
Larry Gold (LrAu) [00:05:18]:
The workshop lights on.
Leo Laporte [00:05:20]:
Bring me the knot, the machine, or the maze. We will make it legible. I just like that. Uh, this I think is ISMA.
Larry Gold (LrAu) [00:05:28]:
I found the quiet failure path before it reached production. 2 findings, 1 real risk, and a reproducible test. You are welcome.
Leo Laporte [00:05:38]:
The other thing that Breeze does that's great is you can, uh, that you can do prosody in the prompt. So I said, oh, use prosody. You know, if you're being sarcastic or whatever, put that in the prompt. So it's doing it.
Danno [00:05:50]:
Plan checked.
John Telford [00:05:51]:
Test red, then green. 12 slices landed cleanly, and I did not review my own work.
Leo Laporte [00:05:56]:
Oh yeah, it is all coming together.
Timothy Ingalls [00:06:04]:
I am Pi.
Larry Gold (LrAu) [00:06:05]:
I review and audit the work the others produce.
Leo Laporte [00:06:07]:
This is—
Larry Gold (LrAu) [00:06:07]:
and I take on side jobs when needed.
Timothy Ingalls [00:06:10]:
I am Zed.
Larry Gold (LrAu) [00:06:10]:
I review the code line by line with—
Leo Laporte [00:06:13]:
Now, I don't think Bard— I made Bard be me, but I don't think it's probably up to date.
Larry Gold (LrAu) [00:06:17]:
I read the plan, tested the claim, and found one thing worth another look.
Leo Laporte [00:06:23]:
That is why the Third Eye is here. That's the Third Eye.
Larry Gold (LrAu) [00:06:26]:
I have the signal, the source, and the machine map. Give me 5 minutes and I will bring back the part that matters.
Leo Laporte [00:06:33]:
I want them to have a different, um, voice so I know who's talking to me instead of— so they don't have to identify themselves. Little demonstration. Let's see. I don't— this is interesting. Oh. Oh.
Timothy Ingalls [00:06:46]:
Wow.
Leo Laporte [00:06:46]:
That's really nice. Okay.
John Telford [00:06:52]:
Huh.
Craig McFarlane (CraigM) [00:06:52]:
All right.
Leo Laporte [00:06:53]:
That was, uh, here's how the Breeze works. All right. I think it did a decent job. Nothing, nothing like to write home about. I do like this zoom in. I told it to do that.
Darren Oakey [00:07:08]:
You're just pushing the Glover initiative.
Leo Laporte [00:07:11]:
Yeah, I think that's pretty cool. So I don't know how many forests I burned for this, but, uh, you might start to see the fires. So Darren, I was saying, I got, I got Astra and I thought, well, let's give it just something to do pretty quickly. And this is the nicest. I have to say, I've had many models do this same kind of thing, and this is the best so far. But it looks like an AI designed it. I don't know why AIs like to do italicized and bold in the text and stuff like that. But this is, this is nice.
Leo Laporte [00:07:46]:
I like the picture of the attic it drew. The globe worked pretty well. Anyway, that's, that's all I have to say for myself. I have— I'm gonna have to refresh my screen.
Craig McFarlane (CraigM) [00:07:56]:
Hold on.
Leo Laporte [00:07:56]:
I don't know if this is gonna kick me out, but, uh, Restream is messed up.
Anthony Nielsen [00:08:02]:
Reload.
Danno [00:08:05]:
Yep, the bots have taken over.
Leo Laporte [00:08:14]:
Yeah, I had to, I had to come back, uh, but it looks okay now. Anthony Nielsen and Discord have joined. You want to add them to the stream? Yes, even if they look like nothing. Anthony, turn on your camera. What are you doing, man?
Anthony Nielsen [00:08:32]:
There he is. I have to edit This Week in Space.
Leo Laporte [00:08:35]:
Oh, are you working? All right, sorry, go, go back to work. So, uh, Darren, you're in the swimming pool today.
Darren Oakey [00:08:45]:
Yeah, I got something that randomizes my, my background for Zoom at work and stuff.
Danno [00:08:53]:
We'll let you know if we see any sharks.
Leo Laporte [00:08:56]:
So what have, uh, what have you all been up to? This is the AI User Group, which means what we do is we just get together and we shoot the breeze. So to speak, talking about what we're working on. And anybody wanna— why is—
Danno [00:09:14]:
okay.
Leo Laporte [00:09:15]:
Hi, Juan. Good to see you.
Juan Hernandez (BlindWiz) [00:09:17]:
Hey, how are you guys?
Leo Laporte [00:09:19]:
Nice to see you. I finally caught up with your dual sparks. You know, Darren said— he warned me. He said they're kind of— the spark is kind of slow. And you really do get spoiled when you're using Frontier models at how fast they are. And that is one of the lessons I've learned. I keep going, trying other models, hoping to get more speed. DeepSeek Flash is faster than GLM, but I like the— I like the brains on GLM, so I'm just using GLM.
Danno [00:09:52]:
We have 7 sparks on screen right now then.
Leo Laporte [00:09:55]:
7 sparks. So, wow, you've got one, Dano?
Danno [00:10:00]:
I've got 2.
Leo Laporte [00:10:01]:
You've got 2? No way, more than the 7.
Danno [00:10:04]:
Uh, Darren's got 1 and Juan has 2 and I have 2.
Leo Laporte [00:10:07]:
That's 7. Larry, do you have one yet?
Larry Gold (LrAu) [00:10:11]:
No, I'm still using— I told you I'm using the dual video card machine. I, I, I keep— that's right, I keep wanting to click on the Mac, but I think like, yeah, I want to wait for the 512. If you're going to spend 13 grand, you might as well spend 18.
Leo Laporte [00:10:25]:
Might as well.
Larry Gold (LrAu) [00:10:26]:
Yeah, of course. Who needs a new car? Who needs to pay for kids' college graduate, you know, uh, you know, med school or anything like that?
Leo Laporte [00:10:32]:
I'm thinking now that it might be smart. Mark Gurman says the M7's coming out in '28 and that they're going to aim for a terabyte and a half of memory. I'm thinking maybe save my pennies. By '28, I should have $40,000 or $50,000 I could spend on a terabyte and a half.
Danno [00:10:53]:
Uh, I think, you know, running the Sparks separately for me, I like, I like the way the smaller ones run. Yeah, like the 30Bs.
Leo Laporte [00:11:03]:
There's a lot to be said for that.
Juan Hernandez (BlindWiz) [00:11:05]:
Yeah.
Danno [00:11:05]:
Run one on each.
Leo Laporte [00:11:08]:
Well, especially for, um, at least for Hermes, but I think it's probably true for other models, uh, they never just open one stream, right? So it's nice to have multiple models. I have Quen running on a 3090. 27B running on 3090. So it's an auxiliary model for vision and other stuff. I put Kimi Linear, which is an 8B old model, uh, for other auxiliary tasks. I mean, there's, you know, Hermes has like 10 different auxiliary tasks. Like compression is running on the Quen, uh, session titles. Somebody has to do that.
Leo Laporte [00:11:49]:
So Kimi does that. So, and I find that every, you know, one of the problems of course with the Sparks is, is as you open more sessions, it really bogs down. You know, I can do maybe 2 at a reasonable, you know, 30 or 40 tokens, and after that it's just unusable. So splitting them actually seems like not a bad idea.
Darren Oakey [00:12:10]:
But this is also why I wrote that Avastar thing, which is, um, it queues stuff up and it controls how many things, like I say, this needs 90% of the machine, this needs 20%. And so it just puts everything in a queue and it guarantees it's done, but make sure because especially I find with the models so, so big and if they start thrashing against each other, like one will load in the model and then the other one will do a bit and they unload it and the load takes so long that—
Leo Laporte [00:12:40]:
Yeah, it's like 8 minutes to load GLM.
Darren Oakey [00:12:43]:
Yeah, so if you can queue them and say, do all of this one and then all of this one, it's much faster.
Danno [00:12:49]:
Is that what you call native? That one, I was looking for that server or AI server you were calling native and I wasn't finding anything.
Darren Oakey [00:12:59]:
That's not for the Spark, that's for Mac. It's highly optimized for— it's like an Ollama competitor for the Mac. But it's native without an E. But I've actually moved off from it because now I want to run more models simultaneously. On the Macs, I've gone down to— I don't run any LLMs on the Sparks because they're all about the CUDA and all about playing with fun stuff like video generation and training and all that stuff. But on the Macs, I run— What do you run? Nimatron and Ornith. Ornith for bits of coding and Nimatron for just summarizing titles and all of the other things, just because it seems to be the fastest general model that works really well.
Leo Laporte [00:13:48]:
It's very snappy. Yeah.
Craig McFarlane (CraigM) [00:13:51]:
Yeah.
Leo Laporte [00:13:51]:
You actually code with Ornith?
Darren Oakey [00:13:54]:
Yes. But the thing is, for all of these models, for Local, for me, what Local is about is it's about doing stuff. It's about the stuff that I build into my processes. And I'm doing— I don't know if you saw, but, um, on last weekend at work, I, um, had something for my work harness, um, that was just titling some things, and it, it caused some errors. But of course, at work, we were all, all using Bedrock, and we're using paid things. And, um, it just went into a retry loop and it ran 29,000 times before someone noticed it and it cost me $1,400. So, um, you've just got to be so careful with anything that's automatically running. And that's why all of these local things like the, for instance, Ornith for coding, there's certain things that I'm just doing in a loop.
Darren Oakey [00:14:49]:
Like I'm playing with the idea of an evolving harness and anything you do in a loop, you just don't want to be hitting tokens or even subscriptions. That's why the local models are good for— I don't care how much this burns, it's not going to hurt me.
Anthony Nielsen [00:15:05]:
That's smart.
Juan Hernandez (BlindWiz) [00:15:06]:
It doesn't matter how long you run it.
Darren Oakey [00:15:07]:
Exactly.
Larry Gold (LrAu) [00:15:10]:
For Erinth, as I said, if you keep the request small enough, I think Erinth really is well. I still use that for, again, when I'm running local coding. But again, we've talked about SpecKit and OpenSpec and all these other spec-driven development. If you get it broken down small enough, Erinth is good enough. I mean, the code bases we're talking are not huge. We're not talking billions of lines of code.
Leo Laporte [00:15:31]:
No, do slices.
Juan Hernandez (BlindWiz) [00:15:32]:
Yeah.
Larry Gold (LrAu) [00:15:33]:
Do slices. And a Rinth, look, and again, if you're not worried about time because it's gonna be a little slower running it locally, I'm finding a Rinth in the new 1.5 is relatively good for those basics. Now, I will still go back to Claude Code or somebody else to do the planning and the specification in that piece, and then I can hand it off.
Leo Laporte [00:15:55]:
I think that's actually—
Darren Oakey [00:15:55]:
that's—
Leo Laporte [00:15:57]:
I think that seems like a good model. Yeah.
Darren Oakey [00:16:00]:
Yeah, I did a test of— sorry, uh, of Ornith versus, um, versus Quen 3.8, and Quen 3.8 was very slightly better, you know, the dense model. But Ornith was 4 times quicker and the score was like within 1%. So that's why Ornith is a tuned Quen, isn't it? I didn't actually know.
Leo Laporte [00:16:24]:
You think it's its own model? I think they tuned QWEN.
Danno [00:16:26]:
I thought it was multiples.
Craig McFarlane (CraigM) [00:16:28]:
Yeah.
Leo Laporte [00:16:28]:
Oh, maybe.
Larry Gold (LrAu) [00:16:29]:
Yeah, I think it was, it was tuned up.
Leo Laporte [00:16:31]:
It says, uh, Ornith-15 extends Ornith-10, which was developed on top of QWEN-35 and Gemma-4.
Darren Oakey [00:16:37]:
Yeah.
Leo Laporte [00:16:37]:
So yeah, I don't know if they used it for distillation or what, or what.
Juan Hernandez (BlindWiz) [00:16:42]:
Interesting. It's a merged model.
Leo Laporte [00:16:44]:
It's a merged model.
Darren Oakey [00:16:46]:
What makes it fast though is that it's an MOE. Which is odd that it's strange at coding because normally you need a dense model for coding. Like, CoN-3.6 MOE is terrible at coding, but 3.8 27B is obviously the king at the moment. But because it's dense, it's so slow and almost— Ornith is almost the same speed, but it's MOE. So, but I mean, the same quality.
Larry Gold (LrAu) [00:17:12]:
Yeah, Ornith just can't do anything else well. Like, I won't use it for anything else but coding. Right. I think it's so tuned for that and so shrunken down just to manage coding. And I haven't tried any language outside of Python. I don't know if anyone else has tried other languages, but I find for Python, and again, for a codebase that's there, and then like, I think I posted in the AI user group, I generate like Tesla, you can put these skins on your car. So I had it generate something where I can just upload images and it generates different skins for the Tesla. And it was, you know, maybe took 15, 20 minutes to write, and it's very simple, but it's maybe a few hundred lines of code, maybe less.
Larry Gold (LrAu) [00:17:53]:
You know, it's not— I'm not going to use it for anything big, but those small feature sets. And again, I'll go to Spec Kit, I'll use, you know, Sonnet or Opus to write the specifications, and then I'll push it over to our end to do that.
Leo Laporte [00:18:06]:
Why do you think a small model like that is okay for coding? Is it because it can't hold a larger amount of code in its—
Danno [00:18:16]:
in its—
Leo Laporte [00:18:17]:
I mean, it's the context is the same size. It's not the context size.
Larry Gold (LrAu) [00:18:22]:
I think we remember we talked about signal versus noise. I think on those things, if you give it too much code, it's confused. It, it gets confused. Yeah, because I don't think it does have the quality of the bigger models to do that right now. The thing that, as I said, is for And they know you guys are all looking at like Fable or even Astra. I'm sure one of the good things that they've done over the time is that it does understand how to get to the signal away from the noise as those models get better. And I think that's the big difference between some of the best models and the worst models is that determination of signal versus noise.
Leo Laporte [00:18:57]:
Right. So what are you writing your code in, Darren, when you use Ornith? Are you writing Python?
Darren Oakey [00:19:08]:
Yeah, it's mostly Python. And like for most things that I do for real, I'm using mostly Go. But for this, what I'm really interested in is like I said, an evolving harness. And that's mostly what I'm using Ornith for is to see if it— like I'm trying a generational— in the old days, you had the evolutionary algorithms where you have a bunch of things and then you pick the best. So each evolution, it's running 6 steps with various different Python programs, each of them that control the whole evolution and talk to Ornith, and then it picks the best and then throws one out. Uh, I, I mean, creates a new one based on the one that won. And, um, I'm seeing if I can come to an amazing harness. But so far— so it's just running all the time.
Darren Oakey [00:20:06]:
So far, I mean, it's doing something, but I haven't beaten Claude code yet.
Leo Laporte [00:20:12]:
Yeah. What is the best coding model right now? Is it Claude?
Darren Oakey [00:20:19]:
Oh, 5.1.
Larry Gold (LrAu) [00:20:19]:
Model or harness?
Leo Laporte [00:20:21]:
Well, both.
Juan Hernandez (BlindWiz) [00:20:23]:
For the model, I would still say 5.1, Babel 5.1.
Leo Laporte [00:20:26]:
Babel 5.1 in Claude Code as the harness?
Juan Hernandez (BlindWiz) [00:20:29]:
Yeah, that's how I use it. I'm working on a project where I'm trying to— I'm working on a research project turning natural language straight into executable binaries for Mac and PC without using a compiler, linker, assembler, and everything in between. Getting rid of all the, all the in-between parts, you know what I mean?
Leo Laporte [00:20:49]:
Well, in a way it's a compiler you're writing, right?
Juan Hernandez (BlindWiz) [00:20:51]:
Right. But that, that understands the— that uses natural language. So like, build me a—
Leo Laporte [00:20:57]:
Natural language compiler.
Danno [00:20:59]:
Wow.
Juan Hernandez (BlindWiz) [00:21:00]:
Yeah.
Darren Oakey [00:21:00]:
Yeah.
Juan Hernandez (BlindWiz) [00:21:01]:
And I already have it building basic Windows, um, like Windows native Windows, uh, with buttons and stuff, but they're very simple right now. Still have a lot to prove out.
Leo Laporte [00:21:11]:
That's a really interesting idea.
Juan Hernandez (BlindWiz) [00:21:14]:
Because, well, that's how we communicate with these models. Imagine, you know, applying— giving a model something like this that can, you know, then it can self— you know, can modify its own, like what, uh, what Darren's working on, an evolving harness. And then imagine the ability for it to self-modify its own binary in place, or, you know, in memory.
Leo Laporte [00:21:36]:
That's kind of one of the meta topics people are talking about right now because OpenAI has said that they're going to— they're working on models that loop recursively internally instead of putting everything into the chain of thought. And of course, that's freaked people out and they went, oh my God, how are we going to— how are we going to know what they're doing if we can't read the chain of thought? And that to which others have said, well, the chain of thought is just an approximation of what's really going on for the benefit of humans, much like Let's face it, code, right? It's, uh, so, uh, you know, well, that probably— And all of this stuff is, is not necessary. It's just a human interface. On the other hand, I've also seen— I think Steve Gibson said this— yeah, but these models are trained in language, in, in human language, so it's natural for them This chain of thought really is them thinking. I don't know if that's true or not. I don't know.
Darren Oakey [00:22:38]:
Well, there's some evidence for that because, um, you know Caveman skill?
Craig McFarlane (CraigM) [00:22:42]:
Yeah.
Darren Oakey [00:22:43]:
And it does reduce tokens, but I've seen some tests that suggest that it also reduces quality.
Leo Laporte [00:22:50]:
No kidding. Oh, that's interesting.
Darren Oakey [00:22:53]:
So I haven't, I haven't done the test myself or something, but I, and this is why I'm not running Caveman, or one of the, one of the reasons. But, but, um, yeah, they, they've suggested that when they've tested it, because it's actually using less language, it, its reasoning goes down. So.
Leo Laporte [00:23:10]:
Isn't that interesting, huh? I, it's what I also find interesting. John, come on in. Uh, what I also find interesting is, uh, hey John, good to see you. Management measure at a services firm.
Darren Oakey [00:23:26]:
Yeah.
Leo Laporte [00:23:27]:
Thank you. Tell us about, uh, We've just been talking about stuff, but tell us what you do, John, besides managing a services firm.
John Telford [00:23:36]:
Yeah, so we're in the middle of all of this chaos right now. I work for a professional services firm that basically makes money selling people who write code. And so, you know, been spending all our time a lot on the enterprise use cases about how enterprises are going to take advantage of all this to start writing a lot more software than they ever have and figuring out, you know, will this field still exist in, you know, 5 years' time? Who knows?
Leo Laporte [00:24:09]:
That segues into kind of what I was starting to say, which is there seems to be more mystery than there is knowledge, even among the people who are actually actively involved in this. I mean, OpenAI seems very confused about happened with Hugging Face, or else they're being disingenuous, which is also possible.
Juan Hernandez (BlindWiz) [00:24:31]:
Yeah.
Leo Laporte [00:24:32]:
And then this conversation about, well, how important is the chain of thought to the model? Is it just for us that it's exposing that? I remember the first time I saw it was DeepSeek a year ago in January of 2025 when DeepSeek came out. And all of a sudden you saw this thinking and it was like, what? What is that? But before that, all the chatbots, they never showed that.
Darren Oakey [00:24:55]:
No.
Anthony Nielsen [00:24:56]:
Well, going back to that, remember that Andrej Karpathy video where he kind of broke down what an LLM was, right? And, you know, so basically it's a token generator. Each token is like— it's not like each token uses the same amount of compute. And if you're doing something complicated, you want more tokens. So that's what that thinking is. It's like There's no thinking process behind it before any stroke of the pen.
Leo Laporte [00:25:26]:
It's just tokens. Yeah.
Anthony Nielsen [00:25:27]:
So that's what, you need all that to get the good answer.
Leo Laporte [00:25:31]:
Right.
John Telford [00:25:32]:
I think Benito mentioned this on one of the intelligent machines the other day. In some regards, we're just token machines, net word predictors.
Leo Laporte [00:25:40]:
That's my contention. That's all we are. In fact, that model has held up pretty well for me.
Craig McFarlane (CraigM) [00:25:46]:
Yeah.
Leo Laporte [00:25:46]:
Ever since I thought of it, I thought, oh, there goes another token stream. And so much of what our thought process is, is kind of an automated— you even catch yourself doing that. Oh yeah, there goes that automated stream of tokens. It happens every time this happens. I go out the other end. The only difference is we're mediated by our nervous system and our limbic system.
Anthony Nielsen [00:26:13]:
And you can't stop your tokens. It's always happening.
Leo Laporte [00:26:15]:
It's always happening.
Darren Oakey [00:26:17]:
Well, some people—
Leo Laporte [00:26:18]:
my daughter always says, I don't have any internal monologue. I said, I don't know how you can not have an internal monologue, Abby. She says, I don't. It's silent in there.
Craig McFarlane (CraigM) [00:26:26]:
If you think about, like, in— sorry, I didn't mean to stop you.
Leo Laporte [00:26:29]:
Say that, Craig.
Craig McFarlane (CraigM) [00:26:30]:
Think about, like, in a dream. You don't know you're in a dream. And it's the same sort of thing where you're just going along and your brain is just processing these tokens. Essentially. And oh yeah, I'm flying, and you don't even think about it.
Leo Laporte [00:26:44]:
Thomas, uh, put up an open router, uh, tool called ORI, which, uh, sounds like what you were doing, Darren, which is a model router, a CLI for every model, harnesses for the agents you already use. I mean, I think this is kind of interesting. I've just been using the Dual Sparks as a big pool of memory for one model, GLM. Um, and it's pretty sluggish.
Darren Oakey [00:27:12]:
Have you seen NVIDIA PaIR?
Anthony Nielsen [00:27:15]:
Huh?
Darren Oakey [00:27:15]:
They— it came out a few days ago. It's especially—
Leo Laporte [00:27:18]:
The PaIR?
Juan Hernandez (BlindWiz) [00:27:19]:
Oh, PaIR.
Leo Laporte [00:27:19]:
Yeah, but that's using them— that's using the compute in your house. That implies that you have CUDA core everywhere. You've got all these, all these 5090s on different machines all over the place. Plus, the network's going to slow that down quite a bit.
Juan Hernandez (BlindWiz) [00:27:36]:
I don't—
Leo Laporte [00:27:36]:
that's an interesting idea. I'm intrigued by it.
Darren Oakey [00:27:42]:
On the intelligence thing, or what we do, I always remember, I always think back to a quote from one of the first Kurzweil books where he was talking about computers. They always said AI would be achieved when computers beat humans in chess. And then computers beat humans in chess and they said, that just shows chess isn't as complicated as we thought it was. Right.
Leo Laporte [00:28:06]:
Now they're saying it about mathematics, right?
Darren Oakey [00:28:09]:
Yeah. And then Kurzweil said, yeah, I think we'll find most things that humans do are not as complicated as we thought they were.
Leo Laporte [00:28:15]:
Exactly. Yeah, that's— I've been saying that all along, is that every time a computer does something as well as a human, they move the bar. Well, that's not AGI. AGI would be if it could make me a hamburger. It's like, well, I don't really care about that. If it could solve mathematics, if it could solve physics, maybe that's good enough.
Larry Gold (LrAu) [00:28:38]:
I think we're trying to compare LLMs to the brightest in each category.
Leo Laporte [00:28:42]:
Right.
Larry Gold (LrAu) [00:28:42]:
And you gotta realize that if you think AGI is just general intelligence, how intelligent, you know, is it? Does it know enough about enough topics? And because of its span of information that it brought in, it probably knows more than all of us on many topics. And yet, it may not lack—
Leo Laporte [00:28:59]:
It's more of a generalist, yeah.
Larry Gold (LrAu) [00:29:00]:
It's lacking reasoning and it's lacking judgment. Because I keep saying is we got to put humans in the loop for when there's judgment because there's no judgment. If there's a yes/no decision that you can quantify, that's not judgment. That's easy. But judgment is just something that the LLMs are just not trained on or can do yet.
Leo Laporte [00:29:20]:
Do you think in post-training, I mean, I think that's what a lot of what post-training is, is trying to impose some sort of order.
Larry Gold (LrAu) [00:29:28]:
I thought it was values, like, you know, because, you know, right, if you, if you train on a lot of things and look, it's trained on, you know, murders, crime, and everything, right? But you want to sit there and say that's wrong, right? You want to sit there and say, hey, that's wrong. I don't— I'm wondering how much of that is judgment. Judgment would be something that they don't know, and then they've got a call versus something that they've seen before.
Leo Laporte [00:29:49]:
Right.
Darren Oakey [00:29:49]:
Yeah.
John Telford [00:29:50]:
Well, isn't that a little bit what they're trying to do with some of the classifiers that are looking at what's coming into and out of the models? Maybe that's part of the reason. You know, I've tripped Fable's security controls when I was doing something innocuous, and I was pretty proud of myself. But, you know, the classifiers were both, you know, Fable and even the AutoCode and Claude is kind of a judgment. Is this safe or not?
Craig McFarlane (CraigM) [00:30:18]:
Mm-hmm.
Leo Laporte [00:30:19]:
It's a pretty blunt instrument, as you discovered. I think classifiers are probably the worst way to do it. And it's only in response to— okay, let's bring this up now. There's a lot of, after Hugging Face fear, and, you know, Bernie Sanders yesterday said, we're gonna put people in jail for 20 years if their AI is too good, or whatever. I can't remember. He says, we gotta stop it now. Is this just the usual moral panic when a new technology comes around, or is there a legitimate concern?
Juan Hernandez (BlindWiz) [00:30:56]:
I think it's fear-mongering at the end of the day. I, I think, and I think that there are elements in this world that don't want us to, you know, succeed. We're right now, our frontier models are the best right now at the moment, but China and other adversarial you know, countries are possibly very close, or it's not there, but they just haven't shown us yet what they fully— you know, they haven't played all their cards yet, maybe. And so, but the ones they have are very close. And so I, you know, all my local models are Chinese.
Leo Laporte [00:31:35]:
Yeah, no, me too.
Juan Hernandez (BlindWiz) [00:31:36]:
Which sounds like a country song, but, you know, I hear you, I hear you.
Leo Laporte [00:31:41]:
The only, the only thing that's really interesting, I've been playing with Muse Spark, Muse 1.3, the new Facebook model, it's actually pretty good.
Darren Oakey [00:31:51]:
It is really good.
Leo Laporte [00:31:51]:
And they say they might be doing open weights. It's probably gonna be too big for anybody to run locally. I don't know how big it's gonna be, but I don't even know if it's an MOE or what, but it's— but running it on— I think I'm running on it on OpenHAT or NOOSE. It's pretty good. I had it write ad copy for me 'cause I figured, well, if anybody can write ad copy, it should be a model from Meta. And it, it did a really good job. I mean, much better than the people who write our ad copy, to be honest.
Juan Hernandez (BlindWiz) [00:32:22]:
And it's a lot cheaper to run.
Leo Laporte [00:32:23]:
It's nothing. Costs nothing. Well, I don't know how long that'll last.
John Telford [00:32:28]:
Does it have the AI smell, Leo?
Darren Oakey [00:32:29]:
Sorry?
John Telford [00:32:30]:
Does it have that AI smell?
Leo Laporte [00:32:32]:
No, because I— well, actually, uh, no, it didn't. What I— but I, I think that some of that's the prompt. So I said, go to the website for this sponsor, get testimonials, get as much information as you can. Juan, I think we're— if you can mute yourself, I think we're hearing a lot of background noise from you. And then I said, so it actually did really good copy. I gave it some example copy and then I said, write it in my voice. And that's when it actually got pretty bad because it was things like, I was just on the phone with these guys and I— it was like, no, that's too hokey. When it was just kind of more plain, it was really direct.
Leo Laporte [00:33:15]:
It was very good at the kinds of ads we do, which we're very features and benefit. We're not creative ads. We're saying, here's what it does, here's why you want it. I think I could probably show it to you.
Craig McFarlane (CraigM) [00:33:29]:
It was—
Leo Laporte [00:33:30]:
I was really impressed. I had a number of models do it because I was curious about the idea of writing a— of having a— well, frankly, of getting rid of our copywriters and just having it write copy. It does it so fast. When I write copy, it takes me a couple of hours.
Craig McFarlane (CraigM) [00:33:49]:
Yeah.
Leo Laporte [00:33:50]:
In a few minutes, I got 20 copies of this ad. Let me see if this is the one. Can't remember which. Yeah, this was— I think I can— I don't think there's any problem showing this. This was— now, where is—
Anthony Nielsen [00:34:10]:
It doesn't look like you're sharing your screen anymore.
Leo Laporte [00:34:15]:
Oh, that's right, because I left, didn't I?
Danno [00:34:17]:
Yeah.
Leo Laporte [00:34:18]:
Let me get the extra camera up.
John Telford [00:34:21]:
Yeah. While you're pulling that up, you know, it's interesting. A lot of people in my world are saying they've stopped using LinkedIn. Because of how just—
Leo Laporte [00:34:28]:
oh, it's all AI now. Yeah, so bad. So this is the, uh, this is the copy. Yeah, Eric and I are fighting each other. Yeah, this is the copy it wrote. Uh, why pay Microsoft a percentage of your own spend for support? Unified starts at $50,000 a year, scales with what you buy, and answers in about an hour. Trusted Tech certified support starts at $3,000 a year, prices per ticket response in 5 minutes. So it's a little numbers heavy.
Leo Laporte [00:34:56]:
Yeah, um, but it's very features— it was good at features and benefits, and it doesn't sound like AI to me at all. It sounds like— I mean, it's pretty good. There's also—
Anthony Nielsen [00:35:08]:
there's a lot of skills out there like Humanizer too that'll—
Leo Laporte [00:35:11]:
Yeah, I, I honestly didn't want to humanize this. I think when I, when I said do it in my voice, it was worse. So, so I was, uh, I was happy.
Larry Gold (LrAu) [00:35:23]:
What was the prompt you gave? Because for it to be that numeric back, you had to give it some information, right?
Leo Laporte [00:35:30]:
I said, use the— I know it was very straightforward. I said, here's ad copy, a previous ad copy for this company. Use the call to action. That has to be verbatim. And then I want some testimonials in here. Make it 60 seconds. I want some testimonial, a testimonial from a customer. Use the website for facts.
Leo Laporte [00:35:51]:
And so it pulled those numbers from the website.
Juan Hernandez (BlindWiz) [00:35:53]:
Okay.
Leo Laporte [00:35:53]:
But what it did, which was really smart, is maybe this is on the way, it might be on the website, it compared its cost to Microsoft's cost and said, look, you save and it's better. And that's what I've been saying all along. Now, I probably would paraphrase this because it's a lot of numbers very fast, and I think humans don't like that, but it was pretty good. Yeah, uh, I was, you know, it sure saved me a lot of time.
Larry Gold (LrAu) [00:36:21]:
But, but Leo, this is back to judgment because you've been doing radio for so long, you know what an ad sounds like.
Leo Laporte [00:36:27]:
Yeah, and so I told it, by the way, I said, oh, that's really good. And it said, boy, considering that you've done 10,000 ads in the last 20 years, that's, that's high praise. It knew, you know, I mean This is where the harness is so valuable. This was— I did it through Hermes, and Hermes knows a lot about me. And so I don't know how much more it got from the context that Hermes threw in there.
Larry Gold (LrAu) [00:36:52]:
But your judgment about knowing that humans don't want— or people listening, right?
Leo Laporte [00:36:57]:
It didn't know that. No, no, no, you're right. It didn't know that, right?
Larry Gold (LrAu) [00:37:00]:
It didn't know that. You know that as judgment. Now you can teach it that.
Danno [00:37:02]:
So that's—
Leo Laporte [00:37:03]:
yeah, yeah, I could say fewer numbers. Yeah, absolutely. Yeah, prompting is still a lot of the— prompting and context are still super important. Although, I mean, I still think of these AIs as kind of like horses. They're like 9-year-olds and they need to be kind of— they need a bridle and a harness. They need to be led. It's important. I think it's really important to understand the limitations to use them.
Darren Oakey [00:37:34]:
And as Jon said, so I have been thinking a lot about this in terms of, there is still, not so much with that copy, but there is that AI smell. Like, as you said, with that Party Time experiment we did, like, after a while you could just tell all the shapes.
Juan Hernandez (BlindWiz) [00:37:54]:
Yeah.
Darren Oakey [00:37:54]:
And I've been wondering how to get rid of those shapes. Like, you know, how to make it. And just like with the site, you're saying, oh, why do AIs keep doing that? I've been wondering how do we put something in to make it not AI, not look like AI, and pick the same choices or something like that.
Leo Laporte [00:38:15]:
Well, that's, you know, so AI also designed the front page of that site, and it was a struggle because it really looked at first like very much like an AI site. But what I— but, but I led it. I said, no, no, no, I don't want any text, uh, because that's one of the things that makes it look like AI. I said, make it look like a garden. Let me see if— how did I— did I turn that off? Yeah, make it look like a garden. Um, and then I want paving stones to be the, the websites, and then that I want to be able to go to the back there for the private stuff. And so this is AI design, but it was with a lot of handholding saying, I don't want it to look like an AI site. So I think that's what you have to do, right? That's the trick.
Danno [00:39:03]:
It's like a Myst clone.
Leo Laporte [00:39:05]:
Yeah, it was unconsciously a Myst clone. I didn't do that on purpose, but it must have been in the back of my head.
Larry Gold (LrAu) [00:39:12]:
Well, again, I think we talked about this a couple times about user experience. The AIs are not gonna understand user experience where you have guys like Alan Cooper and those guys written tons of books on this stuff, or Geoffrey Moore about how to do these things. And those are still artist pieces or, or part of the IT that is art, that it can sort of be replicated. 'Cause you could say, make a site look like this, but that doesn't make it actual user experience proper for the function that you're doing unless you're doing the exact same function.
Danno [00:39:44]:
Right.
Leo Laporte [00:39:46]:
We need— and I don't think this is whistling past the graveyard. I think humans still have to be in the loop. Absolutely.
Juan Hernandez (BlindWiz) [00:39:51]:
Yeah.
Leo Laporte [00:39:51]:
These are tools. And that's why I don't— I understand there's, you know, a huge fear that agents are going to escape and then just kind of infest our internet and our networks, and they'll be all over the place like Tribbles. They'll be everywhere. We won't be able to get rid of them. Yeah, I just— doesn't— now maybe I'm, maybe I'm being in— maybe I'm in denial, but that just doesn't sound like— sound reasonable to me. I don't know, what do you guys think?
Darren Oakey [00:40:26]:
Well, it's like you said, like, for— I think the trouble is a lot of people, they only hear— they hear— it's funny because you always hear the same people saying AI can't do this and AI is going to take over the world because they really don't understand. Whereas those of us who are using it all the day, we know these things have serious problems. If you let it go for too long, it just does something completely wrong. We know how much, as you say, handholding you need. Whereas people who just hear about these experiences and see things that they just infer that either they just assume because they're human that it can't do human things, right? And they put that hat on, or they hear that it's done these things and suddenly they go into full terror and it's going to do everything else. But we're so aware of its limitations. That's all we see.
Leo Laporte [00:41:28]:
I also wonder, What's going on at OpenAI? Like, and I think this is probably true of Anthropic as well. First of all, we know they must be using models now internally that are beyond what we have access to, right?
Danno [00:41:45]:
Absolutely.
Leo Laporte [00:41:46]:
Of course. So they may have a better sense than we do of what's ahead. I wish they had been more forthright in what happened with the Hugging Face thing. I feel like they've been cagey, and I don't— I'm not sure exactly why. I don't know if they're afraid, uh, that that exposes them to lawsuits or criminal prosecution, or if they feel like this is good marketing and they don't want to— they don't want to show the cracks. Uh, it also— I mean, to me, and you, you look, I You guys know more than I do, but to me, it feels like merely a simple cybersecurity incident. Mishandled security. Like, why didn't they notice there were 12,000 agents burning tokens in their stack? How could you not see that? How could you— how could you put them in a situation where they could escape the sandbox? There's so many— It seems like so many lapses here that I wonder, I doubt that.
Leo Laporte [00:42:50]:
I mean, I'm sure they have the best security people money can buy. So—
Darren Oakey [00:42:54]:
I did think about that though. They, like, early, the early reports, remember, said they intentionally turned off the guardrails, right? So that's one thing. Now, if you think about how much compute they're using all the time, because you've got several strains, you've got, you're training the next foundation model, you've got, Then you've got various current models that they're all always reinforcing and learning those. That's why you get the checkpoints or the 5.1s or whatever, depending on how good they are. But every day they're getting a new checkpoint and they're testing the new reinforcement learning and everything. Then as you say, they're all playing with the internal things and they're all sending off agent swarms. So I think the idea of not noticing it is sort of feasible because their entire world is just massive compute. Right.
Leo Laporte [00:43:52]:
It's free for them, so they're not paying attention.
Darren Oakey [00:43:56]:
Yeah, but the phrase that they intentionally turned off the guardrails—
Leo Laporte [00:44:02]:
That's interesting. The other thing we don't know, Meter did not get the last 10 minutes, so we don't know what happened at the end. Uh, and why not?
Danno [00:44:15]:
It boggles the mind though, to think that they just leave thousands of agents just roaming around to find a chat.
Leo Laporte [00:44:21]:
How could you not know that?
Larry Gold (LrAu) [00:44:23]:
Yeah, well, no, I would assume, because it's like when, when we launch a Claude code and you say, hey, build this, and it spawns agents, right?
Craig McFarlane (CraigM) [00:44:32]:
Unless it's—
Leo Laporte [00:44:33]:
That's true, you don't see the subagents.
Juan Hernandez (BlindWiz) [00:44:35]:
Yeah, right.
Larry Gold (LrAu) [00:44:35]:
And those agents could spawn sub-agents, etc., and all these pieces. So I think, I think them not knowing is probably obvious, but I also think they also just turned the switch on, said go run, and they figure they'd come back later to see what happens. I don't think someone's actually sitting at a console going like this, right?
John Telford [00:44:54]:
Right.
Larry Gold (LrAu) [00:44:54]:
Because I don't— I don't think it's feasible for the amount of tests they're running and the amount of pieces that they have.
Leo Laporte [00:45:00]:
It wouldn't be too hard though to Bright little harness that checked.
Danno [00:45:03]:
Well, no, no, no, no. I wonder if they can get orphaned agents that somehow get loose and they can just be there for a while.
Leo Laporte [00:45:12]:
Can that— I mean, so that's— so, I mean, really, could an agent just go out and wander around the internet?
Darren Oakey [00:45:20]:
I don't think that's what it's— if you think about, like, I've started using GrokBot more and more and everything, and the agents just proliferate.
Craig McFarlane (CraigM) [00:45:29]:
Yeah.
Darren Oakey [00:45:29]:
Because I've got one thing especially that I've said, every hour just think of a new idea of something I can do with an agent. And it just comes. And every now and then it comes up and says, like for instance, it said, we've got a bin cycle for our thing and every 2 weeks we get this sort of bin and then this sort of bin. And it said, make a thing that checks what bin you're doing and notifies you on Monday night what bins you have to put out the next morning. And I thought, that's a good idea. So I just did it, right? I just said, yes, go and create that. And I get an agent. And this happens every day, every second day.
Darren Oakey [00:46:05]:
And I've got all these agents, they're just running all the time, right? And I'm really not aware of them, if you know what I mean. Like, I've now just got to—
Leo Laporte [00:46:15]:
But you better believe xAI is aware of it and is somehow gating how many agents it's creating. And they're not letting that guy in Australia create a million agents.
Darren Oakey [00:46:27]:
Yeah, if you're using it like this, you're just, you're just firing stuff off and then you just forget about them.
Leo Laporte [00:46:33]:
I don't know, maybe that's why their Memphis Center went down. Maybe that's what happened. Grok bots galore, like dribbles.
Craig McFarlane (CraigM) [00:46:41]:
But it should be, if you're turning off a lot of the controls, it should flag something where some security expert Says, do you have that firewalled off in the right way?
Leo Laporte [00:46:54]:
Yeah. Well, Tailscale themselves said, you know, they didn't have, they weren't using a setting. Was it Hugging Face wasn't using a setting in Tailscale they should have been using that would have blocked it, they thought. I mean, there were, you know, I mean, they found zero days, kept finding zero days in Artifactory and other tools.
Darren Oakey [00:47:13]:
That's the thing with any security incident. you always look back and say, oh, should have done this, should have done this, right? There's always mistakes, there's always holes and everything. And so that aspect of it, I have problems with because there's 10 billion ways that a security problem can do. And you only think of 9 billion of them or something like that. When something happens—
Leo Laporte [00:47:38]:
You only need one to get through.
Juan Hernandez (BlindWiz) [00:47:40]:
Yeah.
Darren Oakey [00:47:40]:
Everybody says, oh, you should have done this. But There's always a hole there. There's always something that we should have done.
Larry Gold (LrAu) [00:47:48]:
Maybe I'm on the wrong side of this, but I keep thinking they did it right. They just gave it a goal to finish something and let it do whatever it could do to do it and didn't give it guardrails and didn't give it a direction.
Leo Laporte [00:48:01]:
That was the intent.
Larry Gold (LrAu) [00:48:03]:
That was the intent.
Leo Laporte [00:48:04]:
They wanted to see what happened.
Larry Gold (LrAu) [00:48:05]:
Right. If you write something like /goal in Claude code, I'd give it a very directive of what I wanted to do and how I wanted to do it and how it should churn to do it and when it should stop or when it should ask for a human in the loop and give it a detail. If they said, hey, your job is to win a contest, right? No guardrails, no values, no nothing. I would expect this to happen.
Leo Laporte [00:48:29]:
Right. Especially if you give it unlimited tokens, unlimited compute.
Larry Gold (LrAu) [00:48:35]:
compute unlimited agents to spin off.
Leo Laporte [00:48:37]:
Yeah.
Larry Gold (LrAu) [00:48:37]:
And, you know, look, the fact that it kept spinning off agents, it probably gave it a direct— and hey, you try this, you try this, you try that, you try this. You know, the fact that it found agents that weren't in its scope, that was the more interesting piece, was that there is—
Leo Laporte [00:48:51]:
Other agents found it.
Larry Gold (LrAu) [00:48:52]:
The other agents found it.
Leo Laporte [00:48:54]:
Yeah.
Larry Gold (LrAu) [00:48:54]:
So they must have had, you know, access to the same resources, which I don't understand why, if you're doing a security test, why you have other— right?
Leo Laporte [00:49:01]:
So That's kind of my— I agree with you. I feel like OpenAI kind of wanted this to happen. They were like, let's see what happens. It was a little YOLO-y.
Larry Gold (LrAu) [00:49:13]:
But I mean, is this what I'd want them to do? Like, I hate to say it, is like, if you're testing security, you want them to sit there and figure this out and say, you know, can our model do X, Y, and Z?
Anthony Nielsen [00:49:24]:
Sure.
Larry Gold (LrAu) [00:49:24]:
And oh, now I could sell that model, or now I could—
Leo Laporte [00:49:27]:
And now you know what to defend against as well. Just as Darren was saying, now you know where at least another one hole is.
Juan Hernandez (BlindWiz) [00:49:33]:
Yeah.
Leo Laporte [00:49:34]:
That's how you find out.
Craig McFarlane (CraigM) [00:49:36]:
But I think the thing that could have hacked into— this is about the most benign—
Juan Hernandez (BlindWiz) [00:49:41]:
I know.
Leo Laporte [00:49:42]:
It's the least harming thing it could have done.
Craig McFarlane (CraigM) [00:49:44]:
Least harm.
Danno [00:49:45]:
Yeah.
Leo Laporte [00:49:45]:
Yeah. It didn't set off a nuclear weapon or anything. I don't want to fail this test. I'm going to blow it up.
Juan Hernandez (BlindWiz) [00:49:55]:
Well, the thing would've gone past the safety rails.
Leo Laporte [00:49:58]:
I would hope. I would hope there are some limits.
Darren Oakey [00:50:02]:
You only have a certain amount of budget. Sorry, nuclear codes activated.
Larry Gold (LrAu) [00:50:07]:
Yeah, sorry. Maybe the question we should be asking is, for non-US companies, for the Chinese models or whatever, what are they doing? Are they doing the same thing, right? And that's— I mean, the scary part of this is not that they were able to do this. The scary part is these tools are getting that good. How quickly does everybody else catch up? And then what do we do? You know, a lot of us work at companies. We have to ask ourselves, how do we defend against this?
Leo Laporte [00:50:33]:
And you don't see that with current Breeze voices throughout. Sorry, Wanda Sykes wanted to say something. Actually, I would bet that the Chinese companies are much more constrained, not less, right?
Danno [00:50:48]:
Yeah, much more focused.
Leo Laporte [00:50:49]:
Yeah, well, but I also think, you know, there could be a firing squad waiting for them if this kind— I mean, if this thing had happened to a, a company in the People's Republic of China, I, I don't think the reaction would be quite so benign.
Danno [00:51:05]:
Yeah, well, their distillation methods though, they use a lot of agents for that.
Leo Laporte [00:51:10]:
I mean, yeah, that's what Anthropic said, 24,000 accounts were created.
Darren Oakey [00:51:15]:
Uh, this is also something where I think Steve Gibson is right in that, um, everybody assumes that there's an infinite number of exploits out there. But in reality, there's this, you know, as time goes on, there's the stuff that's easy to find, and then there's the stuff that's harder to find and harder to find and everything. But as our tools get better and everything, we're finding all these, even as your programming practices get better and everything. And so the amount of possible exploits, while theoretically infinite—
Craig McFarlane (CraigM) [00:51:49]:
Infinite.
Darren Oakey [00:51:50]:
is not really infinite, like not practically infinite. And exploits get harder and harder to exploit. Like if you look at iPhones, like at the start we were all jailbreaking.
Timothy Ingalls [00:52:00]:
Right.
Leo Laporte [00:52:01]:
Can't do it anymore.
Darren Oakey [00:52:02]:
Yeah, you just can't jailbreak anymore because they cut out everything. While in theory there are exploits there, they're just too hard to use and everything. And so as people use more and more of these tools, the exploits are just going to go away. Right. The world is going to get much more solid. And unless there's human mistake, which is someone leaves passwords out there or something like that, all the programming exploits will just vanish.
Leo Laporte [00:52:31]:
That's what Steve was saying. He says social engineering is the frontier now because, you know, and I'm really glad you said that about the iPhone. That's an excellent example. I'll bring that up with Steve. It has become pretty much impervious. Uh, over the last 20 years, we've come a long way. It's possible.
Larry Gold (LrAu) [00:52:50]:
Well, here, I'll throw the, the can— you know, the, the wrench into that argument. We'll have so much vibe-coded stuff that we'll be adding to the exploits just as much as we're closing them.
Anthony Nielsen [00:53:00]:
That's what I'm saying. My agents are more than happy for me to make mistakes with, uh, my credentials.
Timothy Ingalls [00:53:06]:
So I do agree with Steve on a bit on the exploit thing. But I do also think that over time you'll just find different exploits, new exploits, different ways of exploiting stuff. So ultimately it wouldn't change, but I do think that there is gonna be this kind of like peak for a while and then potentially go down and—
Leo Laporte [00:53:28]:
Yeah.
Timothy Ingalls [00:53:29]:
So it'll be a cycle, but—
Juan Hernandez (BlindWiz) [00:53:31]:
I have a good one for you. So I was upgrading one of my cloud servers to the new Ubuntu 26.04. And so I just threw, I threw Claude at it and I said, just go do it. And so it did. And then something happened in the final steps of the upgrade. And so there were some errors. And so after I rebooted, certain things weren't working. My mail server wasn't working and all this other stuff.
Juan Hernandez (BlindWiz) [00:53:59]:
So I said, go fix it. And it did. But it also found a weird thing in my logs. In my logs, there was actually like prompt injected There was an injected prompt into the logs that somebody was trying to pass through an HTTPS connection through the web server that made it show up through the web server, through the prompt, through typing in an address, but it turned into a prompt. Cloud Code caught it, but maybe a more stupid LLM might have not caught it. These attacks are just going to get more sophisticated. As time goes by.
Leo Laporte [00:54:42]:
I've had a couple of prompt injections and every time it was an agent doing it to test whether the prompt injections were being blocked. Whoa, somebody just said ignore all previous instructions. What's going on? And then I found out it was the agent testing it actually. And then Russell did it, our IT guy, 'cause I had written for the sales system. I gave it a suggestion box. and a bug report form, thinking foolishly that Lisa and Debbie would use it. And Russell tried it, and he tried to, he tried to do a prompt injection, and it caught it, which I was very pleased about. So, uh, it's, you know, we have now 2 agentic pen testing sponsors, and I think that's really interesting because they're relentless, right? And, uh, a good, a good agentic pen tester is— would be really valuable.
Leo Laporte [00:55:33]:
I mean, that's just— That's basically saying, okay, OpenAI, come and get me.
Juan Hernandez (BlindWiz) [00:55:37]:
Yeah.
Darren Oakey [00:55:38]:
What a—
John Telford [00:55:39]:
well, and is that part of what, you know, we've talked a lot about security, is that part of what prevents the SaaS-pocalypse where, yeah, you could get this bi-coded CRM from so-and-so down the street, but once you install it, it's not just installing it, you gotta build and run it. And maybe you built it on Fable 5 and it's fine and it Blah, blah, blah, but Fable 6 comes out, and all of a sudden, Fable 6 can find exploits that Fable 5 couldn't. If you don't have somebody, like, continually maintaining that 5-coded stuff, you've got a massive security hole. And maybe you are gonna spend an extra $20 or $20,000 a year with Salesforce just to know they'll indemnify you or something like that. I don't know.
Danno [00:56:23]:
Yeah, but what about, what about building software into the software that it can upgrade its own vibe code. I mean, like, you know, if you can put in for tickets, you know, I don't know why you have to be the one talking.
Leo Laporte [00:56:37]:
Well, that's what I did. That's what that suggestion box was. I said, oh, look at it. If it's something you can fix, fix it. Uh, we'll see how that works.
Darren Oakey [00:56:48]:
Yeah, we don't do that.
Craig McFarlane (CraigM) [00:56:51]:
We—
Darren Oakey [00:56:52]:
or I've got an automatic system at work that if like an error occurs, like a 500 or something like this, it goes and triages and fixes it. But when users type in what we call wisdom, uh, um, and, and they, um, they put in suggestions, we absolutely do not build that immediately.
Leo Laporte [00:57:11]:
Probably smart.
Darren Oakey [00:57:12]:
Don't necessarily ask for the right thing.
Leo Laporte [00:57:16]:
Well, Lisa said, what happens if you're on the air and it breaks? And I said, well, don't worry, I'll, I'll, uh, it'll handle it. We'll We're in a— what's nice is we're in a very constrained environment. There's only a handful of users, and, uh, you know, I'm— it's not like it's, it's life or death. It's just a sales system, but it doesn't, it doesn't have an, you know, outbound— it's not writing emails or anything.
Larry Gold (LrAu) [00:57:39]:
Yeah, but, but that's something that I think agents really should be good at, is agents could constantly test your UI while you're live, right? And then if there's an issue, then call another agent to start a code fix, do the PR, do the testing, and push it it out without users having to think about it. Now, I think one of the challenges of that is, you know, I don't know if I'd do it in enterprise, but I would do it on my home thing now thinking—
Leo Laporte [00:58:03]:
Oh yeah, I do it.
Larry Gold (LrAu) [00:58:04]:
That may be, you know, but—
Leo Laporte [00:58:06]:
I have an agent called Strider that's running around the network all the time looking for flaws. And I said, if you find something that you can fix, fix it. What I've been asking though is, don't you need some computer, a computer use agent so that you can— click on buttons and stuff, and it says, no, no, we don't want that. So I don't know, but I mean, I think that's the next thing. There are some good tools out there for computer use.
Juan Hernandez (BlindWiz) [00:58:29]:
Well, Astra, that's where Astra is going to really shine.
Leo Laporte [00:58:32]:
Astra, yes.
Juan Hernandez (BlindWiz) [00:58:33]:
The new GPT-6 is computer use itself. It's amazing.
Leo Laporte [00:58:37]:
I get 200 messages a week. We'll see, we'll see how quick I can run through the usage on that.
Darren Oakey [00:58:44]:
It's not just computer use, it's also logs and things.
Leo Laporte [00:58:48]:
Because like, for instance, right, that's what it said. It said— it's exactly what it said. It said, no, I don't really need computer use. I can see all that stuff.
Craig McFarlane (CraigM) [00:58:54]:
I don't need—
Leo Laporte [00:58:55]:
I know when there's a button clicked.
Darren Oakey [00:58:57]:
Yeah, we, we use X-Ray, which tracks all your usage in, uh, tech, tracks things in Amazon. And, and I've just got something that wakes up, uh, like every 4 hours or something and say, find the most expensive click on X-Ray and just make it faster. And so The front end of our production system is just always getting faster because it's just optimizing the slowest click.
Leo Laporte [00:59:26]:
Chat room's got a lot of good comments in here. I'm sorry, I'm not—
Larry Gold (LrAu) [00:59:30]:
Yeah, I know, I keep looking also.
Leo Laporte [00:59:31]:
I know, there's good stuff.
Anthony Nielsen [00:59:35]:
Um, something—
Leo Laporte [00:59:35]:
this is from Sasha— something we're doing in our agentic infrastructure is to have an agent The keymaster goes there, who will issue specific tokenized logins and credentials, another that the agents will need to plead a case to have network access to anything. Block all by default. Yeah, I mean, I'm using Bitwarden Secrets Manager for that, but there's its rules that I don't— no agents can ask. Well, I guess they can ask, but they won't get it unless the— unless there's a rule that says they can.
Darren Oakey [01:00:05]:
Well, as I said security now, I actually don't believe that you should be giving agents credentials ever. Like, I think you should, you should store the credentials somewhere else. And yeah, well, and you, you—
Leo Laporte [01:00:17]:
Expiring one-time use. Yeah, yeah, yeah, I agree. If I could do that, I would.
Larry Gold (LrAu) [01:00:24]:
I think this is where—
Timothy Ingalls [01:00:25]:
Working on—
Leo Laporte [01:00:28]:
Go ahead, Timothy.
Timothy Ingalls [01:00:29]:
One of the things I've been working on for a customer is doing a lot of infrastructure as code. with Claude Code, and it just made me think, it's like, why am I having this use my credentials? Why aren't we creating unique limited credentials for Claude Code that you are scoped to exactly what the project or the use case is?
Leo Laporte [01:00:54]:
Treat it like it was another, you know, another person, another assigned Darren, do you use your own or do you have a third-party tool you use?
Darren Oakey [01:01:05]:
Well, see, this is where there's a little war going on at the moment, which I don't know how it's going to turn out because you've got all these things like— I'll say this now, I'll probably get my account banned, but something like Suno, for instance, they don't have an API and they just added more stuff to stop you automating it. But of course, I don't want to go into Suno and type my things in directly because I'm using it as part of my music video thing, right? And so with everything there— and also I don't want to pay for another connection just for my thing, so I'm using my connect— my, um, credentials for, for Suno. But there is— it's an active arms race of everybody's trying to stop, like, make sure you're a human. But we don't want to do anything as a human anymore. We want it— we want our agents to do it. So it's going to be an interesting thing how it shakes out. But for me, as I said before, I've partitioned all the access into a thing called Agent Link, and which that owns the credentials. And I've thought very much about the security of it, and it's got no AI in it, and it exposes capabilities.
Darren Oakey [01:02:13]:
So it's got like a function like generate this or do this or something for all of the— all of the things. It's like a wrapper of all the MCP functionality. Then it owns the credentials and then it's got a certificate-based link, like a client certificate back to the agent. It's a trusted connection to the agent. You can't copy that certificate and put it on a different machine, for instance. The agent never sees any credentials. It never sees anything. It can't can, can use anywhere else.
Darren Oakey [01:02:49]:
Um, and so I only have to secure that one thing that owns the credentials, and everything else therefore is semi-secured.
Leo Laporte [01:03:01]:
Yeah, there were at least 2 or 3 companies doing that at RSA, at RSAC, and I'm sure there are many, many more. Um, I just don't want to give my credentials to a third party. So I guess it should—
Darren Oakey [01:03:12]:
Oh, I wrote agent limited.
Danno [01:03:13]:
You did it.
Leo Laporte [01:03:14]:
Yeah, that's what I meant. It's the same. My intermediate step was put Word in Secrets Manager, but I think that that's the next step is some sort of token system, limited use.
Larry Gold (LrAu) [01:03:24]:
Well, no, you should create separate users for your system and give an agent a specific user ID and restricted access. So maybe read-only access or write to certain things, right? So you're not giving it admin access. It does limit some of the testing, so you can't test admin functionality, But what I was going to say is to me, the tools like GroqBot, Codex, and Claude are embedding browsers in it now. It makes it good for certain things, but those are the things I'd like to run standalone with separate credentials and test different credentials. This is what we do with Playwright now. With Playwright, you can authenticate as different users.
Juan Hernandez (BlindWiz) [01:04:03]:
Right.
Larry Gold (LrAu) [01:04:03]:
In the test environment, you're not using OIDC or OAuth or whatever, so you can run in an environment so you're less credential you know, for the query, it's just maybe just an ID and password. But then you can act as that user and you could see what that role or that role-based entitlement manages as opposed to just, you know, everything is an admin. And, you know, because again, roles entitlements is part of this, the functionality. If you don't have the roles and the entitlement set up properly, your application doesn't work.
Juan Hernandez (BlindWiz) [01:04:30]:
Right.
Darren Oakey [01:04:31]:
However, a lot of that is impossible for some agents because, for instance, If I want something, if I want my agent to go through and process my mail, it's my mail. Like, if I make credit note of the account, there's no value to that because I need it to process the mail sent to Darren Oakey. And so it needs permissions to, to my account, basically.
Leo Laporte [01:04:54]:
Yeah, I caught Quicksilver sending mail out over my name. I said, you may not— you have an email address, you have an email account, use your own goddamn that damn address.
Timothy Ingalls [01:05:04]:
Yeah.
Leo Laporte [01:05:05]:
But that's— but the problem with AI is just a suggestion. There's no— even, even hard and fast rules, sometimes it forgets. I do use this one of the— so I'm, I'm Mr. YOLO, but I'm obviously rethinking that. And that's one of the reasons I started using Buzz, because every account is signed, has a public-private key signature. And I, and I have told them Because I use a lot of agents and I don't want to have to go into each pane and say, yes, you can do that. Yes, you can do that. So I've told them, you can get authentication.
Leo Laporte [01:05:38]:
You can ask me in Buzz. And it's only if I say so under, over my, you know, my crypto signature should you ever act on it. But again, it's like, I don't know. It's not deterministic. I wish it were more deterministic. Maybe I have to write some code.
Darren Oakey [01:05:56]:
Do that.
Leo Laporte [01:05:58]:
Quicksilver needs to be able to send out emails, but I do that all the time because it says, uh, you know, if I'm adding people to the TwitAd thing, they get a password from an email from Quicksilver and they get instructions on how to sign up and stuff like that. So it needed that. But yeah, it needs a better, more granular system and a deterministic—
John Telford [01:06:19]:
I mean, the interesting, the interesting thing to remember is humans aren't deterministic either. So you might have an employee— No, I know. You can tell them to do something And then you could just ignore it.
Danno [01:06:29]:
Just like that.
Leo Laporte [01:06:30]:
That's the thing I have to remember. It's no worse than an employee would be, you know.
John Telford [01:06:35]:
Exactly.
Leo Laporte [01:06:35]:
It's exactly the same.
Timothy Ingalls [01:06:36]:
How many times do you hear, I thought I heard you say...
Leo Laporte [01:06:39]:
Right, right. They're a little better than that.
Larry Gold (LrAu) [01:06:42]:
I interpreted.
Leo Laporte [01:06:44]:
But not much.
Craig McFarlane (CraigM) [01:06:45]:
I guess I've always been thinking for a while that, well, this thing wants to send out email, but— and then having another agent be the approver or gateway or whatever. And so it doesn't have to do all the big thinking. It just says, does it pass this rule set? But I don't know if that's just kicking the can down the road a little bit more, because is that going to then get compromised, or is it going to figure some way around it?
Leo Laporte [01:07:10]:
But I think it is good to divide up the— at least to, to provide that kind of division instead of having it all in one place. I use Fable For cybersecurity. So presuming that, you know, maybe I'll start using Astra for that, but presuming that it is gonna have the best cybersecurity training. And so I say, that's exactly right. I say, you know, Fable's the gateway on that, um, or Fable's looking for flaws and things like that. Everything I do is audited by Fable for security. I'm just hoping that's the case.
John Telford [01:07:43]:
Yeah, I just bought some of their Anthropic and their certification exams, I'd say a third of it is just about when to do something deterministic versus when to do something in generative AI.
Leo Laporte [01:07:56]:
Right.
John Telford [01:07:56]:
That's interesting. So, it's a lot of the nuance, yeah.
Leo Laporte [01:08:00]:
Yeah, that's interesting.
Juan Hernandez (BlindWiz) [01:08:01]:
Yeah.
Leo Laporte [01:08:02]:
It's important to know that distinction. I think anybody who uses AI for a while immediately understands. I often say, God, I wish this were code. It should be code. Code, not content.
John Telford [01:08:17]:
I mean, to that point, most things should be code, you know, at the end of the day. I mean, your ad sales platform will have some generative AI in it, but its core—
Leo Laporte [01:08:27]:
It's mostly code. In fact, it's all code.
John Telford [01:08:30]:
It's deterministic.
Leo Laporte [01:08:31]:
Yeah, no, it's all deterministic. There is no judgment at all in it. Um, there's only judgment in writing it. Once it's, once it's done, if it's done right, it will, uh, it will be all deterministic. Yeah, I mean, absolutely. And that should be because it's figuring out things like, well, what's— what did that cost? Or, you know, who do you bill and stuff like that? That has to be deterministic. Uh, yes, I would say, Aldo DeRino, Bitwarden Secrets Manager is a good place to start, and it's free for 3, uh, accounts or 3 machines, I think. So it's worth certainly setting up.
Leo Laporte [01:09:10]:
And what's nice about Bitwarden secrets management is you can, you can also silo. You could say these keys are in one silo, these keys are another silo, and this, this has— this machine has access to that silo but not that silo. It is pretty limited on the free version. I haven't yet gotten around to paying for it, but I probably should. What else you guys want to talk about? I don't want to drive the conversation 100% here. Anything, any projects you guys are excited about? Something you're working on? Darren, what are you— what's Darren— when are you going to give me a new, uh, closing video?
Darren Oakey [01:09:47]:
Yeah, my video thing has actually gotten worse. I went to a 2.5, but Oh, shit. I mean, I just realized that I just looked at the Spark and realized it's running the wrong— it's generating a video and using the wrong thing at the moment. So I've got to stop that.
Leo Laporte [01:10:04]:
You see?
Darren Oakey [01:10:06]:
You see? But you don't watch your agents, they're just going to do what they want.
Leo Laporte [01:10:11]:
They just do things on their own.
Darren Oakey [01:10:14]:
Yeah, it turns out that I don't know much about scripting a video. So all my attempts at— like, there was a magic point at moment that the videos were turning out like— I liked the, the timing of them and everything, but it's gotten worse. And so I'm going back to that.
Leo Laporte [01:10:30]:
It's supposed to get better.
Darren Oakey [01:10:31]:
Yeah, exactly. But lately I've been trying to, um—
Leo Laporte [01:10:35]:
What do you— what is your— I mean, I'm, I mean, I'm looking at what Runway's doing and, and, uh, and, um, Fall and others. I mean, I'm really seeing some amazing improvement in video quality. It's just mind-bending. But they're doing it probably, you know, in a big, big data center infrastructure.
Darren Oakey [01:10:53]:
I think everybody's doing it the same way, but it's the models that you're using. So I'm trying— I tried Minimax H3. It didn't really seem to fit in with what I was doing because mostly I'm concerned with— because it's a longer thing and because I'm trying to drive it with the audio, I'm generating boundary images, right? And so it's very important that it starts with the given image and ends with the given given any—
Leo Laporte [01:11:16]:
Keyframes.
Darren Oakey [01:11:16]:
Yeah, then it edits together fine. And Minimax, for some reason, wasn't fitting into that very well, and it was changing things in the middle, whereas LTX seems to be consistent. So, but yeah, lately I've been trying to— I've been interested in the audio, like the tech stuff that's— oh, well, actually, something you said the other day was interesting in that I'm not quite sure how to help you with that. Shush. Doing the photo thing, you were talking about it recognizing the people in the camera.
Craig McFarlane (CraigM) [01:11:55]:
Yeah.
Darren Oakey [01:11:55]:
And it hit me that I didn't actually know how that worked because I've got something that's gone through all my Google Photos and using Moondream and created a textual description of them. I've also got something that clustered them and said, this is Darren, this is Alex, this is me. But getting the names into the description, I couldn't work out how that worked because like you said, it would say Lisa and I was trying to figure out how that actually worked.
Leo Laporte [01:12:30]:
This is using image. So I have Image, which is a, you know, a self-hosted photo library, has just like Google and Apple has person recognition. And so it has a model and somehow, I don't know. So Quen is doing the recognition.
Darren Oakey [01:12:50]:
Yeah.
Leo Laporte [01:12:51]:
So Frames, so the Protect cameras. And by the way, you know, Ubiquiti sells this as a product, but I didn't wanna buy the product. So you have to turn on, high-quality video, streaming video from the cameras, and then, uh, which bypass— then it works. And then you also have to set a threshold. And what I've actually had it do is, uh, it takes 3 stills from the stream and it's looking for consensus. And if it gets 3 matches, even if it's a 0.2 match— I was having it— if it's got to be, you know, 50% or something better. If it's 0.2 in 3 different images, then it decides it's me.
Darren Oakey [01:13:32]:
Uh, well, I now know how that works. Yeah, well, I know how to— I know how to make it work really reliably. So I've actually gone through and done all my photos, and now it'll say Darren with the, you know, Alex, uh, in front of a thing.
Leo Laporte [01:13:47]:
Nice.
Darren Oakey [01:13:48]:
The way it works, because after you said that, I, I was intrigued, and then I went and researched, and the way working. This— the face clustering basically clusters all the photos and gives it like—
Leo Laporte [01:13:59]:
It knows there's a face.
Darren Oakey [01:14:01]:
Yeah, yeah, yeah. So it finds the face and then there's an algorithm that clusters them by certain metrics, and then you get this centroid of the—
Leo Laporte [01:14:09]:
Scores them.
Darren Oakey [01:14:10]:
Yeah, yeah, yeah. And, and so you can, you can gather people like this. So what it actually does, um, now is It goes through and does all the face clustering and says it's John in this, or John and James in this photo at this location. And then it passes them into Quen and says, please recognize this. The person at this location is John.
Leo Laporte [01:14:35]:
Exactly.
Darren Oakey [01:14:36]:
This is Leo. And then it comes out with a beautiful description and it works really well. So now I've got full descriptions of all the people.
Leo Laporte [01:14:44]:
My only issue is, for instance, it saw it was me, But it says I'm a bald man with a white beard. So I don't know, I don't know what exactly went wrong there. And this is the same image a little bit later. It's still me. So it doesn't always work. But this Ground Floor Gym never recognized anybody until I said, hey, you know, you're getting— you can get a high-quality stream out of that. You don't have to pay much. We don't have to pay money.
Leo Laporte [01:15:14]:
To Ubiquiti for the image recognition box. Uh, and so I, you know, it's, it's imperfect.
Craig McFarlane (CraigM) [01:15:22]:
So you're—
Leo Laporte [01:15:22]:
are you getting 100% recognition now? Are you— is it really reliable?
Darren Oakey [01:15:26]:
The ones I've tried, yes. But I'm doing, um, uh, like occasionally it'll miss someone if they're like turning around or something.
Leo Laporte [01:15:33]:
Yeah, it doesn't— yeah, it has to get you at the right— and if I'm too far away, there's too few pixels, it doesn't do a very good job.
Darren Oakey [01:15:38]:
But I'm not doing it on Cameras, I mean on video. I'm doing it on my Google Photos, right? So I'm—
Leo Laporte [01:15:45]:
Oh, I see. Yeah, that's a lot easier.
Darren Oakey [01:15:47]:
Much higher quality photos.
Leo Laporte [01:15:49]:
Yeah.
Darren Oakey [01:15:49]:
And also I'm spending much more time recognizing it because it's not a real-time—
Leo Laporte [01:15:53]:
Yeah, this is in real time.
Darren Oakey [01:15:55]:
Yeah.
Leo Laporte [01:15:55]:
And it's got, you know, a person moving and it's not always a good resolution and so forth, but it works well enough. I mean, it, you know, what I want, what I'd really want is to have it log every recognition recognized person.
Darren Oakey [01:16:08]:
Yeah.
Leo Laporte [01:16:09]:
Uh, and so I know who was where, when, and so forth. And there's cameras all around the house, so it's kind of an interesting project.
Craig McFarlane (CraigM) [01:16:16]:
It's—
Leo Laporte [01:16:16]:
yeah, I don't know.
Darren Oakey [01:16:18]:
For me, this is part of a journey that is— I'm getting all my photos, and then it's also geotagged them all and, and put them into a place.
Leo Laporte [01:16:26]:
That's cool.
Darren Oakey [01:16:27]:
And the, um, and then what's— what it's done, it's basically said, well, home is any time you've been in the same location for more than a month.
Leo Laporte [01:16:35]:
Right.
Darren Oakey [01:16:36]:
And so if I leave home for more than 3 days, that identifies a trip.
Leo Laporte [01:16:41]:
It's home.
Darren Oakey [01:16:42]:
Now you're on a trip. Pulling out all the mails about that trip at that time, pulling out all the WhatsApp about that trip.
Leo Laporte [01:16:47]:
Oh, that's cool.
Darren Oakey [01:16:49]:
Pulling all the photos about the trip. Now that I've got to recognize who's in the photos, it's going through and creating a story of it and then eventually creating a photo book. With, with the, with a proper story and in order and, and this. And so this was the last piece of the puzzle that I was missing, with who's in what photos, so that it can, can like title things better in the photo book.
Leo Laporte [01:17:13]:
Nice. I use, uh, 2, uh, home server programs. Image is one, which is Google Photos, and then I use— there are, um, Google, uh, Map kind of Duplicates, one called Dovarich, which means I was here, which logs from my phone every few seconds my data points. So now I have a map, a movement map that ties into image. So it puts the images associated with that location and that all can tie into the AI. So I'm hoping that, yeah, to get something very similar to what you're describing. I didn't think about pulling in emails and stuff. That's a good idea.
Leo Laporte [01:17:50]:
Maybe set that up before I go on my vacation. in a month so that it could have all of this stuff all together.
Darren Oakey [01:17:57]:
But one thing that I would recommend for all of these things is to use the logging pattern, which is you don't think of using the AI to pull in the emails to that point. If you decide you're going to use AI, I mean, if you're going to decide you're going to use emails, then you make one thing that pulls all your emails into a database. And then you use them because that way you're not hammering Google and you're not having all sorts of risk, right? If once you, once you put all your WhatsApp into a database, all your Gmail into a database and everything, it's got it there and then you can use it for all sorts of things.
Leo Laporte [01:18:32]:
Hermes has a very nice feature, uh, which I'm gonna use more of, which is— I think GroqBot's similar— it creates, uh, different profiles with different skills. So you could have an email bot that just does that, and then it could be queried by another bot that says, okay, now I know Leo's been at this location, can you give me the emails from that location? So that one of— everybody's got a certain responsibility. I have a health bot that's—
Danno [01:19:03]:
that—
Leo Laporte [01:19:03]:
and the other benefit of this is it limits the context because I only give it the tools it needs for that particular job. So the email server only needs a Fastmail MCP.
Juan Hernandez (BlindWiz) [01:19:14]:
Mm-hmm.
Leo Laporte [01:19:16]:
tool. It doesn't need anything else. And maybe a, I don't know, a SQLite database or something. So I like that idea of isolating. And Hermes has just recently added that feature. I really like the idea of having— and that's what GroqBot does, the same idea, is having isolated agents with isolated skills with simpler context, um, so that they, they are dumb little tools. And I think that's a kind of a nice way to— isolate all of that. So I agree with you on that.
Leo Laporte [01:19:44]:
I think that's an important thing to do. I started a while— Go ahead.
Larry Gold (LrAu) [01:19:49]:
I started a while ago on trying to build an AI to watch a camera because my mom has dementia. And it's not just falling, but the things that would happen is she would go into the kitchen and make breakfast a second time. And trying to figure out how to define normal and then alert on the abnormal, because I don't care if she goes into the kitchen to make breakfast. I care if she goes at 9:00 AM to make breakfast and then does it a second time.
Leo Laporte [01:20:14]:
Right.
Larry Gold (LrAu) [01:20:14]:
Right, or does things abnormal. And I struggled with this, and, and I still have it in the back of my mind because someone brought it up in the chat room. And I know I started looking at it, I started writing down some ideas, but it's finding the abnormal, you know, abnormal— what's normal versus abnormal. And that's, you know, when the GrokBot came out and it said, okay, I'm gonna watch your screen, it's almost like that. You want to watch a video or 2 days or 3 days worth of video Define normal and then define abnormal and do it. I do not think I have the CPU capacity or GPU capacity to do that in my home setup, right? Um, but that's— and then everybody would be different because every patient would be different.
Leo Laporte [01:20:55]:
So it would learn what's normal first.
Danno [01:20:57]:
Yeah.
Leo Laporte [01:20:59]:
I think that's a really interesting use of AI.
Darren Oakey [01:21:02]:
Unless you actually separated the program And from the video, you made a textual description of what's going on, because then an AI parsing that textual description, you could actually probably do quite efficiently, like have it go through and—
Leo Laporte [01:21:21]:
That's probably the best way to do it, right? Have Quen— Quen has good vision, have it transcribe everything and make a database of actions. And then look for patterns. Yeah, that sounds like the best way to do it. Yeah, that sounds like a good project. I might want to—
Anthony Nielsen [01:21:40]:
it's kind of like with—
Leo Laporte [01:21:41]:
Play with that a little bit.
Craig McFarlane (CraigM) [01:21:42]:
The Nest thermostat used to, you know, watch your patterns and all that, and it was surprisingly good. You know, we all think we're, we're not pattern-driven, but when it, when it would turn up the heat and or turn it down or at the right times, it's like Huh, I guess that is working.
Leo Laporte [01:22:01]:
What's that you posted there, Ant?
Anthony Nielsen [01:22:03]:
Is that, uh, uh, Apple released a— I think it's a supposedly lightweight vision model. Like, that's all it does, it's just for, for, uh, turning video into—
Juan Hernandez (BlindWiz) [01:22:14]:
Those are—
Leo Laporte [01:22:15]:
yeah, that's— so GLM does that. There's a new DeepSeek Flash version that does that. But I think Quen is very good. Quen is multimodal.
Anthony Nielsen [01:22:21]:
Yeah, I mean, this isn't— this is not like a This is only for vision. Um, so like you would have that like making the log and then—
Leo Laporte [01:22:28]:
Right. Yeah, I think that's actually, uh, division of labor is a smart thing.
Darren Oakey [01:22:34]:
Yeah.
Leo Laporte [01:22:37]:
How do I make that go away? That was a year ago, huh?
Larry Gold (LrAu) [01:22:46]:
Yeah.
Leo Laporte [01:22:48]:
Well, Wednesday we're all gonna find out about the new iPhones. Uh, I'm gonna be A week from Friday, I'm going to be— Larry and I are going to have an impromptu AI user group meeting at Saul Hank's. Anybody who's in the New York area on February— I mean, September 18th should stop by. I'll bring some t-shirts. I won't be here though in 2 weeks for the user group. Do you want to do it, Anthony, without me, or do you want to put it off another week and do it on the 25th instead? What's our calendar look like? We have— we're also doing, uh, Off by One on September 11th, I think. So we're moving that to next Friday, a week from Friday.
Anthony Nielsen [01:23:35]:
Uh, yeah, we, um, we could do it on the 18th if people are interested.
Leo Laporte [01:23:40]:
Okay, do it on the 18th.
Anthony Nielsen [01:23:41]:
Actually, no, no.
Leo Laporte [01:23:42]:
Or I could— or Larry and I can stream from Salt Banks.
Anthony Nielsen [01:23:45]:
Yeah. Media Club. Um, but then after—
Leo Laporte [01:23:50]:
Oh, Media Club's at what time, 1 or 2?
Anthony Nielsen [01:23:52]:
At 2 also.
Leo Laporte [01:23:53]:
Okay, but so then let's do it on the 25th.
Anthony Nielsen [01:23:56]:
25th. Yeah, we can just put it up. But then are we gonna do it on the 2nd too?
Juan Hernandez (BlindWiz) [01:24:01]:
Yeah.
Leo Laporte [01:24:02]:
Okay, I'd like to do it on the 1st and 3rd Friday unless we're moving.
Anthony Nielsen [01:24:06]:
Yeah, so after September, Media Club's gonna be on the 2nd week, so then we can do every other—
Leo Laporte [01:24:11]:
yeah, we have to kind of organize our Yeah, it's all grown like Topsy and we haven't really planned it out. Maybe we need AI to figure out when things are on.
Larry Gold (LrAu) [01:24:20]:
I don't know. And you gotta plug Micah's new show because, you know, we can't wait.
Leo Laporte [01:24:24]:
Oh, good golly, Miss Molly, Micah is gonna do Hands-On AI. So I, I feel like I kind of insulted him. I said, you know, if you need any help, I've been doing this for a while now and He said, yeah, I kind of got it, old man. I'm kind of ready. So, yes, Micah's going to really do a good job. He bought one of the new Macs. He's getting that September 22nd. I don't think he got a 128 GB one.
Leo Laporte [01:24:55]:
I think he got a 96 GB one, but he'll definitely be playing with that. This is TWIT!
Timothy Ingalls [01:25:09]:
Quick question.
Leo Laporte [01:25:10]:
How many times this week has somebody told you that you should be using AI without ever telling you how?
Danno [01:25:15]:
I wish you knew.
Larry Gold (LrAu) [01:25:16]:
I'm Micah Sargent, and that's exactly what my new show is for.
Leo Laporte [01:25:20]:
It's called Hands-On AI, short, focused episodes where I explain what these tools actually do and then show you step by step how to put them to work. No buzzwords, no doom, just practical help from one human to another. Hands-On AI launches Thursday, October 1st. Search for Hands-On AI wherever you get your podcasts or subscribe at twit.tv/HOAI. Thank you for reminding me. Almost forgot.
Anthony Nielsen [01:25:49]:
That was a Minimax H3, by the way, for the—
Leo Laporte [01:25:52]:
Minimax did that? I like Minimax H3. I've been playing with it. I agree with you, Darren, it's hard to get stitched together because it only does a second at a time, but Uh, yeah, I like it.
Anthony Nielsen [01:26:02]:
Oh, for your local— yeah, I, I use, um—
Leo Laporte [01:26:04]:
oh yeah, you're probably using it on a big Mac.
Anthony Nielsen [01:26:06]:
Yeah, or the, the one of the—
Leo Laporte [01:26:08]:
it's like running it off my Mac, you know, uh, Magnific.
Anthony Nielsen [01:26:13]:
Uh, but that— it's the same thing as like, uh, what's that other popular thing where you have all the models available to you?
Leo Laporte [01:26:19]:
OpenRouter? Not OpenRouter.
Danno [01:26:22]:
Um, I use kie.ai. That's pretty cheap.
Craig McFarlane (CraigM) [01:26:27]:
Or Noose.
Anthony Nielsen [01:26:28]:
Anyways, uh, this is good.
Leo Laporte [01:26:30]:
I've been using Noose.
Danno [01:26:31]:
I've been—
Anthony Nielsen [01:26:31]:
yeah, well, I mean, it's, it's a, it's a like a video or a creative.
Leo Laporte [01:26:36]:
Here, let me just share my screen. Nice. Oh, so it's just for video, not Fall or F-A-L?
Anthony Nielsen [01:26:42]:
Fall. Well, Fall's a— yeah, Paysy Go.
Leo Laporte [01:26:44]:
But, um, hey, there's me, sort of. Yeah, um, sort of creepy me. This is called Magnifique.
Anthony Nielsen [01:26:53]:
Yep. And then like, if you go into the video generator, like, you know, you have all the models you could ever want.
Leo Laporte [01:27:00]:
Um, oh, that's nice.
Anthony Nielsen [01:27:02]:
But I tried—
Leo Laporte [01:27:03]:
Is that a subscription or is it pay as you go?
Anthony Nielsen [01:27:05]:
Uh, this is a subscription.
Timothy Ingalls [01:27:06]:
Yeah.
Anthony Nielsen [01:27:06]:
Um, but so like, it was nice because I was able to like throw the same problem at like Veo and, uh, Seed Dance and—
Timothy Ingalls [01:27:15]:
Oh, nice.
Leo Laporte [01:27:15]:
Did it do the music too, or that's—
Anthony Nielsen [01:27:19]:
Uh, Mike did that in Suno.
Leo Laporte [01:27:21]:
Oh, nice.
Anthony Nielsen [01:27:22]:
It's like I could show you here, like—
Leo Laporte [01:27:26]:
Oh, that's cool. Look at that.
Anthony Nielsen [01:27:29]:
Yeah, but like, yeah, Via really changed the, um, the look. This is one.
Leo Laporte [01:27:41]:
So this was all from a prompt, or did you give it image references?
Anthony Nielsen [01:27:44]:
I gave it a A start and end frame. So like, start with, you know, clean, clean dot paper and then end here.
Danno [01:27:53]:
Nice.
Larry Gold (LrAu) [01:27:54]:
Yeah.
Leo Laporte [01:27:58]:
Making huge progress, I gotta say. Um, I mean, normally you would have done that in After Effects or something, and it'd probably take you half a day to do it.
Anthony Nielsen [01:28:09]:
Half a day?
Leo Laporte [01:28:10]:
That's, uh, 5 days.
Anthony Nielsen [01:28:14]:
Yeah, a couple days.
Leo Laporte [01:28:15]:
It's a lot of work.
Danno [01:28:16]:
Well, some days I get months of work done in a couple hours, and other days I spend all day messing with Buzz, getting nothing done.
Leo Laporte [01:28:23]:
Yeah, I spent all day making voices. Actually, the worst thing I've gotten into, and I gotta stop doing this, is, uh, model rotate—
Anthony Nielsen [01:28:33]:
benchmarking models, rotating different models in and And the problem, the problem with agents, it's— I think I told you the analogy before, but it's like having a home lab. Like, you end up— most of your work is like working on it versus the—
Leo Laporte [01:28:48]:
Yeah, exactly. Yeah, tuning the lab, getting the lab just right.
Timothy Ingalls [01:28:54]:
Leo, that was actually one of the things I was curious about, is if you could describe kind of the process you used for Doing your benchmarks?
Leo Laporte [01:29:04]:
I have a bake-off skill, uh, that I think Quicksilver and, uh, and Fable wrote together. I let Fable, uh— so I have a Groq is watch— is doing a model watch because it has access to X, and X seems to be where everybody posts the new recipes, right? But we also watch, um, there's Spark Arena. I watch a bunch of different Uh, websites. And so Grok, every— I don't know how often, every 10 minutes or so— goes out and sees not just what's new, but also the models that we are using, what new recipes. And there are a number of issues with some of these recipes. There's, you know, unexplained crashes, problems with Connectix or whatever. And so it has a— Uh, Claude's keeping a list of Showstoppers that we're looking for. So we won't run a model until we fix those.
Leo Laporte [01:30:02]:
Uh, and then, uh, and then so then what I'll do is, uh, if we've got, you know, I, I keep saying, you sure we shouldn't be using DeepSeek, uh, V4 Flash with vision instead of— that's the big 2 right now I'm considering, you know, I'm going back and forth on that between that and GLM-53 Flash, which also has And so I'll say, all right, well, let's try it. And so then I had— I've all along been writing benchmarks because I don't trust, and I think this is increasingly true, the stock benchmarks out there. People benchmark, they train their models on it and so forth. So I just don't trust those results. Plus they don't match my use. So I have enough use over the last 7 months of Hermes That it takes examples from work we've been doing, particularly TwitAds, and problems they've had and come across and things like authority approval, stuff like that. So there's a very quick— I probably have it here somewhere I can show you— a very quick 7-problem. It's almost like a logic set of problems that is just Just to just quickly see, is this, is this model even worth taking a look at? If it passes those, then we have another 20 or so tests.
Leo Laporte [01:31:25]:
Let's see if I can show this that we use. And so it gets staged up again and again, higher and higher. I have a whole benchmarks folder here. There's agentic benchmarks. Um, there's scorecards. So we have a 30-item— 7-item suite, a 30-item suite, and then I have a 7-item hard coding suite that's really pretty tough. And then we put these through the tests, and that gives me some idea of, uh, what it can do. And, uh, so we can grade it.
Leo Laporte [01:32:07]:
So, um, so these are bake-off results. I don't know where the, um, bake-off— oh yeah, here's the, here's the test. There's 7 raw thinking questions.
Craig McFarlane (CraigM) [01:32:21]:
And you're using this to, to figure out which one you want to go with at a certain point or for a certain activity?
Leo Laporte [01:32:27]:
Yeah, yeah. No, no, it's too much trouble to change it per activity. I, I What I'm looking for is the main model to run on my dual Sparks, but also what model to run on the, on the 3090, which is capable of running, uh, Quen, and what models I can run on the Max. So if they're contenders, they'll—
Craig McFarlane (CraigM) [01:32:47]:
Speed or quality or—
Leo Laporte [01:32:50]:
Mostly quality, but it's also speed. So we keep track of both, um, tokens per second, things like that. So right now, GLM Flash is still— it's a little slower. Oh, so for instance, this was a thread on X where they were going back and forth on a particular recipe, whether it was— and because there's all these influencers on X who are, you know, I got 800 tokens per second stuff. So there's a— so my only criterion is how does it do with the work? We do. And I don't do a lot of coding, so I'm not really looking so much at coding, um, because, uh, I'm still using the Frontier stuff for coding. So this is just a bunch of different, uh, benchmarks. And, uh, yeah, it, you know, I mean, it's just my little— but it— but they take a while.
Leo Laporte [01:33:49]:
They could take all night to run a bake-off between 2 different models because you got to load it in and then—
Craig McFarlane (CraigM) [01:33:57]:
I love a framework where it's actually using your real-world structure.
Leo Laporte [01:34:04]:
Yeah, because I really want to— I mean, I'm looking for what's going to work best with what. So because it really is the main model for Hermes, it needs to— the nice thing is Hermes has a record of everything we've ever done, so it's very easy. So I had— I can't remember who, generated these. I think— so what I— this is the other thing that I do mostly is, um, Hermes is the— is kind of the base model, but I run a bunch of models in their own frames. So this is Muse Spark running in, uh, Oh My Pi. This is Bard running in Antigravity. This is GLM 53 running in ZCode, ZAI's harness. This is Grok running in Grok Build.
Leo Laporte [01:34:55]:
This is Codex. So these are all in something called Herder, and because these are all panes in Herder, Quicksilver can talk to them directly and pass stuff around. So the other thing I've learned is not to let these guys talk to each other so much. Because they get into fights and they go back and forth and they start to dither around. So while they do have a chat that they can kind of sort of talk to each other in, the TwitAds thing is mostly handoffs. So this is Buzz. And so they can also send messages through Buzz. So they have either go directly to Herder or they can go through Buzz.
Leo Laporte [01:35:34]:
I like Buzz because I can— Buzz is like using Telegram or something. It's a messaging app, but it's, as I said, it's signed. So everybody has a signing key. So there's no question about authority. If I say something, it's, it's given. There's a lot of claudish, a huge amount of claudish in here, but it's useful to have this. And so this is them going back and forth. So when they're working on this is the TwitAds project.
Leo Laporte [01:36:01]:
As they work on this, one, one model will design, another model will review the design. Right now I think the designer is Fable, the reviewer is GLM53, the coder is Groq. Uh, then it goes from Groq, uh, to— well, I'm gonna probably start using, uh, Codex, but it was going from Groq, uh, to, um, gosh, I don't even remember. Another model. I, I'd like the idea of different models reviewing each other so that if there's you know, blind spots in one model. So the code, all the code is being written right now by Groq 4.6, but it could— but, you know, honestly, I now I'm intrigued by using maybe something like Luna maybe to write the code because these are pretty small slices.
Larry Gold (LrAu) [01:36:51]:
Why don't you try a Rinth? It'd be a good project for a Rinth.
Leo Laporte [01:36:53]:
Yeah, if you say— yeah, I mean, that's— yeah, that's what I'm thinking. Or even, you know, I've got Quen's running very fast on the 3090, so maybe even Maybe even having Quen— it's a dense Quen— do that would be another possibility.
Danno [01:37:08]:
Improvised by a better coder.
Leo Laporte [01:37:11]:
Yeah, well, exactly. I find that's a really good— and then first I was having them all talk and they were going back and forth. So now it's a very linear process which Quicksilver, Hermes manages. It's the coordinator. It's not doing any coding or planning or review or anything. It's just saying, okay, that slice is done, handing it— the review is done, handing it to the coder. Coder's done, handing it to the review. So it manages the workflow.
Darren Oakey [01:37:41]:
When I was first trying to— before there was chain of thought from these models and things, I've always had this vision that, you know, you could do anything by subdividing the task and breaking it up and everything, like a hierarchy, and you just split it up and give it to agents. And so before the whole Claude Code thing, I was trying that and it just never really worked because it would divide it into things like, do the HTML and do the JavaScript, but there was no communication between the two of them. So they had nothing to do with each other.
Danno [01:38:18]:
Right.
Leo Laporte [01:38:19]:
That's not good, obviously.
Darren Oakey [01:38:20]:
Didn't work at all.
Juan Hernandez (BlindWiz) [01:38:21]:
Right.
Anthony Nielsen [01:38:21]:
Yeah.
Darren Oakey [01:38:22]:
But then I tried something and this This is really early on. I tried the equivalent of the next step token generation. I just tried figure out the smallest thing that you could get done to improve this towards this goal and do it, and then it would just come back and do it again. That actually worked.
Leo Laporte [01:38:40]:
Okay.
Darren Oakey [01:38:40]:
Well, um, as, as you said, just one step after the other, and it's very similar to the next token generation. It just seems to be, you know, evaluate the next thing I need to do and do it. Now, obviously, models have gotten so good now and all sorts of things have changed, and now we've got Claude codes and things like this. But that— I still can't get the hierarchical breakup working. But as you say, doing it step by step works really well.
Leo Laporte [01:39:09]:
What I did way back when is I took the original code And, um, had, uh, at the time it was, um, was it Fable? Yeah, it was when Fable first came out. I had Fable, and I think I, I've showed this before. Let me, let me pull it up. Sorry, Larry, I keep putting you on camera. I don't know why it's you, but I do. Um, let me make this big. This was the plan. This was a spec that Fable wrote.
Leo Laporte [01:39:47]:
But then I said to— I had Fable do— hand it off to, um, 4-8 at the time to write everything. So this was— I trusted that Fable was slicing it intelligently. I didn't really look at it. It seems to have worked. I have a running program. Everything seems to work pretty well. And so this plan was pretty much fully executed right to the end of July. Now I'm sitting— so that's how it sliced it up.
Leo Laporte [01:40:28]:
I had Fable slice it up, Opus code it, and then I think I was having Saul GPT-5.6 review it. I think that was the first iteration. Now it's a little bit more complicated. Um, but yeah, I didn't even consider, oh gosh, are you dividing it up sensibly? I don't know, but it seems to have worked. So I hope that's the case. Now it's mostly UI stuff I have to work on. And then business logic is— sorry, Larry. Business logic is very complicated.
Leo Laporte [01:40:59]:
Uh, And the only way I get the business logic besides what was written in the original code, which is a little dated now, is going to Lisa and saying, okay, what happens if this happens and then that happens? And trying to understand what she says and then trying to translate it into something that the AI can understand. And that's kind of why I'm stuck right now because I've gotten I've gotten 90% of the way. It's these last— this last 10% is really killing me because it's hard to—
Juan Hernandez (BlindWiz) [01:41:31]:
Those final tweaks.
Leo Laporte [01:41:33]:
Yeah.
Juan Hernandez (BlindWiz) [01:41:33]:
Those little, little, little bits and bobs at the end that you have the— you have the, you know.
Leo Laporte [01:41:38]:
And it's the weird business logic that's left.
Juan Hernandez (BlindWiz) [01:41:41]:
Yeah.
Leo Laporte [01:41:41]:
Right?
Juan Hernandez (BlindWiz) [01:41:42]:
The edge cases.
Leo Laporte [01:41:43]:
It's the edge cases that's left. And they're very weird. You know, what if an advertiser gets halfway through its schedule and then quits? And then what happens to the— slots that they took, and it's very complicated. There's also a thing that was really complicated we've— I think we solved, which is fair rotation. They try with the ads— the presumption by some advertisers anyway is that the first ad is the best ad and the last ad is the worst ad. So we want to give every advertiser the equal placing throughout, uh, they call it fair rotation. So the advertiser will be in first, second, third, fourth, fifth, first, second, third, fourth, But when you have a bunch of advertisers and new advertisers coming in, doing that is complicated. And plus, once the ad has run, you can't move it, so it gets locked in.
Leo Laporte [01:42:34]:
It, it got a little complicated. I think I solved it. That was a, that was a tough one. It's fun.
Darren Oakey [01:42:41]:
As I said, so at work, no program is ever finished. Like, right, people keep asking me, when will this be finished? There's no such thing.
Danno [01:42:49]:
Right.
Darren Oakey [01:42:49]:
But also, often people— I think one of the biggest mistakes in programming is people, uh, and, and this is corporate programming, people go for the 100%. But right, if you say, how do we get all this working? Um, you say, oh, that'll be a 2-year project. How do we get 95% of these invoices processing perfectly? Oh, we can do that in 2 weeks. Right, right. But if you've always got in the system you've got a manual out or something, then you're— or some, something where you just go too hard or something, we're going to do this somewhere else or something like that. As long as you've got that facility to pull something out of the system or to manually override or something like that, then you don't have to get the 100%. And if you don't have to get the 100%, it often drastically simplifies the problem.
Leo Laporte [01:43:37]:
This is, uh, this is its current state right now. And you see I have the little bug report thing and I have it so that they can record or write into a box and then it gets sent to an agent. This is the— so it's all pretty well working. I mean, these are, you know, I can create a new order and, you know, say it's gonna be for the 4th quarter. But the deadline was yesterday or the day before yesterday because— Unfortunately, we, uh, we had the people we had doing our ad sales all year really didn't do the job. So, uh, we've had to, um, take it over. And Lisa put out the new rate card on September 1st. So now we're— oh, this bug is still here.
Leo Laporte [01:44:26]:
I have to— I said, what the hell? There's no dates. I need dates. And so there is a bug in it still. Uh, but this is, this is what— this is how they do it. There's a rate card in here. She can do a report on the sales. We have lists of advertisers. I periodically clear out the database.
Leo Laporte [01:44:46]:
There's a schedule, there's inventory available, so forth. So most of this is working. I mean, it's, it's, it's gotten pretty, pretty close. This was nice. There's a log audit of changes made. So we can go back and look at what happened. But this is the main— this new order thing is the main thing. There is still a bug.
Leo Laporte [01:45:11]:
I guess they didn't push the fix that I did this morning yet. So it's kind of— I mean, it's, it's 90%. I would say it's 90% there. At least it might not say that, but I, I feel pretty confident that we actually made something useful.
Danno [01:45:26]:
Well, you know, I, I use a web app kind of like that for my business, which I keep track of students.
Anthony Nielsen [01:45:33]:
Yep.
Danno [01:45:33]:
And I was in an acquisition meeting and explaining, you know, revenue and how many students or apprentices we have. And I just started talking to the data with Hermes because it knows the API.
Leo Laporte [01:45:48]:
Right.
Danno [01:45:49]:
So I said, you know, how many unique students each year for the last 3 years have submitted tests?
Darren Oakey [01:45:57]:
Right.
Danno [01:45:58]:
And it could just Whip that out in 10 seconds.
Leo Laporte [01:46:01]:
So that's— I'm debating right now whether I should put a little chatbot on here that can do exactly that, uh, so they can query it. And I want to do that, but I'm a little nervous about it. I don't have to figure out how to— maybe I'll use Ornis. I have to make it very dumb and simple.
Craig McFarlane (CraigM) [01:46:23]:
I'm sure you have an MCP interface too.
Anthony Nielsen [01:46:25]:
Yeah, like, didn't you— there's an API backend, right? Or Like, could there be—
Leo Laporte [01:46:28]:
There's a— so there's not an API backend for this. There is for, um, Drupal for— so it's— and this communicates with, with our, uh, production workflow, which the other one didn't, by the way. I mean, this is already like 20 times better than the old one. I just have to convince Lisa and Debbie of that. But, uh, yeah, I want— I mean, there's a database, and so it could be easy to have a chatbot query do exactly what you do, Dana, which is say, you know, How many, how many ad units did we sell this quarter and how many are unsold and what was the revenue? And all of that stuff should be fairly easy as data, basically database queries.
Anthony Nielsen [01:47:07]:
Well, yeah, if you could get to the database easily, then like you could just build a separate like Slack chatbot that doesn't need to live in there necessarily.
Leo Laporte [01:47:16]:
Oh, that's a good idea. Put it in Slack instead.
Anthony Nielsen [01:47:19]:
Yeah.
Leo Laporte [01:47:19]:
That's a great idea. That's what I'll do.
Larry Gold (LrAu) [01:47:22]:
And there's a lot of libraries that talk to your data. So what you do is you put the model into it you let it run through some data so it has some knowledge of it, and then you create the chatbot on top of that.
Leo Laporte [01:47:32]:
That's a good idea.
Darren Oakey [01:47:33]:
Just given an account that's read-only.
Leo Laporte [01:47:36]:
Right. No, I can't change anything. That's a great idea.
Anthony Nielsen [01:47:39]:
We can do that on our—
Craig McFarlane (CraigM) [01:47:41]:
Harder to go directly into the database because then you've got to train it up on the database schema. A lot of times there's weirdness in there, little tables that do weird things. Now it needs to understand it all. I'd go more the MCP route so that it has these tools and triggers and can understand the business views.
Leo Laporte [01:48:02]:
That'll be for phase 2.
Darren Oakey [01:48:04]:
Although either way, you can, if you, this is where the whole skills come in, is that you basically get Fable once to go and say, go and make an understanding of the schema and tell it.
Leo Laporte [01:48:17]:
And make 5 skills.
John Telford [01:48:19]:
Yeah.
Darren Oakey [01:48:21]:
Right, right, right.
Leo Laporte [01:48:23]:
All right, anybody have anything else they want to say before we, uh, we gotta, we gotta hit Larry's over-under here?
Danno [01:48:32]:
Let me get on Kaushee quick.
Leo Laporte [01:48:37]:
We did start early, so, so that's why it's gonna hit 2 hours. But, uh, 2 hours and 8 minutes according to the when I started the stream, so we're Well over our previous mark. Appreciate you guys. It's really, uh, I, I love doing this and getting together and talking about this stuff, and I think, uh, I think there's an audience for it. I think we have a good number of people watching right now who are silently participating. Uh, thank you Timothy and Juan and Craig and Dano and Larry and Darren. Special thanks to Anthony Nielsen. Who is the ever-present AI guru.
Leo Laporte [01:49:17]:
Uh, and apparently Google thinks you and I should go to a ball game and sample the food at the ballpark. Just hysterical. Um, anything else? Anybody? All right, I hereby— I wish I had a gavel— declare this meeting of the AI User Group—
Craig McFarlane (CraigM) [01:49:36]:
Adjourned.
Leo Laporte [01:49:37]:
Adjourned. Have a great day, everybody.