Hands-On AI 1 Transcript
Please be advised that this transcript is AI-generated and may not be word-for-word. Time codes refer to the approximate times in the ad-free version of the show.
Mikah Sargent [00:00:00]:
All right, here we go. Step 1, going up to the Wi-Fi and turning it off. No network? What's gonna happen next? Let me type something into this prompt here. I'm about to use a torque wrench for the first time. What do I need to know? Enter. All right, I'm asking Quen 3.5 what I need to know about using a torque wrench. And look, it's already thinking. It's processing.
Mikah Sargent [00:00:38]:
I remind you, there's no Wi-Fi, and yet the answer is coming in. No internet, still talking. This chatbot, it's not in the cloud, it's in this laptop. So exciting. I'm Micah Sargent, and this is Hands-On AI. Podcasts you love, From people you trust. This is TWiT. Hello and welcome to Hands on AI.
Mikah Sargent [00:01:13]:
I'm Micah Sargent, and I have been waiting for this. Oh, I'm so excited to kick this show off and get to talk with all of you. And our first episode is getting right into it. We are digging in to AI in a big way, and anyone with a laptop and a little bit of time can make this happen. We're going to talk about running an AI system on your machine. And, you know, you may be wondering why in the world would you want to run it on your computer whenever you could just go to the browser and talk to Gemini or ChatGPT or go on your phone and talk to Muse? Well, there are some really good reasons, so let's get into it. First and foremost, well, running it on your own machine means that your prompts stay on the device. Yeah, it's not getting sent away to some computer somewhere, some company somewhere.
Mikah Sargent [00:02:15]:
Instead, it's private. So what you type stays on the machine and what it says back never leaves the computer. I mean, you saw at the beginning there, I was able to talk with the chatbot and have it respond to me, right? Also, No account and no subscription means that you're only paying for the power that is required to run the AI system. You don't have to worry about, you know, running outta credits, paying every single month, uh, getting more capability and, and, and usage and all those, all those words. It's just easy peasy lemon squeezy. Oh, by the way, that also means that it works with the Wi-Fi off like we talked about. So if you're on a plane and you don't feel like paying for bad Wi-Fi network there. Well, you can just run it locally on your machine.
Mikah Sargent [00:03:09]:
The question, of course, is, is it as capable as the big cloud models? And well, this is where we have to get honest because it is a smaller model. It's gotta be stored locally on your device. I'm— I was running that on a MacBook Air. So yeah, it's gotta be smaller. Those big cloud assistants are a lot more powerful. And we will Get into and find out how all of that works in a few minutes. So how do we go and get this thing set up? Well, we'll dig into that next. One of the easiest ways to do it is with an app called Ollama.
Mikah Sargent [00:03:50]:
Yes, it's got a great name and it works on the Mac. This is Ollama, and I was using it earlier in order to Ask about a torque wrench. And by the way, it thought for 89.6 seconds, and here is its response about the torque wrench. So this is a machine— this is an app rather, uh, that goes online and it downloads models and it does check for updates, but your conversations stay put right here. These conversations that I'm having, it's not going anywhere. It's right here on my machine. And In order to download this model, in order to use it locally on your device, you do need to use the terminal. Gasp.
Mikah Sargent [00:04:33]:
But I promise you, I promise you, it's not going to be difficult. It's really just one line. So let's, let's dig into how we make use of that. You open up terminal, which again, you can hold down the Command key on your Mac, hit the spacebar, and Type in T-E-R and boom, the terminal is going to be there. And then we're going to download a specific model. So in this case, Ollama, after you've got Ollama installed and you will type in ollama pull and we're, we're installing Quen 3.5 and it is a 9 billion parameter model. Now the reason we're using this one is because this model is Perfect for this smaller machine. And despite being such a small model, still is quite capable for what it can do.
Mikah Sargent [00:05:31]:
So it'll download over time. Of course, this is, you know, sped up. 6.6 gigabytes downloaded to your machine. And once it gets there, well, then you can start chatting pretty quick. All you have to do is choose Ollama, run, Quen359B, and look, you can even just copy and paste that part, but that is what's required to get this thing working. Like, how cool is that? How cool that it's so simple to do, right? So neat. And when you're using this system, right, let's, let's kind of cut back to what it looks like. on our device.
Mikah Sargent [00:06:17]:
You know, I have it selected here in Ollama, and now I could type in something like, how many Rs in strawberry? This is a very common test for models and their capabilities, and we'll see how this does. It starts to talk about it, and it says, you know, we're going to analyze the request, All of this is the thinking that the model is doing in the background. It's spelling out the word strawberry. It's looking at the recurrences of the letters. It seems to think that there are 3. There are. And it goes letter by letter saying, no, this is not an R. Yes, this is an R.
Mikah Sargent [00:06:59]:
Yes, this is an R. It's also checking for trickery. It's asking if it's a joke. It's doing all of this work to make sure That it gets 3 as the answer and that it does not overdo it. It even takes some time to think about it and go, let me just go ahead and double-check. I just want to be absolutely certain before I give the answer. And I think that's pretty incredible that again, this is just happening locally on the device. It's not needing to do a Google search to try to come up with this, anything like that.
Mikah Sargent [00:07:33]:
It's even doing some refinement and self-correcting right there in the chat and being able to see. The thinking, so to speak, that the model is doing is a really cool aspect of that. So once we're in the, the app, of course, you're able to work with this model. You can also choose different models. You can download other models and search for models right there. The terminal is the easiest way to get this going. If you know the exact model name, All you have to do is type it into the picker there and then it will download for you. Now, every model comes with a label and it, you know, to be honest, it looks like alphabet soup.
Mikah Sargent [00:08:17]:
You can see it there, Quen359B, and that can be a little confusing. It's hard to kind of break that apart. So I want to figure out exact— I want to help you figure out exactly what this means. So let's take a look at a model direct from the Ollama library, and you will walk away understanding how these things are broken down. So the model is Quen 3.5. It's Quen from Alibaba, version 3.5. The size is the 9B part. That's what 9B part— it means that it has 9 billion parameters.
Mikah Sargent [00:09:02]:
You can sort of think of them as like the knobs that the model learned to set during training. The more knobs it has to sort of work with, the smarter the model, but it also means that the file will be bigger. And then frankly, this is my favorite part, quant or quantization. You can sort of, for nerds out there, think of it as a bitrate for AI. Each of the numbers gets squeezed down to about 4 bits. Q4. The 8-bit version of this model is 11 gigabytes. Okay.
Mikah Sargent [00:09:42]:
And that's where that size comes in. The number that matters the most, 6.6 gigabytes. That's what has to fit in your computer's memory, which then makes me go and probably makes you go, will it fit? I think that the best thing you can do is start with a model file that's about half of your memory. Okay? So that is going to give room for everything else that you need to have your computer doing, that it's not all being taken up by this model. And so if you have 8 gigs of memory. Well, we want a 4 gig file. If you have 16 gigs of memory, obviously you want an 8 gig file. So look for the 4B version of Quen 3.5.
Mikah Sargent [00:10:41]:
It's called Quen 3, or it's called 3.4. If you've got 16 gigs, again, 8 gigs. Uh, so that's 6.6, and that will fit with plenty of room to spare. If you've got 32 gigs of memory, you can guess where that's going to go. That gives you so much space to do what you want. On a PC with a graphics card, you are able to, to go by the card's memory instead. And so you don't necessarily have to worry as much about how things are sort of taking up. So I think it's time to put the local AI model to the test and see how it compares to a more powerful model that exists online.
Mikah Sargent [00:11:26]:
So I've decided today to use a rather available tool that a lot of people have access to, which is Google's Gemini, and compare that with the local Quen model that we're running on our MacBook Air, an old MacBook Air with 16 gigs of RAM, and see how they compete. So to do that, we're going to do some tests, and these are just a set of prompts. So I'll read each one out loud so you can play along if you would like to also participate. Let's start with emails. So I'm going to give each of the models an email. And this email, funny enough, it came from my local— from— it came from my Muse, which is of course another AI system. I set it up with its own email. And I said, I need you to send me a short story.
Mikah Sargent [00:12:26]:
And, uh, so it wrote me this little short story. Micah, once there was a small mossy creature who lived inside a mailbox, et cetera, et cetera, et cetera. So here's how we're going to test these 2 models. The prompt is going to be, summarize this email in 3 sentences, then tell me if I need to reply today. So we'll head over to macOS and give this a shot. So let me copy the email and then we'll start with Quen via Ollama. And I will say, summarize this email in 3 sentences, then tell me if I need to reply. Today.
Mikah Sargent [00:13:19]:
And then I will paste the prompt. We'll go ahead and copy that and I will hit enter and then I will send it over to Gemini. And so first we can see that immediately Gemini has our answer while we can see that Quen is taking some time, but it is going through the process. It is thinking about things. It's analyzing the request while we wait. For, uh, for Quen to respond, let's go ahead and take a look at what Gemini had to say. Uh, Gemini says in response, Marshall sent a short whimsical fable about a mossy creature living in a mailbox who cherishes incoming letters. The story concludes with the creature receiving a note confirming it works.
Mikah Sargent [00:14:03]:
In closing, Marshall clarifies that the message was simply a test to confirm the mailbox is functional. Need to reply today? No. The email is a playful test confirmation and contains No questions or urgent action items. And now we'll wait for Quen to finish. You know, this could be a feature. Local models give you time to get up and move around. It's sort of like commercials in television. Oh boy.
Mikah Sargent [00:14:39]:
Sometimes users paste nonsense to see if the AI hallucinates. The user previously asked about torque wrenches and then switched to counting Rs in strawberry. Oh my word. I'm gonna hit pause on the recording. It's gonna be at this for a while. I have to tell you, a local model's got a lot of space to sort of spread out because it doesn't have to worry about tokens. And it did a lot. It did a lot of thinking in the background and was very paranoid, I have to say, about me testing it.
Mikah Sargent [00:15:12]:
It did say, this email recounts a whimsical tale about a mossy creature inside a mailbox who evaluates every incoming letter for both its message and the weather it traveled through. Upon receiving a note that simply states, it works, the creature becomes content and returns to its usual routine of waiting. The story concludes with a lighthearted confirmation that everything is functioning as intended. No, you do not need to reply today because this appears to be a creative narrative or joke without any deadlines or action items attached. So speed is the biggest difference between the two, but they both arrived at similar summaries and similar answers. And in my opinion, are both correct. No, I don't need to worry about replying today. All right, let's move on to the next test.
Mikah Sargent [00:16:06]:
This time in round 2, we're going to take a look at this week's news. And here's what's great about this. There's, there's no way for Quen to tell me about this week's news because it doesn't have access to the internet. So I'm going to simply ask, what were the biggest tech news stories this week? All right, so let's pop this into Gemini. What were the biggest tech news stories this week? And into a new Quen 359B chat here. And Boom. Google's already got— Gemini's already got some answers for me. Autonomous agents under scrutiny.
Mikah Sargent [00:16:46]:
Yes, that does continue on. OpenAI facing fallout over rogue web crawling agents. Hardware and smart home resets. Apple expanding its hardware lineup. Google resetting its smart home ecosystem by phasing out legacy Nest stuff. Aviation and critical infrastructure AI. And then it even asks, If you want a deeper breakdown of the security concerns surrounding autonomous AI agents. So a little bit of a follow-up at the end there.
Mikah Sargent [00:17:16]:
And I have to tell you, I'm very excited. Last round, Quen took a long time to respond to me. This time, Quen was quick to respond and said, I appreciate your question. However, as an AI model, my knowledge is current only up to 2024, and I don't have real-time access to news or events after that. For recent news, I recommend checking, and then it lists a few different sites where I can check that news. It goes on to say, let me know if you'd like insights on a specific topic, for example, AI, space tech, or sustainability, and even includes a little emoji afterward. So I have to say, I'm appreciative of the fact that it was able to come up with a way to be helpful, even though it had no access to the internet and therefore could not answer my question. I have to, of course, give this round to Google Gemini, but I am pleasantly surprised at what Quen was able to do without any further access to tools that it might need to do so.
Mikah Sargent [00:18:29]:
Now, our final test is a little bit of a test. This one is confusing, And also complex. And I have to imagine that it's going to take Quen a little while to get it right, whereas Gemini will probably hop on. This is a plan for the afternoon. And the prompt is going to be, it's 10 AM. I need to walk the dogs for 45 minutes, bake bread that rises for 3 hours and then bakes for 40 minutes, and be on a video call from 2 to 3 PM. Give me a schedule where nothing overlaps. This is math.
Mikah Sargent [00:19:10]:
This is time math, which I hate. This is multiple bits of context all at the same time. There's a lot involved here, so it should be interesting to see how Quen handles things against Gemini with its greater capabilities. I've popped the prompt into each of these models, and we will give it a go. Starting with Gemini, uh, in absolute mere seconds, it had a response for me. It gave me a graph, uh, or a table rather, that has all of the information from 10 to 10:15. What activity will you do? You'll mix and knead the bread dough. And then it says that there's a note, start the 3-hour rise by 10:15 AM.
Mikah Sargent [00:20:03]:
From 10:15 AM to 11 AM, well, go ahead and walk the dogs. That's 45 minutes while the dough is rising. From 11 to 1, free time while the dough continues to rise. 1:15 to 1:20, shape the dough and preheat the oven. And by the way, I'm reading out this prompt because I can see that Quinn is still working in the background. With the rise ending at 1:15 PM, so that makes the 3 hours elapsed. From 1:20 to 2, bake the bread. 40 minutes in the oven, out before your call.
Mikah Sargent [00:20:34]:
And then from 2 to 3, video call. Dedicated call block with zero kitchen tasks. And even a little note at the bottom: this schedule assumes a 15-minute window from 10 to 10:15 AM to mix and knead the dough before the rise starts, and leaves the finished loaf cooling safely while you take your call at 2 PM. Very good! Let's see how Quen is doing. Well, it's going to take a minute, so we'll let it churn on this for a bit. Mind you, in this instance, Gemini immediately had its response. Well, I have to tell you, I, I had hoped, I had hoped that eventually Quen was going to get us somewhere. But some 30 minutes later, and Quen is still thinking about the bread and the dogs and the walking.
Mikah Sargent [00:21:40]:
And so the fact is that for some tasks, it's just not possible for these small models to be able to do what's needed. But Then what does that leave us with, right? What, what, what is the use of these small local models? Well, if you need help with, you know, everyday writing, you need help with coding, stuff that has a lot of patterns to it, then that is where these local models can be helpful. If you need to know what happened yesterday, and if you need help with complicated scheduling and things that require a lot of context at a given time, Well, that is where the cloud still wins. But each and every day, I mean, mind you, that is a 6 gigabyte file on an M2 MacBook Air that is spitting out all of this thinking. Yeah. And coming up with answers. It's kind of mind-boggling. It's bizarre.
Mikah Sargent [00:23:01]:
And so one can only imagine that these models will continue to get better. It's just this tiny little thing sitting on your drive. Everything that the chatbot knows fits in that one file. How is it even possible? That's next week, and I'm really excited to talk about that. But until then, I've got some homework for you. Go ahead, download Ollama locally to your device. Pick a model that fits your machine. We talked about the math that you should do there, half of what's available.
Mikah Sargent [00:23:50]:
And then ask it something with the Wi-Fi off just so you could tell that it's definitely working locally and it's not tricking you. And then tell me how it went. I wanna hear how you're using local models. And in the meantime, and then some, send your AI questions to handsonaia@twit.tv and you just might hear yours answered on the show. Hit subscribe or follow wherever you're watching or listening. And of course, head to twit.tv/HOAI. If you want the show ad-free, well, that's where Club Twit comes in. twit.tv/ClubTwit.
Mikah Sargent [00:24:30]:
I have been and will continue to be Micah Sargent. And I think I'm going to end the show with a little tagline every time. Because it's fun. Don't just wonder about AI, get your hands on it. I'll see you next time.