Skip to content
-
Subscribe to our newsletter & never miss our best posts. Subscribe Now!
AI Brainbox logo AI Brainbox logo AI BrainBox

AI Tools | Automation | Workflows

AI Brainbox logo AI Brainbox logo AI BrainBox

AI Tools | Automation | Workflows

  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy
  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy
Close

Search

Trending Now:
Free AI Voice Generator Claude code for free Free AI tools Local AI models
  • FaceBook
  • X
  • LinkedIn
  • Instagram
  • YouTube
Subscribe
AI Brainbox logo AI Brainbox logo AI BrainBox

AI Tools | Automation | Workflows

AI Brainbox logo AI Brainbox logo AI BrainBox

AI Tools | Automation | Workflows

  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy
  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy
Close

Search

Trending Now:
Free AI Voice Generator Claude code for free Free AI tools Local AI models
  • FaceBook
  • X
  • LinkedIn
  • Instagram
  • YouTube
Subscribe
Home/AI Tools/Run an AI Model From a Pen Drive (Windows, Fully Offline)
local ai from usb pen drive
AI ToolsAI AutomationAI CodingLocal AIOpen Source AI

Run an AI Model From a Pen Drive (Windows, Fully Offline)

By AI BrainBox
August 16, 2026 3 Min Read
0

A portable AI setup: one executable, one model file, one launcher script. No installation, no Python, no internet after setup. Plug the drive into any Windows machine and run.


What You Need

Pen drive8 GB free minimum, USB 3.0 recommended
Free RAM~4 GB
OSWindows 10 or 11
InternetOnly for the two downloads below

An external SSD works with the exact same steps and runs noticeably faster.


Step 1 — Format the Pen Drive as exFAT

Right-click the drive → Format → File system: exFAT → Start.

Formatting erases everything on the drive. Back up first.

exFAT is required because FAT32 caps individual files at 4 GB, which breaks larger models.


Step 2 — Download llamafile

https://github.com/mozilla-ai/llamafile

Go to Releases → Assets → download the main llamafile binary (v0.10.5 used here).

Copy it to the pen drive.


Step 3 — Add the .exe Extension

Windows won’t execute the file until it’s renamed.

  1. In File Explorer: View → check File name extensions
  2. Rename the file to: llamafile-0.10.5.exe
  3. Accept the “file might become unusable” warning

Step 4 — Download the Model

https://huggingface.co/pramodlohra/Qween3_4B_thinking_finetune/tree/main

Open the Files tab and download:

qwen3-4b-thinking-2507.Q4_K_M.gguf

Direct download link:

https://huggingface.co/pramodlohra/Qween3_4B_thinking_finetune/resolve/main/qwen3-4b-thinking-2507.Q4_K_M.gguf?download=true

Size: 2.5 GB. Copy it to the pen drive, in the same folder as the .exe.

Your drive should now contain:

llamafile-0.10.5.exe
qwen3-4b-thinking-2507.Q4_K_M.gguf

Step 5 — Create the Launcher

Right-click on the pen drive → New → Text Document. Paste this in:

@echo off
cd /d "%~dp0"
llamafile-0.10.5.exe --server --model qwen3-4b-thinking-2507.Q4_K_M.gguf
pause

Save, then rename the file to ai.bat and accept the extension warning.

What each line does:

  • cd /d "%~dp0" — runs from the folder the script lives in. Without this, the script breaks whenever the pen drive gets a different drive letter on a different computer.
  • --server — starts the local web server and chat UI
  • --model — points to the GGUF file
  • pause — keeps the window open so you can read any error instead of it flashing shut

If your filenames differ, edit lines 3 to match exactly.


Step 6 — Run It

Double-click ai.bat.

Wait for the load to finish (slower on first run — it’s reading 2.5 GB off USB), then open:

http://127.0.0.1:8080

Leave the terminal window open. Closing it stops the server.


Optional Flags

Change the port if 8080 is taken:

llamafile-0.10.5.exe --server --model qwen3-4b-thinking-2507.Q4_K_M.gguf --port 8081

Try GPU offload (works only if your build includes a matching backend; falls back to CPU otherwise):

llamafile-0.10.5.exe --server --model qwen3-4b-thinking-2507.Q4_K_M.gguf -ngl 999

Do not add --host 0.0.0.0 unless you know what you’re doing. The default binds to localhost only, which keeps the server off your network.


Troubleshooting

Terminal flashes and disappears — you’re missing pause, or the filenames don’t match. Run dir in the folder to check exact names.

'.' is not recognized as an internal or external command — you used ./file.exe. Windows CMD needs .\ or no prefix at all. Use the script above as written.

Model not found — cd /d "%~dp0" is missing, or the .gguf didn’t finish downloading. Confirm the file is 2.5 GB.

Loads then closes — not enough free RAM. Close other apps and retry.

Very slow generation — USB 2.0 port or a slow drive. Copy the .gguf to your internal drive and point --model at that path.

Warnings in the terminal about CORS or control-looking token — harmless. The first is expected on a localhost server; the second is a metadata quirk in the GGUF that doesn’t affect output.


Notes

The model runs entirely on your CPU. Text you type never leaves the machine.

A 4B quantized model is capable but limited compared to large cloud models, and it can be confidently wrong. Verify anything that matters.

llamafile is Apache 2.0 licensed. Model licenses are listed on their respective Hugging Face pages.

Tags:

ai from pen drivelocal aiportable ai
Author

AI BrainBox

Follow Me
Other Articles
Previous

Unsloth Desktop on Windows: the setup guide I wish I’d had

Next

Ornith-1.5: The Open Model That Beats Claude

Recent Posts

  • Ornith-1.5: The Open Model That Beats Claude
  • Run an AI Model From a Pen Drive (Windows, Fully Offline)
  • Unsloth Desktop on Windows: the setup guide I wish I’d had
  • Run Claude Code on Free AI Models with OmniRoute (Beginner Setup Guide)
  • DeepSeek V4 Flash 0731 Just Dropped | Run it FREE | It’s Really Insane
Hey, I’m Umair. I’m a Data Scientist passionate about AI, exploring emerging technologies, and sharing practical tutorials that help people work smarter with AI.
  • X
  • Instagram
  • Facebook
  • YouTube
Open to AI Projects
Get In Touch

Recent Posts

  • Ornith-1.5: The Open Model That Beats Claude
    by AI BrainBox
    August 21, 2026
  • clone-voice-free-voicebox-tutorial
    Clone Any Voice for Free with Voicebox: Full 2026 Guide
    by AI BrainBox
    June 26, 2026
  • chatgpt-prompt-engineering
    ChatGPT Prompt Engineering for Beginners: How to Get Better Results Every Time
    by AI BrainBox
    June 27, 2026
  • how-to-run-ai-locally-ollama-guide
    How to Run AI Locally on Your PC with Ollama (No Cloud, No Subscription)
    by AI BrainBox
    June 27, 2026
ai brainbox

We're exploring the latest and greatest AI tools and techniques, providing you with everything you need to know to keep up with this rapidly evolving field. Passionate about making AI accessible and understandable to everyone.

  • Facebook
  • X
  • Instagram
  • LinkedIn

Latest Posts

  • Ornith-1.5: The Open Model That Beats Claude
  • Run an AI Model From a Pen Drive (Windows, Fully Offline)
  • Unsloth Desktop on Windows: the setup guide I wish I’d had
  • Run Claude Code on Free AI Models with OmniRoute (Beginner Setup Guide)
  • DeepSeek V4 Flash 0731 Just Dropped | Run it FREE | It’s Really Insane

Menu

  • Home
  • Tutorials
  • About
  • Contact
  • Privacy Policy

Contact

Email

contact@aibrainbox.io

Location

New York, USA

Copyright 2026 — AI BrainBox. All rights reserved.