Run an AI Model From a Pen Drive (Windows, Fully Offline)
A portable AI setup: one executable, one model file, one launcher script. No installation, no Python, no internet after setup. Plug the drive into any Windows machine and run.
What You Need
| Pen drive | 8 GB free minimum, USB 3.0 recommended |
| Free RAM | ~4 GB |
| OS | Windows 10 or 11 |
| Internet | Only for the two downloads below |
An external SSD works with the exact same steps and runs noticeably faster.
Step 1 — Format the Pen Drive as exFAT
Right-click the drive → Format → File system: exFAT → Start.
Formatting erases everything on the drive. Back up first.
exFAT is required because FAT32 caps individual files at 4 GB, which breaks larger models.
Step 2 — Download llamafile
https://github.com/mozilla-ai/llamafile
Go to Releases → Assets → download the main llamafile binary (v0.10.5 used here).
Copy it to the pen drive.
Step 3 — Add the .exe Extension
Windows won’t execute the file until it’s renamed.
- In File Explorer: View → check File name extensions
- Rename the file to:
llamafile-0.10.5.exe - Accept the “file might become unusable” warning
Step 4 — Download the Model
https://huggingface.co/pramodlohra/Qween3_4B_thinking_finetune/tree/main
Open the Files tab and download:
qwen3-4b-thinking-2507.Q4_K_M.gguf
Direct download link:
https://huggingface.co/pramodlohra/Qween3_4B_thinking_finetune/resolve/main/qwen3-4b-thinking-2507.Q4_K_M.gguf?download=true
Size: 2.5 GB. Copy it to the pen drive, in the same folder as the .exe.
Your drive should now contain:
llamafile-0.10.5.exe
qwen3-4b-thinking-2507.Q4_K_M.gguf
Step 5 — Create the Launcher
Right-click on the pen drive → New → Text Document. Paste this in:
@echo off
cd /d "%~dp0"
llamafile-0.10.5.exe --server --model qwen3-4b-thinking-2507.Q4_K_M.gguf
pause
Save, then rename the file to ai.bat and accept the extension warning.
What each line does:
cd /d "%~dp0"— runs from the folder the script lives in. Without this, the script breaks whenever the pen drive gets a different drive letter on a different computer.--server— starts the local web server and chat UI--model— points to the GGUF filepause— keeps the window open so you can read any error instead of it flashing shut
If your filenames differ, edit lines 3 to match exactly.
Step 6 — Run It
Double-click ai.bat.
Wait for the load to finish (slower on first run — it’s reading 2.5 GB off USB), then open:
http://127.0.0.1:8080
Leave the terminal window open. Closing it stops the server.
Optional Flags
Change the port if 8080 is taken:
llamafile-0.10.5.exe --server --model qwen3-4b-thinking-2507.Q4_K_M.gguf --port 8081
Try GPU offload (works only if your build includes a matching backend; falls back to CPU otherwise):
llamafile-0.10.5.exe --server --model qwen3-4b-thinking-2507.Q4_K_M.gguf -ngl 999
Do not add --host 0.0.0.0 unless you know what you’re doing. The default binds to localhost only, which keeps the server off your network.
Troubleshooting
Terminal flashes and disappears — you’re missing pause, or the filenames don’t match. Run dir in the folder to check exact names.
'.' is not recognized as an internal or external command — you used ./file.exe. Windows CMD needs .\ or no prefix at all. Use the script above as written.
Model not found — cd /d "%~dp0" is missing, or the .gguf didn’t finish downloading. Confirm the file is 2.5 GB.
Loads then closes — not enough free RAM. Close other apps and retry.
Very slow generation — USB 2.0 port or a slow drive. Copy the .gguf to your internal drive and point --model at that path.
Warnings in the terminal about CORS or control-looking token — harmless. The first is expected on a localhost server; the second is a metadata quirk in the GGUF that doesn’t affect output.
Notes
The model runs entirely on your CPU. Text you type never leaves the machine.
A 4B quantized model is capable but limited compared to large cloud models, and it can be confidently wrong. Verify anything that matters.
llamafile is Apache 2.0 licensed. Model licenses are listed on their respective Hugging Face pages.