Fit the model to the PC
Record RAM, GPU and VRAM, storage, and normal background use. Start with a model class and context that leave the computer responsive rather than choosing the largest file that can technically load.

LOCAL AI FOR HOME & WORK
A Plain-English Guide to Running Useful Local AI at Home or Work Without Sending Every Prompt and Document to the Cloud
By Mercer Lane · Published by Mercer Lane Press
A plain-English guide to running useful local AI on a Windows PC, choosing models that fit the hardware, understanding what stays local, working with private documents, troubleshooting slow or unstable setups, and deciding when the cloud is still the better tool.
WHO IT IS FOR
Home users, independent professionals, and small-business readers who want practical local AI for drafting, summarizing, extraction, document questions, and other everyday work while making deliberate choices about privacy, hardware limits, and when cloud AI is more appropriate.
THE OPERATING PRINCIPLE
The method begins with the machine you already own and treats privacy, capability, speed, and maintenance as separate questions that need separate checks.
Record RAM, GPU and VRAM, storage, and normal background use. Start with a model class and context that leave the computer responsive rather than choosing the largest file that can technically load.
The main path uses a graphical desktop application, with LM Studio as the worked example, while Jan and Ollama are presented as alternatives for readers with different preferences or integration needs.
Local inference is only one part of the path. Check the active model, app data folder, cloud sync, backups, web tools, plug-ins, remote access, and who else can use the computer.
Drafting from supplied facts, extraction, rewriting, transformation, and source-only document questions are treated as repeatable jobs with explicit rules for missing information and human review.
Close unnecessary apps, start a fresh chat, reduce context, adjust GPU offload where relevant, try a smaller quantisation or model, and only then consider hardware.
Use the same small test pack when comparing models or after significant updates so a working setup changes because it improves real work, not because a new model is fashionable.
PRIVATE DOCUMENTS
Long files may be split into chunks and only selected passages may be sent to the model for a particular answer. The model may not be considering every page every time.
Tell the model to answer only from the attached document, preserve numbers and dates, and say NOT FOUND when the source does not support an answer.
Ask for facts you know are present, facts that are absent, and—where appropriate—conflicting values across controlled test documents before trusting the workflow broadly.
FREE SUPPORTING GUIDES
LOCAL AI GUIDE
Start with the PC you already own. Record RAM, GPU/VRAM and free storage, then choose a model class that leaves useful headroom instead of loading the largest model that can technically start.
Read the guideLOCAL AI GUIDE
Use one desktop application, one modest model, and one repeatable task. Download, load, test, verify, and record what worked before exploring advanced settings.
Read the guideLOCAL AI GUIDE
Local inference answers where model generation happens. It does not answer where chats, files, backups, web tools, plug-ins, or other copies go. Follow the data path before making a privacy claim.
Read the guideLOCAL AI GUIDE
Document chat can be useful without sending the source to a remote AI inference service, but retrieval may show the model only selected passages. Test what it can find before trusting what it summarizes.
Read the guideLOCAL AI GUIDE
Change the cheapest variable first. Close heavy apps, start a fresh chat, reduce context, check GPU offload, try a smaller quantisation or model, then consider updates or hardware.
Read the guideSEARCH BY THE PROBLEM
People searching for local AI on a PC, Ollama, LM Studio, private document chat and running an LLM locally usually need two answers: what their hardware can handle and what actually stays private. This guide addresses both.