Root / Newsroom / Compute & AI
Build Log · Compute & AIChoosing a Local LLM for Knowledge Extraction (2026)
Speed vs. quality on a 16GB card. A field-tested guide from the Frontier // Decoded build sessions — turning a library of books, papers and transcripts into a private, searchable knowledge vault, entirely on your own hardware, privately, and for free.
This whole guide came out of a real problem: I had a mountain of documents — books, PDFs, saved articles, video transcripts — and wanted them all turned into one structured, searchable brain that runs on my own machine. No cloud, no subscription, no data leaving the desk. The rig: a single 16GB GPU (RTX 4070/4080 class) feeding a local stack — LM Studio for the model, an embedding model for retrieval, faster-whisper for audio, all wired into a knowledge vault. Here's what actually worked, and what didn't.
01 / FRAMEThe task shapes the model choice
Not all LLM work is the same, and the "best" model depends heavily on what you ask it to do. This guide is specifically about knowledge extraction: reading source text and pulling out every fact, figure, date, entity, definition and claim into a structured note.
That matters because extraction is a grounded task. The facts already exist in the text; the model's job is to find, organize and faithfully reproduce them — not to invent, reason from scratch, or write creatively. That single property changes everything about model selection:
→ Grounded work rewards instruction-following and structured-output reliability far more than raw reasoning horsepower.
→ Because the answer is in the source, smaller, more efficient models punch well above their weight — a 12B model rarely loses to a 70B at faithfully listing what a paragraph says.
→ The pipeline makes many calls — one per text segment, across thousands of documents — so throughput dominates total time. A model that's twice as fast is worth more than one that's five percent "smarter."
The consequence: for extraction, the sweet spot is a fast, well-aligned mid-size model — not the largest model you can technically load.
02 / ARCHITECTUREWhy Mixture-of-Experts wins speed and quality
The single most important architectural idea for getting both speed and quality out of local hardware is Mixture-of-Experts (MoE).
A dense model runs every one of its parameters for every token it generates. A 30B dense model does 30B parameters' worth of math per token — slow. An MoE model of the same nominal size is split into many "expert" sub-networks, and a router activates only a few per token. A "30B-A3B" model has 30 billion total parameters but only about 3 billion active per token.
The result is the best of both worlds:
For a many-calls task like extraction, MoE is the clear architectural winner — and it's the reason the top picks below are MoE designs.
03 / PICKSRecommended models for a 16GB card
Assuming roughly 16GB of GPU memory and plenty of system RAM, ranked for the extraction use case as of mid-2026. Use Q4_K_M quantization as the default — the standard sweet spot for fitting quality models into 16GB.
-
Qwen3-30B-A3B-InstructMoE
The overall speed-plus-quality champion. Thirty billion parameters of knowledge, 3B active per token. At Q4 it's roughly 18GB, so a few gigs spill into system RAM, but the tiny active footprint keeps it fast. Best accuracy-per-second for extraction. Larger successors in the family (e.g. a 35B-A3B) are stronger still if they fit.
-
Qwen3-14B-InstructDense
The best "no compromises, fits entirely in VRAM" pick. About 9GB at Q4, 35 tokens/sec on a 4080-class card, fully GPU-resident — zero RAM-offload penalty. Excellent instruction-following and clean structured output. If you want maximum reliability with everything on the GPU, this is it.
-
Gemma 4 12BDense + 26B-A4B-QAT MoE
Google's Gemma line is a standout for instruction-following, predictable formatting and chat tone — exactly what structured note-writing needs. The 12B dense version is a rock-solid all-rounder. The 26B-A4B-QAT MoE variant (26B total, 4B active, quantization-aware trained) delivers near-26B quality at roughly 12B speed and fits cleanly — an outstanding balance.
-
Phi-4 / Phi-4-ReasoningDense
Top-tier at instruction-following benchmarks and very predictable structured output — a strong choice when tidy, well-formed notes are the priority.
For maximum throughput when plowing through a very large backlog where nuance matters less, a small fast model such as Gemma 4 e4b or a 3–4B Qwen keeps accuracy high on grounded extraction while maximizing books-per-hour.
04 / CAUTIONA model to be careful with: gpt-oss-20b
⚠ Pipeline hazard
On paper, gpt-oss-20b is attractive — an MoE design that generates faster than many dense 14B models. In practice, running it through LM Studio for long, structured generation surfaced a real problem: intermittent HTTP 400 errors reading "the model produced output that does not match the expected peg-native (harmony) format."
The cause is LM Studio's harmony-format parser choking on some of the model's output during long structured generations. It's stochastic — short prompts usually work, but a full multi-section note fails often enough to stall a pipeline. You can work around it by retrying (the resampled output frequently parses), but for an unattended job that's friction you don't need. For structured extraction specifically, prefer the Gemma or Qwen options above.
05 / OPSPractical lessons from running this at scale
Model choice is only part of the story. Several operational details made a bigger difference than swapping models:
Read the whole document, not the first page
The biggest accuracy trap was silently truncating each source to the first 12,000 characters before sending it to the model — meaning a 300-page book was summarized from roughly its first 2%. The fix: chunk the entire text into 8,000-character segments, extract exhaustively from each, and merge the results into one complete note. This multiplies model calls per document — which is exactly why fast MoE models matter.
Keep models pinned in memory
LM Studio's default "just-in-time" loading unloads a model after an idle timeout, so it reloads (and re-reads gigabytes from disk) on the next call. Pinning the working model plus the embedding model — with no auto-unload TTL — removed repeated multi-second load stalls. Beware duplicate instances: calling "load" again on an already-loaded model can spin up a second copy and blow past VRAM.
Mind VRAM headroom for the embedder
If you also embed for retrieval (RAG), keep one small embedding model resident alongside the chat model. Two large models fighting over 16GB forces heavy RAM offload and tanks speed.
Whisper for audio & video
To bring spoken content into the same pipeline, faster-whisper (GPU-accelerated) transcribes video and audio to text, which then flows through the identical extraction step. ffmpeg pulls the audio; yt-dlp fetches links from the web. Whisper tiers map cleanly to a quality dial: large-v3 (best), medium (balanced), small (faster), base (quick).
Extraction vs. verbatim storage
A structured extraction is a compression of the source — richer than keywords, but not the full text. To truly lose nothing, embed the complete raw text into your search index in addition to the structured note. Then retrieval can quote exact passages, while the note gives the human-readable breakdown.
06 / BOTTOM LINEThe short version
For turning a library into a private, searchable knowledge base on a 16GB card in 2026:
→ Best single pick: an MoE in the Qwen3-30B-A3B class (or your platform's equivalent, e.g. Gemma 4 26B-A4B-QAT) — big-model quality at small-model speed.
→ Best fully-in-VRAM pick: Qwen3-14B or Gemma 4 12B — fast, reliable, no offload.
→ Architecture rule of thumb: for extraction, favor MoE and mid-size instruction-tuned models over the largest dense model you can technically fit.
→ The model is not the bottleneck people think it is — reading the full document, pinning models, and using an efficient architecture matter as much as which specific model you download.
Part of the Frontier // Decoded Compute & AI build log — notes from wiring a private, local knowledge vault. Field-tested, and still growing.