LLM Wiki Pattern¶
Source metadata (verified 2026-09-25)
| Field | Value |
|---|---|
| Title / file | "llm-wiki" gist, file llm-wiki.md, described as "A pattern for building personal knowledge bases using LLMs" |
| Author | Andrej Karpathy (karpathy on GitHub) |
| Published | 2026-04-04, as a follow-up "idea file" to his 2026-04-02 X post "LLM Knowledge Bases" |
| URL | gist.github.com/karpathy/442a6bf555914893e9891c11519de94f |
| Engagement | ~48.6k stars, ~10k forks, ~1.1k comments (2026-09-25). Early reports: 5,000+ stars within days |
| License | None stated: the gist text has no license clause (checked 2026-09-27) |
Informs: LLM Wiki hub, Explanation, How-to Guides, Reference.
Core Concept¶
Authored by Andrej Karpathy (published 2026-04-04), the "LLM Wiki" is a pattern for building personal knowledge bases using LLMs. It proposes a fundamental shift away from stateless Retrieval-Augmented Generation (RAG) toward a persistent, compounding artifact (a wiki).
Instead of an LLM retrieving from raw documents and rediscovering knowledge from scratch on every query, the LLM incrementally builds and maintains a structured, interlinked collection of markdown files. This wiki sits between the user and the raw sources. The gist is deliberately abstract: an "idea file" meant to be pasted into an LLM agent (Claude Code, OpenAI Codex, OpenCode, Pi), which then builds a concrete implementation with the user.
Key Differences from RAG¶
- Stateless RAG: LLM retrieves chunks, generates answers, and forgets. Knowledge does not build up.
- Stateful Wiki: LLM reads new sources, updates entity pages, revises summaries, flags contradictions, and builds cross-references. Knowledge is "compiled" once and kept current.
3-Layer Architecture¶
- Raw Sources: Immutable curated collection (articles, papers, images). Read-only for the LLM.
- The Wiki: A directory of LLM-generated markdown files. The LLM owns this layer entirely (creates, updates, cross-references).
- The Schema: Instructions (
CLAUDE.md,AGENTS.md) detailing conventions and workflows for the agent.
Operations¶
- Ingest: Read a new source -> discuss key takeaways with the user -> write a summary page -> update index -> update entity and concept pages (a single source might touch 10-15 pages) -> log entry.
- Query: Ask a question -> LLM reads index & relevant pages -> synthesizes answer with citations. Answers can take other forms (tables, Marp slide decks, charts). Valuable answers are saved as new wiki pages.
- Lint: Health checks by the LLM (finding contradictions, orphans, missing links, stale claims, concepts mentioned but lacking their own page, and data gaps to fill with new sources).
Indexing & Logging¶
index.md: Content-oriented catalog of all pages, organized by category and updated on every ingest. The entry point for the LLM during queries. Per the gist this "works surprisingly well at moderate scale (~100 sources, ~hundreds of pages)" without embedding-based RAG infrastructure.log.md: Chronological append-only record of actions (ingests, queries, lint passes). Helps the LLM understand recent history. The gist suggests a parseable prefix such as## [2026-04-02] ingest | Article Title.
Associated Tooling¶
- Obsidian: Acts as the IDE for the knowledge base.
- Obsidian Web Clipper: Quick extraction of sources to markdown.
- qmd: Local hybrid search engine (BM25/vector/LLM re-ranking, all on-device) for markdown files, with both a CLI (so the LLM can shell out to it) and an MCP server (so the LLM can use it as a native tool). Useful as the wiki grows beyond simple index files.
- Git: "The wiki is just a git repo of markdown files", so version history, branching, and collaboration come for free.
- Marp / Dataview: Used for generating presentations or querying metadata within the wiki.
Framing and Lineage¶
- The gist's metaphor: "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase."
- It relates the idea to Vannevar Bush's Memex (1945): a private, curated knowledge store where the associative trails between documents matter as much as the documents.
Why It Works¶
The gist puts it as: "The tedious part of maintaining a knowledge base is not the reading or the thinking — it's the bookkeeping." The maintenance burden (bookkeeping, cross-referencing, consistency checks) is typically what kills human-maintained wikis. LLMs reduce this cost to near zero. The human curates and asks questions. The LLM maintains the structure.