Skip to content

LLM Wiki Pattern

Source metadata (verified 2026-09-25)

Field Value
Title / file "llm-wiki" gist, file llm-wiki.md, described as "A pattern for building personal knowledge bases using LLMs"
Author Andrej Karpathy (karpathy on GitHub)
Published 2026-04-04, as a follow-up "idea file" to his 2026-04-02 X post "LLM Knowledge Bases"
URL gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
Engagement ~48.6k stars, ~10k forks, ~1.1k comments (2026-09-25). Early reports: 5,000+ stars within days
License None stated: the gist text has no license clause (checked 2026-09-27)

Informs: LLM Wiki hub, Explanation, How-to Guides, Reference.

Core Concept

Authored by Andrej Karpathy (published 2026-04-04), the "LLM Wiki" is a pattern for building personal knowledge bases using LLMs. It proposes a fundamental shift away from stateless Retrieval-Augmented Generation (RAG) toward a persistent, compounding artifact (a wiki).

Instead of an LLM retrieving from raw documents and rediscovering knowledge from scratch on every query, the LLM incrementally builds and maintains a structured, interlinked collection of markdown files. This wiki sits between the user and the raw sources. The gist is deliberately abstract: an "idea file" meant to be pasted into an LLM agent (Claude Code, OpenAI Codex, OpenCode, Pi), which then builds a concrete implementation with the user.

Key Differences from RAG

  • Stateless RAG: LLM retrieves chunks, generates answers, and forgets. Knowledge does not build up.
  • Stateful Wiki: LLM reads new sources, updates entity pages, revises summaries, flags contradictions, and builds cross-references. Knowledge is "compiled" once and kept current.

3-Layer Architecture

  1. Raw Sources: Immutable curated collection (articles, papers, images). Read-only for the LLM.
  2. The Wiki: A directory of LLM-generated markdown files. The LLM owns this layer entirely (creates, updates, cross-references).
  3. The Schema: Instructions (CLAUDE.md, AGENTS.md) detailing conventions and workflows for the agent.

Operations

  • Ingest: Read a new source -> discuss key takeaways with the user -> write a summary page -> update index -> update entity and concept pages (a single source might touch 10-15 pages) -> log entry.
  • Query: Ask a question -> LLM reads index & relevant pages -> synthesizes answer with citations. Answers can take other forms (tables, Marp slide decks, charts). Valuable answers are saved as new wiki pages.
  • Lint: Health checks by the LLM (finding contradictions, orphans, missing links, stale claims, concepts mentioned but lacking their own page, and data gaps to fill with new sources).

Indexing & Logging

  • index.md: Content-oriented catalog of all pages, organized by category and updated on every ingest. The entry point for the LLM during queries. Per the gist this "works surprisingly well at moderate scale (~100 sources, ~hundreds of pages)" without embedding-based RAG infrastructure.
  • log.md: Chronological append-only record of actions (ingests, queries, lint passes). Helps the LLM understand recent history. The gist suggests a parseable prefix such as ## [2026-04-02] ingest | Article Title.

Associated Tooling

  • Obsidian: Acts as the IDE for the knowledge base.
  • Obsidian Web Clipper: Quick extraction of sources to markdown.
  • qmd: Local hybrid search engine (BM25/vector/LLM re-ranking, all on-device) for markdown files, with both a CLI (so the LLM can shell out to it) and an MCP server (so the LLM can use it as a native tool). Useful as the wiki grows beyond simple index files.
  • Git: "The wiki is just a git repo of markdown files", so version history, branching, and collaboration come for free.
  • Marp / Dataview: Used for generating presentations or querying metadata within the wiki.

Framing and Lineage

  • The gist's metaphor: "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase."
  • It relates the idea to Vannevar Bush's Memex (1945): a private, curated knowledge store where the associative trails between documents matter as much as the documents.

Why It Works

The gist puts it as: "The tedious part of maintaining a knowledge base is not the reading or the thinking — it's the bookkeeping." The maintenance burden (bookkeeping, cross-referencing, consistency checks) is typically what kills human-maintained wikis. LLMs reduce this cost to near zero. The human curates and asks questions. The LLM maintains the structure.