Verdict. This is close to how I run my own notes, with Claude Code working over an Obsidian vault. The data point I keep is that he didn’t need RAG at about 400K words.

what it is

Andrej Karpathy’s workflow for building a personal knowledge base with LLMs. He wrote it up as “LLM Knowledge Bases” on 2026-02-04.

  • Ingest. Source documents (articles, papers, repos, datasets, images) go into a raw/ directory. An LLM incrementally “compiles” them into a wiki, a set of .md files with summaries, backlinks, concept categories and articles linking it all. Web articles come in through the Obsidian Web Clipper, with a hotkey to pull related images locally.
  • Frontend. Obsidian shows the raw data, the compiled wiki and any visualizations. The LLM writes and maintains the wiki. He rarely touches it.
  • Q&A. Once the wiki is big (about 100 articles, about 400K words), he asks an agent complex questions against it. The LLM keeps its own index files and short summaries, and reads the related pages at that scale without a retrieval system.
  • Output. Markdown, Marp slideshows or matplotlib images, viewed back in Obsidian. Outputs get filed back into the wiki so each exploration adds to it.
  • Linting. LLM health checks find inconsistent data, fill gaps with web search, and suggest new articles and connections.
  • Search. He vibe-coded a small search engine over the wiki.

compared to