Sources#
- An open-source spec for Codex orchestration: Symphony.
- Beyond RAG: Building Agentic Document Workflows with LlamaIndex
- How the Open Knowledge Format can improve data sharing
- Knowledge-Centric Self-Improvement
- LLM Knowledge Bases
- llm-wiki
- Muscle Memory for Agents: Compile not Merely Retrieve
- Sidekick's continual learning loop
- The New Physics of Business — Garry Tan, Y Combinator
- The State of Agent Wikis
Summary#
An architecture pattern originated by Andrej Karpathy where an LLM functions as a compiler: it reads raw source documents and incrementally produces a structured, interlinked markdown wiki. Unlike traditional RAG systems that rely on embeddings and vector databases, this approach uses the wiki's own index files and the LLM's context window for retrieval, which is sufficient at personal knowledge base scale (~100 articles, ~400K words).
Details#
Four-Phase Pipeline#
The system operates as a continuous cycle:
- Ingest — Raw content (web articles via Obsidian Web Clipper, papers, repo notes) lands in a
raw/staging directory as markdown files. - Compile — The LLM reads
raw/and builds index files (summaries of all documents), concept articles (organized by topic with backlinks and cross-references), and derived outputs (slides, charts, filed query answers). The LLM auto-maintains the link graph between concepts. - Query & Enhance — Users browse the wiki in Obsidian, ask research questions via a Q&A agent, or search via a CLI/web tool. Critically, all outputs from queries are filed back into the wiki, so every exploration compounds.
- Lint & Maintain — The LLM audits for inconsistencies, imputes missing information via web search, discovers new inter-concept connections, and suggests further questions. After linting, the cycle returns to compile.
Key Design Decisions#
- No vector database — At personal scale, index files + LLM context window are sufficient for retrieval. This eliminates embedding pipeline complexity. At larger scale, a local search engine like qmd (hybrid BM25/vector search with LLM re-ranking, available as CLI and MCP server) can supplement the index.
- Incremental compilation — New raw documents are integrated into existing wiki structure; already-indexed documents are never reprocessed.
- Explorations always compound — Every query answer, chart, and derived artifact is filed back into the wiki. This is the core differentiator vs. RAG: knowledge is compiled once and kept current, not re-derived on every query.
- LLM does the writing — The human rarely edits the wiki directly; the LLM compiles, links, and maintains it. The human's job is sourcing, exploration, and asking the right questions.
- Wiki as persistent, compounding artifact — Cross-references are already there, contradictions already flagged, synthesis already reflects everything read. The wiki gets richer with every source added and every question asked.
Three-Layer Architecture (from Karpathy's Gist)#
Karpathy's original design document makes the architecture explicit:
- Raw sources — curated, immutable source documents (articles, papers, images, data). The LLM reads but never modifies these. This is the source of truth.
- The wiki — LLM-generated markdown files: summaries, entity pages, concept pages, comparisons, synthesis. The LLM owns this layer entirely — creates, updates, cross-references, maintains consistency. The human reads it.
- The schema — a configuration document (CLAUDE.md / AGENTS.md) that tells the LLM how the wiki is structured, what conventions to follow, and what workflows to execute. Human and LLM co-evolve this over time.
Indexing and Navigation#
Two special files help navigate the wiki at scale:
- index.md — content-oriented catalog of every page with one-line summaries, organized by category. The LLM reads this first when answering queries, then drills into relevant pages. Works well at moderate scale (~100 sources, ~hundreds of pages).
- log.md — chronological, append-only record of operations (ingests, queries, lint passes). Parseable with unix tools if entries use consistent prefixes (e.g.,
## [2026-04-02] ingest | Article Title).
Use Cases#
The pattern applies broadly:
- Personal: goals, health, psychology — filing journal entries, articles, podcast notes
- Research: reading papers over weeks/months, building a comprehensive wiki with an evolving thesis
- Reading a book: chapter-by-chapter companion wiki with characters, themes, plot threads (like a personal fan wiki)
- Business/team: internal wiki fed by Slack threads, meeting transcripts, customer calls, with humans reviewing updates
- Any knowledge accumulation: competitive analysis, due diligence, trip planning, course notes
Why It Works#
The bottleneck of knowledge bases is not reading or thinking — it's bookkeeping. Updating cross-references, keeping summaries current, noting contradictions, maintaining consistency across dozens of pages. Humans abandon wikis because maintenance burden grows faster than value. LLMs don't get bored, don't forget to update a cross-reference, and can touch 15 files in one pass.
Intellectual Lineage#
Karpathy draws a connection to Vannevar Bush's Memex (1945) — a personal, curated knowledge store with associative trails between documents. Bush's vision was closer to this than to what the web became: private, actively curated, with connections between documents as valuable as the documents themselves. The part Bush couldn't solve was who does the maintenance. The LLM handles that.
Variations#
Elvis Saravia describes a variant where ingestion is automated: a tuned Skill agent curates research papers daily, indexes them with the qmd CLI tool, and feeds the indexed knowledge base into an interactive artifact generator built with MCP tools. This produces explorable, interactive visualizations across hundreds of papers.
Future Direction#
Karpathy mentions using the wiki to generate synthetic training data and fine-tune an LLM so it "knows" the data in its weights — turning a personal knowledge base into a personalized model.
The downstream half of that idea now has a production instance. Shopify's Sidekick account (Shopify Engineering, 2026-08-05, case-study, first-party and unreplicated) states this page's compile thesis with the target moved from markdown to parameters: a frozen frontier model "has no mechanism for internalizing what production teaches it. Instead, improvements accumulate in the discrete artifacts around it: prompt edits, retrieval examples, routing rules, and harness code" — and their answer is "a continual learning loop that compresses production experience into the continuous space of the model's weights." That is the same immutable-source → compile → persistent-artifact boundary this page is built on, with a lossier, less inspectable, and completely un-citable output layer. It also supplies the number the future direction never had: how much compiled material a domain fine-tune needs before it pays. Their distillation curve runs 13k trajectories (judge score 61.5) to 61k (73.5), crossing the incumbent production system between 26k and 30k and reaching frontier-model parity only at the top — full curve, axis labels and the parity caveat on Agent Quality Flywheel.
Two differences bound how far it transfers to a wiki. The compiled input is repaired production trajectories, not synthetic text generated from documents, so the generation step Karpathy's version turns on is exactly the step this source skips. And what is compiled is behaviour (how to write a correct GraphQL query) rather than knowledge (what is true about a corpus) — a distinction this vault's own Knowledge-Centric Self-Improvement treatment keeps sharply, and one the weights collapse. A fine-tuned model cannot say which source a claim came from, cannot be linted, and cannot mark a claim superseded, which is the citation-granularity debt of the retrieval counter-case below arriving in its most extreme form.
The agent-wiki landscape: four implementations in three months (July 2026)#
A vendor survey (mem0's In Context #17, "The State of Agent Wikis," 2026-07-21, practitioner-opinion — COI: mem0 sells the user-memory layer its closing section advocates) gives the pattern its field name — agent wikis — and its first landscape: within roughly three months of the April 2026 gist, four teams shipped the same three-layer structure (immutable sources; model-written markdown wiki; schema file, "usually CLAUDE.md or AGENTS.md") against four different corpora. The survey's convergence argument: "Four teams solved four different problems and made the same structure. This agreement is good evidence that the structure is correct."
| System | Corpus | Currency | Written for |
|---|---|---|---|
| DeepWiki (Cognition) | any public GitHub repo; 50K+ largest pre-indexed (URL-swap github.com → deepwiki.com) | re-indexed; grounds Devin | agents, and humans browsing |
| AutoWiki (Factory) | your org's repos | CI refresh, every push | engineers and Droids together |
| OpenWiki / Brains (LangChain) | repos (Code Brain) + Gmail, Notion, git, X, HN (Personal Brain) | re-run to refresh | "LLM context, not human prose" |
| GBrain (Garry Tan) | personal sources | manual or scheduled runs | a person and their agent |
(Table from the survey's matrix figure. Its prose flattens the last column to "the reader of the wiki is a model"; the matrix itself contradicts that for three of the four systems — only OpenWiki is model-only.)
What each implementation adds beyond the gist:
- DeepWiki: the wiki as agent infrastructure, not documentation. "The wiki is not the product. The wiki is retrieval infrastructure for the agent" — Devin uses DeepWiki as the compiled layer below its code search. This is the ingest-time counterpart of Repository Exploration Subagent's query-time explorer: both decouple repo grounding from solving, one by precomputing a per-repo artifact every agent shares, the other by searching per task in a disposable window.
- AutoWiki: maintenance moved into infrastructure. Documentation as a build artifact —
/install-wikiwrites a CI workflow that regenerates the wiki on every push to the default branch. Generation is two-pass (structural scan of README/manifests/CI config/entry points, then semantic scan of routes/endpoints/service classes/schemas/feature flags), fanned out across specialized agents so no single agent documents a whole large repo. See Code as Source of Truth for how this composes with checking truth into the repo. - OpenWiki: the leap from repo to everything. Personal Brain compiles Gmail, Notion, git, X, and Hacker News into one local markdown wiki — "documentation of your work," not of a repository.
- GBrain: the minimal-infrastructure proof. No vector database, no service — markdown in git, a schema file, an auto-maintained link graph. (The survey files GBrain as the personal-scale member; Tan's own talk presents it as a company brain with his ~220K-page personal instance as the exhibit — the two agree on the artifact and differ on the ambition. Full treatment below.)
The axis the survey calls "the maturity tell": Factory treats staleness as a build problem and solves it in CI; the other three — and this vault — are exactly as current as the last time someone ran the command.
The survey's four limits, worth recording because this wiki lives them: (1) scale — the gist's ~100-source ceiling before search must be added (DeepWiki, at 50K+ repos, ships search per wiki); (2) compile-time loss — "an early summary can remove a detail from the source. Each later answer has this error. Retrieval from the raw parts does not have this problem. You exchange the cost of repeated work for the risk of lost data" — the architectural risk this vault's parse-warning discipline manages, and exactly the L2-condensation omission that Layerwise Omission Attribution shows how to count with canary taps; (3) staleness — "an incorrect wiki is worse than no wiki. The incorrect information has the format of correct information"; (4) cost — tokens spent compiling pages nobody reads and linting pages that did not change.
A wiki is not memory (the vendor's boundary)#
The survey's closing distinction — self-interested but real: corpus knowledge ("what does this material contain": scoped to a corpus, accumulated from ingestion) versus user/experience memory ("what did this person decide, prefer, try": scoped to an identity, accumulated from interaction, obligated to handle per-user contradiction, staleness, provenance, and deletion). A wiki does the first and not the second: "Your Gmail wiki tells the agent what is in your Gmail. It does not tell the agent that you changed a decision in a conversation on Tuesday. It does not tell the agent that a method already failed for you." The kicker: "The mistake is not choosing a wiki. It is believing you solved memory because you compiled a corpus."
Attribute the boundary's placement to its author — mem0 sells the memory layer, so the line is drawn exactly where its product begins — but the vault's own stack already embodies the split: this wiki is corpus knowledge, the harness's auto-captured memory directory (Agent Context Files's system-captured memory channel) is the interaction-scoped store, and When Knowledge Layers Disagree: Context Files vs Memory, and Conflicting Sources at Compile Time adjudicates when they disagree.
The pattern gets a spec, and the spec standardizes the container not the curation (Google Cloud, June 2026)#
The landscape above is four bespoke implementations that resemble each other by convergence rather than by agreement. Open Knowledge Format (Sam McVeety & Amir Hormati, Google Cloud Data Cloud, 2026-06-12, vendor-claim) is the first attempt to make them interoperable, and it names its lineage in its own second paragraph: an open specification that "formalizes the LLM-wiki pattern into a portable, interoperable format," linking Karpathy's gist directly and quoting the same bookkeeping sentence this page's Why It Works section is built on. The problem it states is the one the survey above leaves standing — "each instance is bespoke… the knowledge encoded in wikis remains siloed within the original teams."
What OKF v0.1 specifies is a container, and it is this vault's container. A bundle is a directory of markdown files, one file per concept, path-as-identity, each carrying a small YAML frontmatter block (type, title, description, resource, tags, timestamp) over a markdown body; concepts link to each other with ordinary markdown links, turning the directory into "a graph of relationships that is richer than the parent/child links implied by the file system." The two reserved filenames are index.md (progressive disclosure as agents navigate the hierarchy) and log.md (chronological history of changes) — the same two special files Karpathy's gist named and the same two this vault runs. The convergence is on the file layout, not merely on the idea, which is a stronger claim than the survey's four-way structural agreement.
And it requires exactly one field. Principle 1 of three, stated: "OKF requires exactly one thing of every concept: a type field. Everything else… is left to the producer. The spec defines the interoperability surface, not the content model." (The other two: producer/consumer independence — a human-authored bundle consumable by an agent, an export-pipeline bundle browsable in a visualizer; and format-not-platform, "it will never require a proprietary account or SDK to read, write, or serve.")
The gap between that surface and everything this page has established is the finding. Read OKF's frontmatter against this vault's:
| OKF v0.1 field | This vault's analog | What the spec has no field for |
|---|---|---|
type (the only required field) | type: entity / kind: / domain: | — |
title, description, tags | title, summary:, tags | — |
resource | the index row's source URL | — |
timestamp (one) | created: + updated: | when a page was written vs. when it was last checked |
| — | sources: + the evidence: tier | provenance, and trust weighting between sources |
| — | supersession marks, staged contradictions | what to do when two sources disagree |
| — | lint.py, the pruning pass | curation of any kind |
Every hygiene mechanism this page has recorded — Tan's provenance/contradiction/pruning doctrine, Wang et al.'s cite-by-id-and-quote protocol with applies_when scope conditions, this vault's own — sits outside OKF's surface, and each of those three arrived at it independently. A fully conformant bundle can carry no provenance, no contradiction handling, and no pruning discipline whatsoever. The spec would answer that this is deliberate: the content model is the producer's business, and standardizing it would kill adoption. That is coherent, and it has one consequence worth stating plainly — conformance is not a quality signal, and it is the consumer half of producer/consumer independence where that bites. A consumer that can parse any bundle has no way to tell a curated one from a dump. Tan's line survives the format unchanged: a brain nobody curates is "a garbage dump with great search."
The single timestamp: is the sharpest instance, and this vault has paid for the lesson. The survey's third limit above is staleness — "an incorrect wiki is worse than no wiki. The incorrect information has the format of correct information" — and the discriminating fact is not when a document was written but when it was last checked against its sources. One timestamp cannot express that, and the field is the hardest one to keep honest: this vault's 2026-08-14 lint pass found 126 of 338 pages whose updated: had drifted from their actual last edit, which is why lint.py now carries an [fmdate] check reading git rather than frontmatter. A format that ships one timestamp ships the field that rots, without the field that would reveal the rot.
The cross-vendor claim is, so far, one vendor at v0.1. Everything shipped alongside the spec is Google's: a BigQuery enrichment agent (walk a dataset, draft a concept per table, second LLM pass to enrich with citations and join paths), a self-contained static HTML visualizer, three sample bundles built from Google public datasets, and a Google Cloud Knowledge Catalog updated to ingest OKF. The post's own bar is the right one — "the value of a knowledge format comes from how many parties speak it, not from who owns it" — and by that bar nothing is yet established; the reference implementations are explicitly "proofs of concept." Agent Context Files records what clearing the bar looks like: Google's Genkit implementing Anthropic's SKILL.md across four language SDKs is a second vendor consuming another's spec verbatim. Google is on the implementing side there and the authoring side here, and only the first of those is evidence.
The governance leg is out of scope, and the use case is not. The retrieval counter-case below makes governance one of three legs — "stuffing the corpus means stuffing documents the requester should not be allowed to see," and "'model promised to ignore' is not a boundary." OKF's advertised wins are portability properties: shippable as a tarball, hostable in any git repo, mountable on any filesystem. Against its own stated corpus — table schemas, metric definitions, incident runbooks, join paths, deprecation notices — "just files" is simultaneously what makes a bundle portable and what makes it uncontainable. The post claims no access model and does not pretend to; it is worth recording only because the enterprise-internal use case is the one it argues for.
A property of the evidence base, not of the argument: this is the second Google Cloud artifact on the compile side of the compile/retrieve axis in two months. Muscle Memory (Google Cloud FDE, 2026-08-10) argues compile-over-retrieve for personalization memory; OKF (Google Cloud Data Cloud, 2026-06-12) standardizes the compiled artifact's format. Different teams, same employer, same architectural bet — so the corpus's two most recent external endorsements of this page's founding move are not independent in the way source count suggests.
Company Brains: the pattern at organizational scale (Garry Tan's GBrain)#
The most prominent in-the-wild sibling of this architecture (July 2026, practitioner-opinion): Garry Tan's company brain — his MIT-licensed open-source GBrain, "effectively Postgres for agents." His formulation is "the library plus the librarian": the organization's full record (email, meetings, decisions and their reasoning, postmortems) is the library, and the load-bearing component is the librarian — the retrieval layer that decides, per task, which "three books" go into the agent's context (his working-memory image: an agent holds ~1M tokens ≈ three Harry Potter books, against the human 7±2; "the question that determines whether your agents are geniuses or goldfishes is who decides which three books are open on that desk"). His personal instance: ~220,000 pages, compiled mostly by his agents from 20 years of email, meetings, and notes.
Tan pre-empts the "this is just RAG" objection exactly as this page does — "retrieval is the primitive, the same way Postgres is just B-trees… Retrieval is easy. Being worth retrieving from is the product." What's worth keeping from his account is the hygiene doctrine, an independent convergence on this vault's own design:
| Tan's failure mode / prescription | This vault's mechanism |
|---|---|
| "Provenance on every fact" | evidence: tiers + per-source citation (Non-Malleable Memory Authority (TMA-NM) proves content-trust without origin-binding is unsound in the adversarial case) |
| "Contradiction checks when new information collides with the old" | the flag-contradictions-explicitly compile rule |
| "A librarian, human plus agent, whose actual job is pruning" | the lint pass |
| "A brain nobody curates becomes a garbage dump with great search" | why compile/lint exist at all — retrieval over an uncurated store surfaces stale facts "with total confidence" |
His summary — "treat the brain like production infrastructure and it compounds; treat it like a dumping ground and you get a very confident agent that is wrong in ways nobody can trace" — is the operational version of this page's "explorations always compound," with the failure branch made explicit. His economic framing is also worth recording: "model quality is rented, but if you build your brain, you own that brain" (Compounding Data Moat at the level of a knowledge store).
The first external empirical corroboration (Caltech, July 2026)#
Until now this page's thesis rested on Karpathy's design document, one practitioner talk, and this vault's own practice — self-referential evidence at best. Knowledge-Centric Self-Improvement (Wang et al., Caltech, arXiv 2607.19592, empirical) is the first controlled measurement of the underlying bet: that a curated, compiled knowledge artifact is worth more than the system that produced it.
The setup is not a wiki — the agents are benchmark solvers and the store is machine-read — but the architectural boundary is identical. Raw experience is immutable and never edited; a compile step turns it into scoped, evidence-grounded claims; the compiled artifact is the only thing that persists. The measured result: a knowledge bundle frozen at generation 10, separated from the tasks and the model family that produced it, lifts zero-shot solve rates on held-out tasks in all eight donor-recipient cells (Polyglot 8.3%→20.0%, ARC-AGI-1 23.3%→43.3% for the strongest pairing) — and generic agents reading it beat DGM, HyperAgents, GEPA and OpenEvolve on five benchmarks at lower dollar cost. "Explorations always compound" now has a number attached to it, produced by people who had never heard of this vault.
What is worth keeping is a second independent convergence on the same hygiene doctrine — this time from a group optimizing for machine consumption, which makes the agreement harder to explain away as shared taste:
| Their protocol rule | This vault's mechanism |
|---|---|
| Distillation is "a selection step, not a generic summarization step" — keep claims that are actionable and scoped, drop advice that does not name its condition | the compile rule against filler; the granularity rubric |
Every Insight carries applies_when / does_not_apply_when | scope conditions in article prose; the supersede-don't-overwrite rule |
| Claims must cite prior posts by id and quote ≥ 40 verbatim characters of grounding | per-source citation and the evidence: tier |
anti_meta_self_check — a schema field that drops posts whose primitive cannot be defended as non-generic | the thin-article check; "dense, precise, no filler" |
| A retrieval gate rejects a post unless the agent has already read the store for that task | read the index first |
Conflicting claims are preserved as FALSIFIED / UNTRIED with both sides' evidence rather than averaged to consensus | flag contradictions explicitly; When Knowledge Layers Disagree: Context Files vs Memory, and Conflicting Sources at Compile Time |
| Per-task and cross-task bundle inputs kept disjoint so local curation is not polluted by global speculation | concept pages vs derived query outputs |
Their formulation of why the last row matters is the sharpest statement of it the corpus has: "when the evidence is genuinely conflicting, the protocol's job is to keep the conflict legible to future agents, not to average it away."
One finding cuts against a natural instinct here. Their knowledge-transfer adapter bounds every field at 0-3 items and is instructed to return short or empty lists when the prior is only weakly relevant — added after they observed that transferring a fixed quantity made recipient memory "overly noisy or detrimental." More compiled knowledge is not monotonically better; what is delivered has to be selected against the task at hand.
The position stated from the memory side, and what it does not measure (Google Cloud, August 2026)#
Muscle Memory for Agents: Compile not Merely Retrieve (Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang & Tanya Dixit, Google Cloud FDE, arXiv 2608.08995, 2026-08-10, empirical) is the first external source whose own stated thesis is this page's founding move. Knowledge-Centric Self-Improvement corroborated the underlying bet without framing itself that way; this paper names the opponent in the first sentence of its abstract:
"Memory for LLM agents has converged on a single architectural pattern: store experience as text, embeddings, reflections, or rules; retrieve at inference time; let a general-purpose orchestrator interpret what to do. This paper argues that the pattern is the wrong default for personalization."
The corpus is different — recurring user intent mined from 250 synthetic multi-turn conversations, compiled into executable Python "mini-agents" that a router dispatches to, rather than documents compiled into a wiki a human reads. The architecture is the same: an immutable raw layer, a compile step emitting a tested artifact, and a runtime whose job reduces to selecting the artifact rather than interpreting fragments. Its four-phase pipeline (Harvest → Analyze → Augment → Evaluate) maps onto this vault's ingest → compile → query → lint cycle stage for stage. Its three principles are this page's design decisions restated as a position: P1 compilation over retrieval ("retrieval is appropriate for one-off facts whose interpretation depends on the live request; compilation is appropriate for recurring procedures whose interpretation is itself the part that should be cached"), P2 specialists over generalists, P3 tested before deployment.
It crosses the boundary mem0 drew, which is the most consequential thing in it for this page. The wiki-is-not-memory section above splits corpus knowledge from user/experience memory and assigns compilation to the first. This paper takes the compile thesis into the second half and argues it wins there too — so the compile/retrieve axis is orthogonal to the corpus/user axis rather than aligned with it. Nothing in mem0's boundary is struck: a wiki still does not know what you decided on Tuesday. What is now contested is the unstated corollary, that user memory therefore has to be retrieved.
P3 is the principle this vault does not have. The paper's compile step is quality-gated before anything deploys: two candidates generated at temperatures 0.2 and 0.35 with a lightweight selector, a post-generation fact-check that revises low-confidence assertions with uncertainty caveats, a critic pass scoring each agent against real conversation history on 10 criteria at a 7/10 threshold (hard-ceilinged at 4/10 if ≥2 contradictions are found), a non-parametric ranking over five weighted dimensions (value ×3, distinctiveness ×2, trigger clarity ×2, quality ×2, frequency ×1) retaining only agents ≥25/50, an overlap merge (embedding cosine plus domain/task-type Jaccard ≥0.73), and a mini-eval running each survivor on 6 historical scenarios that prunes it if average accuracy <2.5 or if it never triggers. This vault's analog, lint.py, runs after a page lands, is report-only, and checks structure rather than content. The paper's argument for why that difference matters is the sharpest statement of the compile thesis's audit advantage in the corpus: "Retrieved-then-interpreted memory cannot be tested in this way; its behavior emerges only at inference, conditioned on the orchestrator and the rest of the prompt, and is therefore difficult to audit, regression-test, or version." Read against the retrieval counter-case below, that is a different answer to Doulcet's auditability leg — compiled memory cannot cite a sentence back to a source region, but it can be regression-tested before it runs.
The results, at their real size. 90 held-out scenarios, 18 per user (10 similar to training intents, 8 different), five personas, one LLM judge (Gemini 3.1 Pro, temperature 1.0) scoring accuracy, helpfulness and personalization on 1–4 scales and declaring a winner:
| User | Persona | Agents | Fired | W-L-T | Acc. Δ | Pers. Δ |
|---|---|---|---|---|---|---|
| user_1 | software engineer | 1 | 5/10 | 3-2-0 | −1.00 | +1.20 |
| user_2 | marketing manager | 6 | 10/10 | 10-0-0 | 0.00 | +2.50 |
| user_3 | ML grad student | 6 | 9/10 | 8-1-0 | −0.56 | +2.23 |
| user_4 | cafe owner | 5 | 6/10 | 6-0-0 | 0.00 | +2.17 |
| user_5 | travel + cooking | 5 | 6/10 | 5-1-0 | 0.00 | +1.67 |
| All | 23 | 36/50 | 32-4-0 | −0.28 | +2.05 |
(Agent counts from Table 1, which docling collapsed into one grid row and which was reconstructed with pdftotext -layout at ingest; the rest from Table 2, parsed clean and reconciled digit-for-digit against the prose and the abstract. Table 2 excludes 8 false-positive firings by its own caption.)
The headline is 32 of 36 wins where an agent fires, an 88.9% win rate, personalization +2.05 (1.67 → 3.72) against an accuracy cost of −0.28 (3.92 → 3.64). Three things the aggregate hides and the per-user rows do not. The accuracy cost is not uniformly zero — user_1 pays a full point on a four-point scale while three of five users pay nothing, and −0.28 is the mean of a bimodal column, not a small tax everyone pays. Firing is the exception, not the rule: trigger accuracy is 72% (36 of 50 similar-domain scenarios), so the 88.9% is conditioned on the 72%, and the false-positive rate on different-domain scenarios is 20% (8 of 40) — the augmented assistant still wins 5 of those 8, which the paper reports as "rarely damaging" rather than harmless. And n is small in every direction: 5 personas, 36 scored firings, one judge, all evaluation synthetic.
The gate pruned the wrong thing, for a reason that transfers. user_1 was cut to a single agent because the non-parametric ranking "penalized agents whose scope overlapped with general-purpose LLM capabilities ('explain Python errors' is valuable to the user but indistinct from the baseline)." A distinctiveness gate scores a compiled artifact against what the base model already does — correct for cost, wrong for coverage — and it is the same user who then pays the −1.00. The paper offers two explanations for that deficit in different sections (over-pruning to one agent in §6; the deficit "correlates with domain technicality rather than architectural isolation" in the trade-off paragraph) and never reconciles them; at five users it cannot.
The measured instance of compile-time loss, and it is this page's own architectural risk with a number attached. The survey's second limit above — "an early summary can remove a detail from the source; each later answer has this error" — is exactly what the paper's error analysis found, in a paper arguing for compilation: "mini-agents produce stylistically personalized responses but hallucinate technical details." The traced case is a statistics agent whose baked-in prompt wrote mannwhitneyu(..., continuity=True) instead of the correct use_continuity, crashing at runtime — baseline 4/4, augmented 1/4. The un-compiled arm, which read nothing, was right; the compiled artifact carried the error into every invocation. The authors' own diagnosis is the transferable part: the quality gates "target general factual grounding … but do not validate domain-specific API correctness," and closing that gap "likely requires tool-augmented verification (e.g., executing generated code snippets in a sandbox) rather than LLM-only critic passes." An LLM critic over a compiled artifact catches fabrication and misses specifics — see Verification as the New Bottleneck.
The scale claim is argued, not measured, and it is this page's scale question it bears on. The paper's cost case is that compilation pays "a lightweight, fixed-size routing cost (one feature-extraction call plus one embedding lookup) that is independent of how many patterns have been compiled," against retrieval's per-call prompt inflation of "typically 500–2,000 tokens of input to the main LLM call on every invocation" that grows with the breadth of the store. Its Discussion states the conclusion directly: routing overhead "is architecturally bounded: it does not grow with the number of compiled agents or the depth of stored history." No latency or token measurement exists anywhere in the paper. It defers precisely that comparison, verbatim: "A full empirical latency and token-cost comparison between compilation routing and retrieval-based alternatives is a valuable direction for future work; we note that such a comparison must account for retrieval's own per-call costs (embedding, search, context injection), which are routinely omitted from retrieval-system analyses." Its own swarms cap at 1–6 agents per user (23 total across 5 users) and are never stress-tested larger, and the bounded-routing claim does not appear among the paper's four stated Limitations — it sits in the Discussion as an argument.
One detail in that routing stage is awkward for this page's no vector database decision above and worth stating plainly: the paper's compiled system does run an embedding index — a 768-dimensional scope embedding per agent (text-embedding-005), cosine-scored against the embedded user message with soft attribute penalties (−0.15 domain, −0.10 task type) and a ≥0.45 threshold, with a binary-questionnaire stage behind it for candidates embeddings cannot separate. So compilation did not eliminate the index; it relocated it, from embedding the corpus and retrieving fragments to embedding the artifacts and routing. That is a design answer to the scale question — the index shrinks from one entry per fragment to one per compiled artifact — and it is not a measurement of where either curve breaks.
Its own limits, and the one it does not list. All four stated: evaluation is entirely synthetic (simulated user agents and an LLM judge, no human participants); a single model family (Gemini) does generation, matching and judging, so the judge shares a family with everything it scores; triggers are static after generation, with no adaptation to runtime feedback or evolving preferences; and there is no cross-user pattern transfer. Evaluation parity is otherwise carefully drawn — neither arm gets web search or external memory, and the only difference is the augmented agent's swarm tool. The limit it does not list is the one above: an unmeasured cost claim presented as an architectural property.
And one internal gap worth recording, because it is the shape this vault's own compile passes can fall into. §5.1 lists three observations "our experiments are designed to validate," the third being that "the pipeline components (behavioral separation, embeddings, hallucination guard, questionnaires) are complementary; no single component subsumes the others." The only evidence offered is a Design lessons paragraph the paper itself fences: "qualitative lessons from iterative development, not formal ablations." A validation target was announced and never run, and nothing in the paper flags the mismatch.
Spec-as-Compilation Source (Symphony's Cross-Language Fuzz)#
The most concrete extension of LLM-as-compiler in the wild so far: OpenAI's Symphony team treated their SPEC.md as the source and asked Codex to implement it in Elixir, TypeScript, Go, Rust, Java, and Python. They then used divergences across the implementations to identify ambiguities in the spec and simplify it.
What this technique does that's genuinely new:
- The LLM is the compiler (markdown → working orchestrator in N target languages).
- Multiple implementations are a spec-fuzzing signal — anywhere implementations diverge, the spec is under-constrained. This is analogous to differential fuzzing in compiler verification, but with English/markdown as the source language.
- The spec is the durable artifact, not the compiled output. OpenAI explicitly said they don't plan to maintain Symphony as a standalone product — it's a reference implementation that users point their own coding agent at.
Implications for this vault:
_system/compiler-prompt.mdis structurally analogous to Symphony'sSPEC.md— both define how an agent should turn one kind of artifact (raw docs / Linear tickets) into another (wiki articles / running orchestrators).- Spec-fuzzing-via-multi-language is overkill for a knowledge base, but the idea generalizes: if
compiler-prompt.mdproduces meaningfully different wikis when run by different model families (Claude vs. GPT vs. local), the divergences point to under-specification. - The schema layer (Karpathy's term) is the same artifact category as
SPEC.md/WORKFLOW.md— repo-versioned markdown that defines agent behavior. See cross-link to Claude Code Best Practices (CLAUDE.md), Hermes Agent (AGENTS.md/SOUL.md), Symphony (WORKFLOW.md).
The retrieval counter-case, and where it actually bites (Doulcet, May 2026)#
This page's founding move is "replace RAG with compilation." The strongest opposing view now in the corpus is Document Parsing as the Retrieval Bottleneck — a 116-slide LlamaIndex workshop (practitioner-opinion, direct vendor COI) whose thesis is that retrieval not only survived the long-context era but won it. Its argument is worth stating precisely, because only one of its three legs is about token budgets:
- Cost — 1M tokens per query at frontier rates; caching helps, the arithmetic still fails for any nontrivial corpus. (Dissolves if inference gets cheap enough.)
- Governance — stuffing the corpus means stuffing documents the requester should not be allowed to see. "'Model promised to ignore' is not a boundary."
- Auditability — "'Why did the AI say this?' Retrieval gives a citation log. Long context gives a vibe."
Legs 2 and 3 are not about capacity at all, and leg 3 is the one that lands here. A compiled wiki answers from the wiki, not from the page. This vault's articles cite [[raw/...]] documents, but an individual sentence in an article does not resolve to a page and a region in the source — which is exactly the property the deck spends its parsing section building (bbox grounding, per-field page citations, "cite back to pixels"). The compile step converts a citable corpus into a readable one and pays for it in traceability.
Where the two architectures agree, and it is the load-bearing agreement. The survey's compile-time-loss limit above — "an early summary can remove a detail from the source; each later answer has this error" — is the deck's structure loss moved one layer later. The deck's version of the sentence is blunter: "none of it recovers a document that was parsed badly in the first place." Both architectures are lossy compressions of an immutable source, both put the loss at a step nobody re-examines, and both are defended by keeping the raw layer immutable so the step can be redone.
But only one of them actually redoes it. The deck's rule is "parsing is a stage, not a step — you will reparse, you will re-extract with new schemas, you will rerun retroactively when the parser improves; build for that, store every intermediate, make parsing idempotent." A retrieval pipeline pays the ingest tax per query and can re-pay it better later; a compiled wiki pays it once and keeps the result. This vault has the precondition (immutable raw/, recorded parse warnings, docling: blocks naming the parser configuration) and not the practice: no compiled article is regenerated when a source's parse is later found to be damaged, and _system/backfill/known-bad.md exists precisely because that backlog has nowhere to go. The reparse-and-recompile path is the concrete thing this page owes the counter-case.
The synthesis is supplied, unintentionally, by the deck's own extract vs parse slide: use schema-first extraction when you know what you want, the fields recur across documents, and downstream is structured — use parse + retrieval when you don't know what you want and the question shapes have a long tail. That is the compile/retrieve boundary in the deck's own vocabulary. A personal wiki over ~100 curated sources read by one person is the first case; an arbitrary-Q&A surface over an enterprise corpus with per-tenant permissions is the second. "Most real systems use both" is the deck's answer, and it does not contradict this page so much as bound it.
Evidence weighting: the counter-case is a vendor talk without measurement, against this page's design document plus one controlled study (Knowledge-Centric Self-Improvement) plus this vault's own practice. It is not authority to demote the compile thesis. What it does supply is a well-specified requirement the compile thesis has not met — citation granularity and reparse currency — and those are checkable regardless of who raised them.
Connections#
- Document Parsing as the Retrieval Bottleneck — the architectural rival, treated above. Retrieval as audit trail rather than capacity workaround; the shared irreducible risk (structure lost at ingest is unrecoverable downstream, whether the consumer is a retriever or a compiler); and the two requirements it raises that this architecture does not currently meet — sentence-level citation granularity, and recompiling when a source's parse is found damaged
- Code as Source of Truth — checking specs/skills into the repo is the compiler-wiki pattern applied to code
- This concept is the foundational architecture of this Obsidian vault (see
_system/compiler-prompt.md) - Agent Harness Engineering — shares the pattern of repository-local knowledge as system of record; OpenAI's AGENTS.md-as-table-of-contents mirrors this wiki's schema layer
- Claude Code Best Practices — CLAUDE.md files serve as the schema layer in Claude Code's implementation of this pattern
- LLM-Driven Vulnerability Research — the vulnerability research scaffold uses SHA-3 cryptographic commitments as a form of verifiable knowledge compilation; Claude Code's agentic capabilities power the discovery pipeline
- Client-Side Agent Optimization — the wiki's compile / query / lint phases are themselves an agent pipeline; different phases could be assigned to different models (cheap model for index drift checks, strong model for cross-reference synthesis) and the combo optimized
- Symphony — the most concrete extension of LLM-as-compiler beyond knowledge bases: OpenAI compiled
SPEC.mdinto 6 language implementations and used the divergences as a spec-fuzzer to remove ambiguity - Ticket-Driven Agent Orchestration — Symphony's
WORKFLOW.mdis structurally the same artifact category as the schema layer here; both are repo-versioned markdown that the LLM "compiles" into action - Agent Context Files — the spec-as-document pattern is LLM-as-compiler applied to a context file; Symphony's compile-SPEC.md-into-6-languages spec-fuzzing is the clearest instance
- Design Concept Grilling — Brooks's "design concept" (shared understanding before any artifact) is the alignment-layer analog: a wiki captures what is true, a grilling session captures what we agree on, both treat the LLM as a partner in compilation rather than a generator of one-shot output
- Andrej Karpathy — originated this pattern (the llm-wiki gist) and, in his May 2026 interview, re-endorses it as his daily practice — building a wiki from articles he reads and querying it
- Software 3.0 — Karpathy's canonical example of a "new information-processing task that wasn't a program before": recompiling documents into a wiki is impossible in Software 1.0/2.0
- Outsource Your Thinking, Not Your Understanding — why this pattern works for Karpathy: "anytime I see a different projection onto information, I gain insight" — the wiki is a tool for building understanding, not just retrieval
- Memory and Context Poisoning — the adversarial threat surface this pattern inherits: any system that lets an agent write to durable memory needs the integrity-validation and source-attribution controls that keep a compiled store trustworthy
- LLM-Assisted Grey-Literature Theory Building — the same architectural boundary run for research synthesis: an LLM compiles thousands of raw documents into a structured, quote-grounded artifact (a causal theory), but the interpretive step stays with the human — automating the codes→theory synthesis produced 15,029 shallow, redundant statements, the negative-result echo of "the LLM does the bookkeeping, not the judgment"
- Garry Tan — GBrain as the organizational-scale sibling (library + librarian, memory-plus-hygiene); an independent convergence on this vault's provenance/contradiction/pruning design
- AI-Native Organization — the company brain is the memory layer of Tan's AI-native org; the org's skills route work, the brain supplies what the org already knows
- Non-Malleable Memory Authority (TMA-NM) — the adversarial-case proof behind "provenance on every fact": authority derived from memory content or lineage is launderable; origin must be bound at write time
- Compounding Data Moat — "model quality is rented, but you own your brain": the curated knowledge store as the durable asset over model access
- Latent vs. Deterministic Space — this vault as the worked example: deterministic generators/linters (
build.py/lint.py) around a latent compiler - Layerwise Omission Attribution — this vault located inside someone else's taxonomy. "OCR table-structure loss" is the first mechanism listed under its L0 ingestion layer, and it is exactly the docling table-collapse/shift failure documented in
_system/compiler-prompt.md; the ingest-timetable-collapseandtable-shiftchecks are L0 checkpoint taps in everything but name. The technique the vault does not have is the canary — a known token planted in the source and exact-matched after parsing, which converts "are these tables suspect?" from a heuristic net into a countable per-document loss rate. Its L2 layer (string slicing, metadata stripping, condensation) is the compile step itself - Knowledge-Centric Self-Improvement — the first external, controlled measurement of this page's central bet: a curated knowledge artifact outperforms the systems that produced it, and keeps working after the tasks and the model family are gone
- Owning Your Externalized Cognition — the ownership argument for running this architecture personally rather than renting it, plus the hygiene doctrine that matches this vault's: "a brain nobody curates is a garbage dump with great search," so the primitive is memory plus provenance, contradiction checks, and pruning. Its "retrieval is easy, being worth retrieving from is the product" is the compile-vs-pile boundary in one line
- Repository Exploration Subagent — the query-time counterpart to DeepWiki's compile-at-ingest: FastContext grounds a solver by searching per task in a disposable explorer window, DeepWiki by precomputing a per-repo wiki all agents share; same decoupling, opposite ends of the ingest/query cost trade
- Agent Quality Flywheel — the compile target moved from markdown to model parameters, and the closest thing the corpus has to this page's own Future Direction running in production: production experience "compressed into the continuous space of the model's weights," with a dataset-size curve for how much compiled material a domain fine-tune needs. The output layer is the extreme case of the citation-granularity debt below — weights cannot cite, cannot be linted, and cannot mark a claim superseded
- Crystallizing Agent Work into Workflows — the same compile move on behaviour rather than knowledge, and the source of this page's missing gate. Muscle Memory's specialists are quality-gated once against history and then frozen; Malik's playbooks earn promotion from a runtime track record and are automatically demoted on regression, which is the control both compiled-memory sources lack and which this vault's report-only lint pass does not supply either
- AI-Assisted Error Analysis — the compile-once premise arriving from the eval side. Shankar's third objection to "here are my traces, go evaluate" is that the run persists nothing: each pass rereads the data and re-derives the failure modes from scratch, with no shared intermediate across runs. Her fix — persist the failure-mode taxonomy in memory or in code — is this page's architecture applied to evals, and her taxonomy-authorship boundary is the same one the grey-literature pipeline draws
Open Questions#
- At what scale does the no-vector-database approach break down? Karpathy's ~100 articles fit in context, but what about 1,000+? Partially answered (2026-08-13), on the design axis only: Muscle Memory for Agents: Compile not Merely Retrieve argues the compile side's per-call cost is fixed — one feature-extraction call plus one embedding lookup, "independent of how many patterns have been compiled" — against retrieval's 500–2,000 tokens of prompt inflation per call growing with the breadth of the store, and it relocates rather than removes the index (one 768-dim scope embedding per compiled artifact instead of one per fragment). It measures none of this. Its swarms cap at 1–6 artifacts per user (23 total), are never stress-tested larger, and it defers the exact comparison to future work in its own words; the bounded-routing claim sits in its Discussion and not in its Limitations. So the architectural argument now has an external statement and the scale number still has no source.
- What's the optimal granularity for concept articles — one concept per article, or clustered by theme? Partially answered (2026-08-03): Knowledge-Centric Self-Improvement answers a machine-read version of this and its answer is neither — granularity is carried by scope conditions attached to each claim (
applies_when/does_not_apply_when) rather than by article size, with two levels of store (per-task and cross-task) whose inputs are kept disjoint. It also supplies a measured caution: transferring a fixed quantity of knowledge made recipient memory "noisy or detrimental," so their adapter bounds delivery at 0-3 items per field and returns empty lists when the prior is weakly relevant. Suggestive, not settling — bundles consumed by a solver agent are not articles read by a human. - How effective is the synthetic training data → fine-tuning pipeline in practice? Partially answered (2026-08-13), on the downstream half only: Sidekick's continual learning loop runs the compile-into-weights step in production and reports a scaling curve for it — judge score 61.5 at 13k compiled trajectories rising to 73.5 at 61k, past the incumbent production system between 26k and 30k and level with a frontier reference at the top — so "does a domain fine-tune over your own accumulated material beat a general model on that domain" has an affirmative
case-studyanswer with a dataset-size number attached. Three things keep it partial. The generation step is absent: their training data is repaired production trajectories, not synthetic text derived from documents, which is the half of Karpathy's proposal that is actually uncertain. The object compiled is behaviour, not knowledge, so nothing establishes that document-derived facts fine-tune as well as action trajectories do. And it is one unreplicated first-party account whose top point clears its frontier reference by 0.4 judge points with no intervals. - Does conformance to a knowledge format predict anything about the knowledge? OKF standardizes the container and explicitly assigns provenance, contradiction handling, and pruning to the producer, so a conformant bundle can be an uncurated dump — and its consumer half has no way to tell. The trigger event that settles it: the first bundle from a producer outside Google. Check whether it carries per-claim source attribution and any staleness signal beyond the single
timestamp:field. If it does not, "portable format" and "worth retrieving from" are fully orthogonal, and the spec's minimalism is why. - Would a pre-deployment quality gate on a compiled page catch errors a post-hoc structural lint cannot? Two sources now gate compiled artifacts before they run (a critic pass plus a held-out mini-eval; auto-generated acceptance tests from traces), and one of them measured the failure the gate misses — a fabricated API argument baked into a specialist's prompt, scoring 1/4 against the un-compiled baseline's 4/4. This vault gates nothing at compile time. The falsifiable version: run a content gate over newly compiled pages (does every quoted figure reconcile against its cited source region) and count how many defects it catches that
lint.pystructurally cannot see.
Resolved Questions#
- How to handle conflicting information across sources during compilation? Answered: When Knowledge Layers Disagree: Context Files vs Memory, and Conflicting Sources at Compile Time — a five-step protocol extracted from this vault's own practice and worked cases: (1) align constructs before declaring conflict (most contradictions dissolve into metric/population/time-axis/unit non-comparability — the Faros-vs-CMU worked example); (2) attach provenance and evidence tier, weigh by method and incentive, never average; (3) stage genuine conflicts explicitly on every affected page, bidirectionally — silent choice is the compile-time form of laundering; (4) convert staged conflicts into tracked open questions with named resolution conditions; (5) resolve at compile/lint time (Tan's librarian), so queries inherit the staged conflict with weights visible rather than re-adjudicating per query.
Derived#
- What Are AI Tools? — query exploring what AI tools are covered in this wiki
- When Knowledge Layers Disagree: Context Files vs Memory, and Conflicting Sources at Compile Time — the compile-time contradiction protocol, grounded in this page's hygiene doctrine and the vault's staged-conflict history
- Authority and Audit Survive Abundance — why the retrieval counter-case's governance and audit legs survive free context, and why they bind this architecture too: compilation and retrieval are both selection-with-a-record, corpus-stuffing is the only option the audit leg eliminates, and the citation-granularity debt this page already concedes is that requirement arriving here
Sources#
- LLM Knowledge Bases
- llm-wiki
- The New Physics of Business — Garry Tan, Y Combinator — Garry Tan, AI Engineer talk (2026-07-17,
practitioner-opinion): the company-brain section (GBrain, library + librarian, memory plus hygiene) - Knowledge-Centric Self-Improvement — Wang, Yoon, Qu, Wang, Sehgal, Mazumdar & Yue (Caltech, arXiv 2607.19592, 2026-07-21,
empirical): the held-out transfer result (§4.4, Table 4) and the curation-protocol rules that converge on this page's hygiene doctrine (§3, Appendix E). Full treatment and parse notes on Knowledge-Centric Self-Improvement - Beyond RAG: Building Agentic Document Workflows with LlamaIndex — Pierre-Loic Doulcet, AI Engineer Singapore 2026 (
practitioner-opinion, LlamaIndex vendor COI): the three-legged case for retrieval surviving long context, and the "parsing is a stage, not a step" rule this page's compile phase does not follow. Full treatment and parse warnings on Document Parsing as the Retrieval Bottleneck — the raw file deliberately contains corrupted parser output as demo material and no number may be taken from it - The State of Agent Wikis — mem0, In Context #17 (2026-07-21,
practitioner-opinion; memory-vendor COI on the wiki≠memory boundary): the agent-wiki landscape, the maintenance-currency axis, the four limits. Technique-matrix and wiki-vs-memory figures viewed; the matrix carries per-system corpus/currency/audience data absent from the prose - Sidekick's continual learning loop — Andrew McNamara & Cody Mazza-Anthony, Shopify Engineering, 2026-08-05,
case-study(first-party, unreplicated, no controlled arm). Cited here for the compile-into-weights framing (the frozen-model / discrete-artifacts passage and "compresses production experience into the continuous space of the model's weights") and for the distillation curve read off the article's GraphQL Distillation figure, whose values appear nowhere in the page text. Full source treatment on Agent Quality Flywheel - How the Open Knowledge Format can improve data sharing — Sam McVeety (Tech Lead, Data Analytics) & Amir Hormati (Tech Lead, BigQuery), Google Cloud blog, published 2026-06-12, clipped 2026-08-14 (
vendor-claim, ~1,900 words). Cited here for the v0.1 design (bundle-as-directory, path-as-identity, the six frontmatter fields, theindex.md/log.mdreserved filenames, links-as-graph), the three principles quoted verbatim, the shipped reference implementations, and the explicit citation of Karpathy's gist as the pattern being formalized. First-party spec announcement with no measurement of any kind — no adoption count, no producer or consumer outside Google, no comparison against any other knowledge format; every claim about what the format buys is a design argument. Spec and sample bundles atGoogleCloudPlatform/knowledge-catalog/tree/main/okf. Parse note: the clipper split the post's two code listings (the directory tree and the sample frontmatter) into one fenced block per line — ~40 consecutive single-line fences. Nothing was lost and the reading order is intact; every structural fact quoted here was read off those fragments and cross-checked against the prose that introduces them. Date caveat:published: 2026-06-12comes from the clipper's frontmatter and nothing in the body corroborates it; mem0's 2026-07-21 agent-wiki survey landscapes four implementations and does not mention OKF, which is mildly odd for a spec published five weeks earlier. Recorded as unresolved, and nothing on any page turns on the date - Muscle Memory for Agents: Compile not Merely Retrieve — Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang & Tanya Dixit (Google Cloud FDE), Muscle Memory for Agents: Compile not Merely Retrieve, arXiv 2608.08995, 2026-08-10,
empirical(12pp, 2 tables, 2 figures; reference implementation atGoogleCloudPlatform/generative-ai/tree/main/agents/personalized-agent-swarms). §3 the three principles and the compiled-vs-retrieved divergence; §4.2 the quality gates; §4.3 the two-stage router and the routing-overhead argument; §5.2 + Table 2 the head-to-head results; §6 the limitations, the accuracy-personalization trade-off and themannwhitneyuerror analysis. First-party-stack COI rather than a product claim: Google Cloud authors, Gemini for generation, matching and judging (the paper's own Limitation 2), and the implementation ships in Google's public repo. Parse notes:verifyreturnedwarn— Table 1 was cell-collapsed (6 logical rows merged into 1) and reconstructed withpdftotext -layoutat ingest, recorded in the 2026-08-13ingestentry in Log; the agent-count column quoted on this page comes from that reconstruction and matches the prose's "23 agents across 5 users". Table 2 parsed clean and every cell was reconciled digit-for-digit against the prose and the abstract.canary-recallreportsokbut did not run — 0 eligible tokens, because the headline statistics (88.9%, +2.05, −0.28) repeat across abstract, intro, results and conclusion, so nothing clears the sampler's exactly-once bar; the manual prose/table reconciliation stands in for it. Both figures' captions are transcribed in full in the raw and no figure magnitude is quoted here, so the image two-pass was not load-bearing (Figure 1 was opened at ingest and matches its caption). Small synthetic eval: 5 personas, 36 scored firings, one LLM judge in the same family as the system under test, and the routing-cost claim is unmeasured
Cited by 36
- Code as Source of Truth×5
Fung states check-it-into-the-repo as a workflow prescription. Google Cloud's Open Knowledge Format…
- Open Questions Backlog×5
Llm As Compiler Knowledge Base (28d) — Would a pre-deployment quality gate on a compiled page catch…
- Agent Context Files×4
The pattern generalizes upward. The same "plaintext spec as load-bearing artifact" instinct shows…
- Andrej Karpathy×4
He closes the interview by tying education back to Llm As Compiler Knowledge Base: "anytime I see a…
- When Knowledge Layers Disagree: Context Files vs Memory, and Conflicting Sources at Compile Time×4
Both answers rest on one principle the corpus establishes independently at the security layer and…
- Where Does the Why Live?×4
Cheap building dissolved specification and relocated alignment into the artifact — but it orphaned…
- Crystallizing Agent Work into Workflows×3
Llm As Compiler Knowledge Base — the same compile move on knowledge rather than behaviour, and the…
- Garry Tan×3
His open-source project (MIT-licensed, built in the open): a company brain — "the library plus the…
- Knowledge-Centric Self-Improvement×3
That last sentence is the compile-time contradiction rule of Llm As Compiler Knowledge Base arrived…
- Owning Your Externalized Cognition×3
"Isn't this just RAG?" — "Sure, and Postgres is just B-trees." Retrieval is the primitive, not the…
- AI-Assisted Error Analysis×2
Llm As Compiler Knowledge Base, arriving from the eval side: the expensive artifact is
- Latent vs. Deterministic Space×2
Llm As Compiler Knowledge Base — this vault as an instance: deterministic generators and linters…
- LLM-Assisted Grey-Literature Theory Building×2
Why it failed is the general lesson: practitioners use inconsistent terms, and the same term for…
- Memory and Context Poisoning×2
This is the adversarial counterpart to the benign persistent-memory designs elsewhere in the wiki —…
- Outsource Your Thinking, Not Your Understanding×2
Karpathy ties the thesis directly to the LLM-wiki pattern he originated: building a personal wiki…
- Software 3.0×2
A subtler point: previous code operated over structured data. Software 3.0 enables operations that…
- Symphony×2
External release — Extracted to a standalone SPEC.md. OpenAI asked Codex to implement the spec in…
- Agent Harness Engineering
Llm As Compiler Knowledge Base — shares the pattern of repository-local knowledge as system of…
- Agent Quality Flywheel
Llm As Compiler Knowledge Base — the same compile boundary with parameters as the output layer…
- AI-Native Organization
Llm As Compiler Knowledge Base — the company brain (library + librarian) is the AI-native org's…
- Authority and Audit Survive Abundance
This leg is not retrieval-partisan, and the corpus's honest entry here is that it binds the…
- Azalia Mirhoseini
AI as a compiler is the direction she names as personally most exciting: generate low-level…
- Claude Code Best Practices
Llm As Compiler Knowledge Base — CLAUDE.md files serve as the schema layer in this vault's…
- Client-Side Agent Optimization
Llm As Compiler Knowledge Base — the wiki's own compile / query / lint phases could be modeled as…
- Compounding Data Moat
Llm As Compiler Knowledge Base — the moat as a knowledge store: Tan's "model quality is rented, but…
- Document Parsing as the Retrieval Bottleneck
Llm As Compiler Knowledge Base — the strongest opposing view in the corpus, and the disagreement is…
- The HTML Artifact Lifecycle: Where Plan History Lives, and When Disposable Becomes Durable
The compiled-store precedent: this vault's own architecture regenerates derived views (index…
- Layerwise Omission Attribution
Llm As Compiler Knowledge Base — this vault sitting inside the taxonomy. "OCR table-structure loss"…
- LlamaIndex
Llm As Compiler Knowledge Base — the architectural rival: this vault's compile-once wiki against…
- LLM-Driven Vulnerability Research
Llm As Compiler Knowledge Base — the responsible disclosure process uses SHA-3 cryptographic…
- Agent Systems & Harness Engineering
Llm As Compiler Knowledge Base — Karpathy's architecture: LLM incrementally compiles raw docs into…
- Non-Malleable Memory Authority (TMA-NM)
Llm As Compiler Knowledge Base — the benign-curation face of the same requirement: Tan's…
- Repository Exploration Subagent
Llm As Compiler Knowledge Base — the compile-at-ingest counterpart: Cognition's DeepWiki…
- Ticket-Driven Agent Orchestration
Llm As Compiler Knowledge Base — the wiki's /compile and /lint are themselves ticket-like work…
- What Are AI Tools?
Llm As Compiler Knowledge Base — architecture pattern for LLM-powered knowledge bases
- What Makes a Self-Improvement Artifact Transfer?
Llm As Compiler Knowledge Base — this vault — is the same bet made for a human-and-LLM reader:…
Related articles
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Agent Context Files
The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…
- Agent Harness Engineering
Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…
- Loop Engineering
Replacing yourself as the agent's prompter by designing the system that prompts it: a recursive-goal loop built from fi…
