Open Knowledge Format (OKF)

Canonical version: Open Knowledge Format (OKF).

The Open Knowledge Format (OKF) is an open specification from Google Cloud for representing the metadata, context, and curated knowledge that AI systems need, in a portable, vendor-neutral, human- and agent-readable form. Version 0.1 shipped on 2026-06-12, authored by Sam McVeety and Amir Hormati of the Google Cloud Data team. The interesting part for me: it is basically the Obsidian vault plus AGENTS.md pattern turned into a standard. Plain Markdown files with YAML frontmatter, living in Git, readable by a human with cat and by an LLM verbatim. If you already keep a knowledge base this way, you already speak most of OKF.

Update (October 2026): v0.2 shipped on 2026-07-24 and added trust signals (provenance, verification, freshness, attested computations). In August 2026 the spec moved to its own repository, GoogleCloudPlatform/open-knowledge-format. The v0.1 details below still hold unless the v0.2 section says otherwise. OKF also reached a much wider audience on launch day, when Marie Haynes posted on X that it "could replace Notion or Obsidian" (about 1M views). I address that claim further down.

Why it matters

  • Every agent builder is solving the same context-assembly problem from scratch, and every catalog vendor is reinventing the same data model. The knowledge itself stays locked behind whichever tool created it
  • OKF formalizes the "LLM wiki" pattern that emerged organically: Andrej Karpathy's LLM-wiki gist (April 2026, ~16M views), the AGENTS.md / CLAUDE.md convention used across 60,000+ open-source projects, and Obsidian vaults wired to coding agents
  • It picks the lowest-friction substrate that already exists (Markdown + frontmatter + Git) instead of inventing a new one. No JSON schema registry, no protobuf, no required SDK

What it is

  • An open spec, not a product or platform. Apache 2.0, published in the GoogleCloudPlatform/knowledge-catalog repo
  • A format, deliberately. "Not tied to any specific cloud, database, model provider, or agent framework. It will never require a proprietary account or SDK"
  • v0.1 is small on purpose: ~451 lines, fits on one page. Explicitly "a starting point, not a finished standard," versioned for backward-compatible growth

File format and structure

  • A bundle is a directory of UTF-8 Markdown files. Ship it as a tarball, host it in any repo, mount it from any filesystem
  • Example layout:
sales/
├── index.md          # directory listing (reserved)
├── log.md            # chronological history (reserved)
├── tables/
│   ├── index.md
│   ├── orders.md
│   └── customers.md
└── metrics/
    └── weekly_active_users.md
  • Reserved filenames: index.md (a directory listing, for progressive disclosure) and log.md (chronological update history). Neither is a concept document
  • Each concept is one Markdown file with YAML frontmatter. The only required field is type (a producer-defined string consumers use for routing, filtering, presentation). Consumers must tolerate unknown types and preserve unknown fields
  • Recommended optional fields: title, description, resource (canonical URI for the underlying asset), tags, timestamp (ISO 8601)
  • Conventional body sections when applicable: # Schema, # Examples, # Citations
  • Relationships are Markdown links. A link from concept A to B asserts a directed, untyped edge, so a bundle is a graph, not just a folder tree. Broken links are tolerated
  • Conformance (v0.1): every non-reserved .md file has parseable YAML frontmatter; every frontmatter block has a non-empty type; reserved filenames follow their structures when present

Example concept document:

---
type: BigQuery Table
title: Orders
description: One row per completed customer order.
resource: https://console.cloud.google.com/bigquery?p=acme&d=sales&t=orders
tags: [sales, revenue]
timestamp: 2026-05-28T14:30:00Z
---
# Schema
| Column | Type | Description |
|--------|------|-------------|
| `order_id` | STRING | Globally unique order identifier. |
| `customer_id` | STRING | FK to [customers](/tables/customers.md). |

Design principles

  • Minimally opinionated: OKF requires exactly one thing of every concept, a type field. Everything else is up to the producer
  • Producer / consumer independence: who writes the knowledge is cleanly separated from who consumes it
  • Format, not platform: no lock-in, no SDK, no account

What v0.2 added

Google's v0.2 post starts from one worry: once agents write most of the corpus, can a consumer still trust it? A human-written wiki page has someone accountable behind it. Ten thousand concepts generated overnight do not. v0.2 answers with optional frontmatter. type stays the only required field, and a v0.1 bundle stays valid.

  • Provenance: sources is a list of entries, each with a required resource (URL, bundle path or a scope like "all queries in project X") plus optional id, title and credibility signals: author, usage_count and last_modified, framed by a shared usage_window. OKF records the signals and leaves any scoring to the consumer. A claim in the body cites its source with a Markdown footnote keyed to the source id ([^ga4-schema]), so agents can reorder the list without breaking attribution
  • Trust: generated: { by, at } records who wrote the content and when it last changed; verified lists who confirmed it. Consumers derive a trust tier from verified: unverified (no key), machine-confirmed, or human-reviewed (at least one human:<id> actor). The tiers are advisory; they do not control access
  • Actors: <producer>/<version> for agents (reference_agent/gemini-2.5-pro), human:<id> for people, process:<id> for automated jobs
  • Lifecycle: status (draft, stable, deprecated; absent means stable) and stale_after, an absolute ISO 8601 instant. Staleness is a plain date comparison, with no TTL to compute
  • Attested Computation: a new concept type that carries a sanctioned computation (inline under # Computation or in a separate file), a runtime (bigquery, dbt, python…), typed parameters, an executor that returns a receipt, and an attester: deterministic code with no LLM that checks the receipt. The agent may only fill in parameter values; it must never write or edit the computation. This targets one very concrete failure: an agent reporting a revenue figure from SQL it improvised
  • Two breaking renames: timestamp became generated.at, and the body # Citations list became sources. Consumers may fall back to the v0.1 forms
  • Since 2026-08-21, every timestamp in the spec is an ISO 8601 datetime with an explicit UTC offset

The spec doubled in size: SPEC.md went from 451 lines (about 15 KB) in v0.1 to 1,006 lines (about 38 KB) in v0.2. Most of the growth is enterprise data machinery (receipts, executors, attesters) that a personal wiki will never touch.

Why it improves data sharing

  • Human- and agent-readable: no SDK or query language between the reader and the content
  • Version-controllable out of the box: bundles live in Git, so pull requests, diffs, blame, and review just work. Knowledge curation becomes normal software engineering
  • Portable and lock-in free: a bundle is a directory
  • Structured where it helps, prose where it matters: frontmatter for the few fields you query or index on, Markdown body for the schemas, prose, and example queries humans and LLMs actually read
  • Composes with existing tools: Notion, Obsidian, MkDocs, Hugo, and Jekyll already speak Markdown plus YAML frontmatter
  • Progressive disclosure built in: auto-generated index.md files let an agent walk the hierarchy one level at a time instead of loading the whole bundle into context
  • Graph-shaped, via normal Markdown links between concepts

Relationship to other standards

  • Complements Model Context Protocol (MCP), it does not compete with it. MCP governs an agent's access to tools and data; OKF describes the knowledge itself. An MCP server can expose an OKF bundle as a knowledge source
  • Does not replace domain schemas like Avro, Protobuf, or OpenAPI. OKF references them rather than subsuming them
  • It is positioned as formalizing an emerging practice, not as a rival to older open-data standards (schema.org, DCAT are not addressed)

Tooling and repo

  • Repo: GoogleCloudPlatform/knowledge-catalog (Apache 2.0, primary language Python, created 2026-05-04, ~3.3K stars at first read). README disclaims it is "not an official Google product"
  • The okf/ directory holds the SPEC.md, three sample bundles (GA4 e-commerce, Stack Overflow, Bitcoin public datasets), and reference code
  • Reference implementations shipped with it:
    • Enrichment agent: walks BigQuery datasets and drafts OKF concept docs with citations, schemas, and join paths. Built on Google's Agent Development Kit (ADK) with Gemini as the model
    • Static HTML visualizer (viz.html): a self-contained graph view using Cytoscape.js and marked.js, no backend
    • kcmd: a TypeScript CLI plus MCP server for bidirectional sync between local OKF files and Google Cloud Knowledge Catalog

Status in October 2026

  • New home: OKF now lives in GoogleCloudPlatform/open-knowledge-format (created 2026-08-11, Apache-2.0, 683 stars and 48 forks on 2026-10-03). The okf/ folder in knowledge-catalog is a "frozen snapshot, no longer maintained". The 9,350 stars on knowledge-catalog count the whole Knowledge Catalog toolbox, not OKF alone
  • No releases or tags in the new repo yet. The reference agent's Python package is version 0.1.0 and its CLI has two commands, enrich and visualize. No conformance validator ships with it (a validate pull request, #238, was closed without being merged)
  • Reference agent requirements: BigQuery credentials (your project pays for query bytes) and Gemini through an AI Studio key or Vertex AI. It runs a BigQuery metadata pass, then a web pass that crawls the seed URLs you give it, capped by --web-max-pages and limited to allowed hosts
  • Samples: a fourth bundle, acme_retail (a fictional retailer), exercises every v0.2 family. The visualizer now shows trust tier, status and staleness next to the graph
  • Knowledge Catalog connector: kcmd (built with Bun) carries seven frontmatter keys (title, description, tags, resource, type, generated, sources) plus the body. Only .md files travel, cross-links "resolve to nothing" in the catalog, tags become labels, and the docs call scale "untested beyond the 14-file demo"
  • Open questions in the tracker: typed relationship edges (#463, #322), an .okfignore convention (#190), an authoritative site for OKF (#231), and whether log.md may carry frontmatter (#286: the acme_retail sample's own log.md has a type: Log block). The spec also names no Markdown flavor (CommonMark, GFM), a gap an HN commenter spotted on day one

Adoption and caveats

  • At launch (2026-06-12) every producer and consumer was built by Google. A v0.1 spec from a single vendor is an invitation, not yet a standard
  • Whether OKF becomes a common format depends on producers outside Google adopting it. Catalog vendors like Atlan, Alation, and Collate are the ones to watch
  • Worth noting the commercial shape: OKF gives away the cheap part (a file any editor can open) and points the demand it creates toward the part Google sells, Google Cloud Knowledge Catalog. That is a smart play, not a knock; open formats with a paid serving layer have a long track record

Community reception

  • The main Hacker News thread ("Google proposes Open Knowledge Format based on Markdown", 2026-06-13, 87 points, 15 comments) was lukewarm. One commenter summed it up as "Markdown with YAML front-matter" with "15kb of spec". Another argued you can't represent knowledge well without labeled relationships. A third joked about revisiting RDF/OWL every ten years. The defense: Markdown wins because it is "the lowest common denominator" for humans and models. Someone also flagged broken links in Google's original post
  • Builders shipped around it anyway. HN submissions between June and September 2026 include okf-lint, signed-okf (signed provenance), Kage (verification and freshness for agent memory), Kiso (a publishing engine), okfctl, an "open knowledge compiler", and MCP Memory, agent memory on OKF plus SQLite FTS5 (70 points). In the MCP Memory thread, people asked the obvious question: why not plain Markdown files and grep? The answers were indexing speed once you have hundreds of notes, and MCP access from surfaces with no filesystem
  • The v0.2 post says contributors opened extension proposals (typed edges, agent-routing hints, an erasure profile, .okfignore) and started cataloging tools built outside Google

How OKF compares

OKF Obsidian vault (OSK) LLM Wiki (Karpathy) llms.txt AGENTS.md
What it is A file format spec An app over a folder of Markdown files, plus OSK conventions A usage pattern described in a gist One index file at a website's root An instructions file for agents
Unit One concept per .md file One note per .md file Entity and concept pages One file per site (plus llms-full.txt) One file per repo or folder
Required metadata type in frontmatter None in Obsidian; OSK adds per-type tags such as type/permanent_note None; a schema doc defines conventions None None
Links Standard Markdown links, /-rooted bundle paths recommended, untyped **wikilinks** by default Cross-references between pages Markdown links to URLs Free-form
Index and log Reserved index.md and log.md at any level No reserved names index.md and log.md at the root The file is the index None
Main reader Agents and humans, across organizations The human, with agents more and more The LLM maintains it, the human reads it LLMs visiting a site Agents working in a repo
  • Obsidian vault with OSK conventions: same substrate (Markdown, YAML properties, folders, Git). The gaps are mechanical. OSK encodes the note type in tags and folders, while OKF wants a type: key. Obsidian writes **wikilinks**, which the spec doesn't define, so an OKF consumer gets no edges from them. Conformance requires frontmatter with a non-empty type in every non-reserved .md file, which a real vault fails on day one: templates, READMEs, Excalidraw drawings, notes without properties. A note named index.md or log.md would be read as a reserved file. The overlaps are good news: OSK's description maps one to one, and created/updated plus ai_generated map loosely to generated: { by, at }. Exporting a vault as a bundle looks like a script job: derive type from the type tag, rewrite wikilinks as Markdown links, generate the index.md files
  • LLM Wiki: Google says OKF "formalizes the LLM-wiki pattern". Karpathy's index.md and log.md became reserved filenames, and his raw sources map to the references/ convention. Two differences: OKF's log.md wants plain YYYY-MM-DD date headings with entries as bullets under them, whereas Karpathy's log puts one operation per heading (## [2026-04-02] ingest | Article Title). And his third layer, the schema document that tells the LLM how to maintain the wiki, has no place in OKF
  • AGENTS.md: it tells an agent how to behave; OKF describes what is known. They complement each other. One catch: an AGENTS.md or README.md inside a bundle is a non-reserved .md file, so it needs type frontmatter to stay conformant. Google lists the AGENTS.md / CLAUDE.md family among the patterns OKF generalizes, yet the spec has no slot for instructions
  • llms.txt: both rely on progressive disclosure through lists of Markdown links. llms.txt is one file per website. OKF's index.md can sit in every folder, carries no frontmatter (except okf_version at the root) and lists * [Title](url) - description entries, ideally pulled from each concept's description. A published bundle could generate its llms.txt from the root index with little effort

Does OKF replace Obsidian?

Short answer: no. OKF is a file format; Obsidian is an app. Marie Haynes' post ("Seems to me this could replace Notion or Obsidian") mixes the two up, and the spec itself says why it can't be true.

  • OKF has no app to replace anything with. Its non-goals include "prescribing storage, serving, or query infrastructure". Google's own README lists Obsidian, Notion and MkDocs as UIs that can consume OKF bundles
  • An Obsidian vault already is a folder of Markdown files with YAML frontmatter. That is most of the OKF spec. Switching from Obsidian to OKF makes as much sense as switching from VS Code to UTF-8. You can do both at once: keep editing in Obsidian and make the files conformant
  • The "agents query and edit a living wiki" part comes from the agent, not the format. The spec defines no query language, no edit API and no sync. Google's tooling is a BigQuery-to-OKF generator, a static HTML viewer and a catalog connector; none of them lets you write a note. The editing, linking, search, graph, plugins, sync and publishing still come from somewhere else, and for many people that somewhere is Obsidian
  • Notion is the more interesting case. Notion keeps your content in its own cloud database in a proprietary format. OKF gives a way out of that kind of lock-in. Obsidian users have had one all along: their files, per the File over app principle
  • Look at who the examples serve: BigQuery tables, metrics, incident playbooks, revenue computations checked by attesters. v0.2 doubled the spec for enterprise data trust. Nothing in it covers daily notes, reading notes, tasks or anything else people do in a PKM system

Where OKF does matter for PKM: if agents and tools start to expect type, description, sources and generated/verified, a vault that already carries those fields becomes readable by any OKF consumer without translation. That's a reason to make your notes convert cleanly. Not a reason to switch apps.

My take

This validates the approach I have used for years. Markdown notes, YAML frontmatter, Git underneath, and an AGENTS.md / CLAUDE.md at the root telling agents how to behave. OKF is, more or less, the standardized version of an OSK vault used as an LLM wiki. If it gets traction, the knowledge bases people already keep in Obsidian become directly consumable by any agent that speaks OKF, with no export step. That is worth keeping an eye on, and worth structuring my own knowledge so it would convert cleanly.

Four months on, I read v0.2 as Google picking a lane. TypedMark goes further on what a type is (schemas, inheritance, validation); OKF v0.2 goes further on whether a machine-written concept can be trusted (sources, verification, staleness). Both sit on the same files. For anyone keeping notes in Obsidian, the useful question remains how cleanly a vault converts to a bundle. Whether OKF replaces Obsidian was never a real question.

References


About Sébastien

Ready to get to the next level?

Found this valuable? Share it with someone who needs it.

Join 6,000+ readers. Get practical systems for knowledge & AI. Free.

Subscribe ✨

Free: Knowledge System Checklist

A clear roadmap to building your own knowledge system. Subscribe and get it straight to your inbox.

6,000+ readers. No spam. Unsubscribe anytime.

Subscribe