Open Knowledge Format (OKF)
Canonical version: Open Knowledge Format (OKF).
The Open Knowledge Format (OKF) is an open specification from Google Cloud for representing the metadata, context, and curated knowledge that AI systems need, in a portable, vendor-neutral, human- and agent-readable form. Version 0.1 shipped on 2026-06-12, authored by Sam McVeety and Amir Hormati of the Google Cloud Data team. The interesting part for me: it is basically the Obsidian vault plus AGENTS.md pattern turned into a standard. Plain Markdown files with YAML frontmatter, living in Git, readable by a human with cat and by an LLM verbatim. If you already keep a knowledge base this way, you already speak most of OKF.
Update (October 2026): v0.2 shipped on 2026-07-24 and added trust signals (provenance, verification, freshness, attested computations). In August 2026 the spec moved to its own repository, GoogleCloudPlatform/open-knowledge-format. The v0.1 details below still hold unless the v0.2 section says otherwise. OKF also reached a much wider audience on launch day, when Marie Haynes posted on X that it "could replace Notion or Obsidian" (about 1M views). I address that claim further down.
Why it matters
- Every agent builder is solving the same context-assembly problem from scratch, and every catalog vendor is reinventing the same data model. The knowledge itself stays locked behind whichever tool created it
- OKF formalizes the "LLM wiki" pattern that emerged organically: Andrej Karpathy's LLM-wiki gist (April 2026, ~16M views), the
AGENTS.md/CLAUDE.mdconvention used across 60,000+ open-source projects, and Obsidian vaults wired to coding agents - It picks the lowest-friction substrate that already exists (Markdown + frontmatter + Git) instead of inventing a new one. No JSON schema registry, no protobuf, no required SDK
What it is
- An open spec, not a product or platform. Apache 2.0, published in the
GoogleCloudPlatform/knowledge-catalogrepo - A format, deliberately. "Not tied to any specific cloud, database, model provider, or agent framework. It will never require a proprietary account or SDK"
- v0.1 is small on purpose: ~451 lines, fits on one page. Explicitly "a starting point, not a finished standard," versioned for backward-compatible growth
File format and structure
- A bundle is a directory of UTF-8 Markdown files. Ship it as a tarball, host it in any repo, mount it from any filesystem
- Example layout:
sales/
├── index.md # directory listing (reserved)
├── log.md # chronological history (reserved)
├── tables/
│ ├── index.md
│ ├── orders.md
│ └── customers.md
└── metrics/
└── weekly_active_users.md
- Reserved filenames:
index.md(a directory listing, for progressive disclosure) andlog.md(chronological update history). Neither is a concept document - Each concept is one Markdown file with YAML frontmatter. The only required field is
type(a producer-defined string consumers use for routing, filtering, presentation). Consumers must tolerate unknown types and preserve unknown fields - Recommended optional fields:
title,description,resource(canonical URI for the underlying asset),tags,timestamp(ISO 8601) - Conventional body sections when applicable:
# Schema,# Examples,# Citations - Relationships are Markdown links. A link from concept A to B asserts a directed, untyped edge, so a bundle is a graph, not just a folder tree. Broken links are tolerated
- Conformance (v0.1): every non-reserved
.mdfile has parseable YAML frontmatter; every frontmatter block has a non-emptytype; reserved filenames follow their structures when present
Example concept document:
---
type: BigQuery Table
title: Orders
description: One row per completed customer order.
resource: https://console.cloud.google.com/bigquery?p=acme&d=sales&t=orders
tags: [sales, revenue]
timestamp: 2026-05-28T14:30:00Z
---
# Schema
| Column | Type | Description |
|--------|------|-------------|
| `order_id` | STRING | Globally unique order identifier. |
| `customer_id` | STRING | FK to [customers](/tables/customers.md). |
Design principles
- Minimally opinionated: OKF requires exactly one thing of every concept, a
typefield. Everything else is up to the producer - Producer / consumer independence: who writes the knowledge is cleanly separated from who consumes it
- Format, not platform: no lock-in, no SDK, no account
What v0.2 added
Google's v0.2 post starts from one worry: once agents write most of the corpus, can a consumer still trust it? A human-written wiki page has someone accountable behind it. Ten thousand concepts generated overnight do not. v0.2 answers with optional frontmatter. type stays the only required field, and a v0.1 bundle stays valid.
- Provenance:
sourcesis a list of entries, each with a requiredresource(URL, bundle path or a scope like "all queries in project X") plus optionalid,titleand credibility signals:author,usage_countandlast_modified, framed by a sharedusage_window. OKF records the signals and leaves any scoring to the consumer. A claim in the body cites its source with a Markdown footnote keyed to the sourceid([^ga4-schema]), so agents can reorder the list without breaking attribution - Trust:
generated: { by, at }records who wrote the content and when it last changed;verifiedlists who confirmed it. Consumers derive a trust tier fromverified: unverified (no key), machine-confirmed, or human-reviewed (at least onehuman:<id>actor). The tiers are advisory; they do not control access - Actors:
<producer>/<version>for agents (reference_agent/gemini-2.5-pro),human:<id>for people,process:<id>for automated jobs - Lifecycle:
status(draft,stable,deprecated; absent meansstable) andstale_after, an absolute ISO 8601 instant. Staleness is a plain date comparison, with no TTL to compute - Attested Computation: a new concept type that carries a sanctioned computation (inline under
# Computationor in a separate file), aruntime(bigquery,dbt,python…), typedparameters, anexecutorthat returns a receipt, and anattester: deterministic code with no LLM that checks the receipt. The agent may only fill in parameter values; it must never write or edit the computation. This targets one very concrete failure: an agent reporting a revenue figure from SQL it improvised - Two breaking renames:
timestampbecamegenerated.at, and the body# Citationslist becamesources. Consumers may fall back to the v0.1 forms - Since 2026-08-21, every timestamp in the spec is an ISO 8601 datetime with an explicit UTC offset
The spec doubled in size: SPEC.md went from 451 lines (about 15 KB) in v0.1 to 1,006 lines (about 38 KB) in v0.2. Most of the growth is enterprise data machinery (receipts, executors, attesters) that a personal wiki will never touch.
Why it improves data sharing
- Human- and agent-readable: no SDK or query language between the reader and the content
- Version-controllable out of the box: bundles live in Git, so pull requests, diffs, blame, and review just work. Knowledge curation becomes normal software engineering
- Portable and lock-in free: a bundle is a directory
- Structured where it helps, prose where it matters: frontmatter for the few fields you query or index on, Markdown body for the schemas, prose, and example queries humans and LLMs actually read
- Composes with existing tools: Notion, Obsidian, MkDocs, Hugo, and Jekyll already speak Markdown plus YAML frontmatter
- Progressive disclosure built in: auto-generated
index.mdfiles let an agent walk the hierarchy one level at a time instead of loading the whole bundle into context - Graph-shaped, via normal Markdown links between concepts
Relationship to other standards
- Complements Model Context Protocol (MCP), it does not compete with it. MCP governs an agent's access to tools and data; OKF describes the knowledge itself. An MCP server can expose an OKF bundle as a knowledge source
- Does not replace domain schemas like Avro, Protobuf, or OpenAPI. OKF references them rather than subsuming them
- It is positioned as formalizing an emerging practice, not as a rival to older open-data standards (schema.org, DCAT are not addressed)
Tooling and repo
- Repo:
GoogleCloudPlatform/knowledge-catalog(Apache 2.0, primary language Python, created 2026-05-04, ~3.3K stars at first read). README disclaims it is "not an official Google product" - The
okf/directory holds theSPEC.md, three sample bundles (GA4 e-commerce, Stack Overflow, Bitcoin public datasets), and reference code - Reference implementations shipped with it:
- Enrichment agent: walks BigQuery datasets and drafts OKF concept docs with citations, schemas, and join paths. Built on Google's Agent Development Kit (ADK) with Gemini as the model
- Static HTML visualizer (
viz.html): a self-contained graph view using Cytoscape.js and marked.js, no backend kcmd: a TypeScript CLI plus MCP server for bidirectional sync between local OKF files and Google Cloud Knowledge Catalog
Status in October 2026
- New home: OKF now lives in
GoogleCloudPlatform/open-knowledge-format(created 2026-08-11, Apache-2.0, 683 stars and 48 forks on 2026-10-03). Theokf/folder inknowledge-catalogis a "frozen snapshot, no longer maintained". The 9,350 stars onknowledge-catalogcount the whole Knowledge Catalog toolbox, not OKF alone - No releases or tags in the new repo yet. The reference agent's Python package is version 0.1.0 and its CLI has two commands,
enrichandvisualize. No conformance validator ships with it (avalidatepull request, #238, was closed without being merged) - Reference agent requirements: BigQuery credentials (your project pays for query bytes) and Gemini through an AI Studio key or Vertex AI. It runs a BigQuery metadata pass, then a web pass that crawls the seed URLs you give it, capped by
--web-max-pagesand limited to allowed hosts - Samples: a fourth bundle,
acme_retail(a fictional retailer), exercises every v0.2 family. The visualizer now shows trust tier, status and staleness next to the graph - Knowledge Catalog connector:
kcmd(built with Bun) carries seven frontmatter keys (title,description,tags,resource,type,generated,sources) plus the body. Only.mdfiles travel, cross-links "resolve to nothing" in the catalog, tags become labels, and the docs call scale "untested beyond the 14-file demo" - Open questions in the tracker: typed relationship edges (#463, #322), an
.okfignoreconvention (#190), an authoritative site for OKF (#231), and whetherlog.mdmay carry frontmatter (#286: theacme_retailsample's ownlog.mdhas atype: Logblock). The spec also names no Markdown flavor (CommonMark, GFM), a gap an HN commenter spotted on day one
Adoption and caveats
- At launch (2026-06-12) every producer and consumer was built by Google. A v0.1 spec from a single vendor is an invitation, not yet a standard
- Whether OKF becomes a common format depends on producers outside Google adopting it. Catalog vendors like Atlan, Alation, and Collate are the ones to watch
- Worth noting the commercial shape: OKF gives away the cheap part (a file any editor can open) and points the demand it creates toward the part Google sells, Google Cloud Knowledge Catalog. That is a smart play, not a knock; open formats with a paid serving layer have a long track record
Community reception
- The main Hacker News thread ("Google proposes Open Knowledge Format based on Markdown", 2026-06-13, 87 points, 15 comments) was lukewarm. One commenter summed it up as "Markdown with YAML front-matter" with "15kb of spec". Another argued you can't represent knowledge well without labeled relationships. A third joked about revisiting RDF/OWL every ten years. The defense: Markdown wins because it is "the lowest common denominator" for humans and models. Someone also flagged broken links in Google's original post
- Builders shipped around it anyway. HN submissions between June and September 2026 include okf-lint, signed-okf (signed provenance), Kage (verification and freshness for agent memory), Kiso (a publishing engine), okfctl, an "open knowledge compiler", and MCP Memory, agent memory on OKF plus SQLite FTS5 (70 points). In the MCP Memory thread, people asked the obvious question: why not plain Markdown files and grep? The answers were indexing speed once you have hundreds of notes, and MCP access from surfaces with no filesystem
- The v0.2 post says contributors opened extension proposals (typed edges, agent-routing hints, an erasure profile,
.okfignore) and started cataloging tools built outside Google
How OKF compares
| OKF | Obsidian vault (OSK) | LLM Wiki (Karpathy) | llms.txt | AGENTS.md | |
|---|---|---|---|---|---|
| What it is | A file format spec | An app over a folder of Markdown files, plus OSK conventions | A usage pattern described in a gist | One index file at a website's root | An instructions file for agents |
| Unit | One concept per .md file |
One note per .md file |
Entity and concept pages | One file per site (plus llms-full.txt) |
One file per repo or folder |
| Required metadata | type in frontmatter |
None in Obsidian; OSK adds per-type tags such as type/permanent_note |
None; a schema doc defines conventions | None | None |
| Links | Standard Markdown links, /-rooted bundle paths recommended, untyped |
**wikilinks** by default |
Cross-references between pages | Markdown links to URLs | Free-form |
| Index and log | Reserved index.md and log.md at any level |
No reserved names | index.md and log.md at the root |
The file is the index | None |
| Main reader | Agents and humans, across organizations | The human, with agents more and more | The LLM maintains it, the human reads it | LLMs visiting a site | Agents working in a repo |
- Obsidian vault with OSK conventions: same substrate (Markdown, YAML properties, folders, Git). The gaps are mechanical. OSK encodes the note type in tags and folders, while OKF wants a
type:key. Obsidian writes**wikilinks**, which the spec doesn't define, so an OKF consumer gets no edges from them. Conformance requires frontmatter with a non-emptytypein every non-reserved.mdfile, which a real vault fails on day one: templates, READMEs, Excalidraw drawings, notes without properties. A note namedindex.mdorlog.mdwould be read as a reserved file. The overlaps are good news: OSK'sdescriptionmaps one to one, andcreated/updatedplusai_generatedmap loosely togenerated: { by, at }. Exporting a vault as a bundle looks like a script job: derivetypefrom the type tag, rewrite wikilinks as Markdown links, generate theindex.mdfiles - LLM Wiki: Google says OKF "formalizes the LLM-wiki pattern". Karpathy's
index.mdandlog.mdbecame reserved filenames, and his raw sources map to thereferences/convention. Two differences: OKF'slog.mdwants plainYYYY-MM-DDdate headings with entries as bullets under them, whereas Karpathy's log puts one operation per heading (## [2026-04-02] ingest | Article Title). And his third layer, the schema document that tells the LLM how to maintain the wiki, has no place in OKF - AGENTS.md: it tells an agent how to behave; OKF describes what is known. They complement each other. One catch: an
AGENTS.mdorREADME.mdinside a bundle is a non-reserved.mdfile, so it needstypefrontmatter to stay conformant. Google lists theAGENTS.md/CLAUDE.mdfamily among the patterns OKF generalizes, yet the spec has no slot for instructions - llms.txt: both rely on progressive disclosure through lists of Markdown links.
llms.txtis one file per website. OKF'sindex.mdcan sit in every folder, carries no frontmatter (exceptokf_versionat the root) and lists* [Title](url) - descriptionentries, ideally pulled from each concept'sdescription. A published bundle could generate itsllms.txtfrom the root index with little effort
Does OKF replace Obsidian?
Short answer: no. OKF is a file format; Obsidian is an app. Marie Haynes' post ("Seems to me this could replace Notion or Obsidian") mixes the two up, and the spec itself says why it can't be true.
- OKF has no app to replace anything with. Its non-goals include "prescribing storage, serving, or query infrastructure". Google's own README lists Obsidian, Notion and MkDocs as UIs that can consume OKF bundles
- An Obsidian vault already is a folder of Markdown files with YAML frontmatter. That is most of the OKF spec. Switching from Obsidian to OKF makes as much sense as switching from VS Code to UTF-8. You can do both at once: keep editing in Obsidian and make the files conformant
- The "agents query and edit a living wiki" part comes from the agent, not the format. The spec defines no query language, no edit API and no sync. Google's tooling is a BigQuery-to-OKF generator, a static HTML viewer and a catalog connector; none of them lets you write a note. The editing, linking, search, graph, plugins, sync and publishing still come from somewhere else, and for many people that somewhere is Obsidian
- Notion is the more interesting case. Notion keeps your content in its own cloud database in a proprietary format. OKF gives a way out of that kind of lock-in. Obsidian users have had one all along: their files, per the File over app principle
- Look at who the examples serve: BigQuery tables, metrics, incident playbooks, revenue computations checked by attesters. v0.2 doubled the spec for enterprise data trust. Nothing in it covers daily notes, reading notes, tasks or anything else people do in a PKM system
Where OKF does matter for PKM: if agents and tools start to expect type, description, sources and generated/verified, a vault that already carries those fields becomes readable by any OKF consumer without translation. That's a reason to make your notes convert cleanly. Not a reason to switch apps.
My take
This validates the approach I have used for years. Markdown notes, YAML frontmatter, Git underneath, and an AGENTS.md / CLAUDE.md at the root telling agents how to behave. OKF is, more or less, the standardized version of an OSK vault used as an LLM wiki. If it gets traction, the knowledge bases people already keep in Obsidian become directly consumable by any agent that speaks OKF, with no export step. That is worth keeping an eye on, and worth structuring my own knowledge so it would convert cleanly.
Four months on, I read v0.2 as Google picking a lane. TypedMark goes further on what a type is (schemas, inheritance, validation); OKF v0.2 goes further on whether a machine-written concept can be trusted (sources, verification, staleness). Both sit on the same files. For anyone keeping notes in Obsidian, the useful question remains how cleanly a vault converts to a bundle. Whether OKF replaces Obsidian was never a real question.
References
- Introducing the Knowledge Catalog: https://cloud.google.com/blog/products/data-analytics/introducing-the-google-cloud-knowledge-catalog
- How OKF can improve data sharing: https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing/
- OKF SPEC.md: https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/okf/SPEC.md
- OKF directory: https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf
- Repo root: https://github.com/GoogleCloudPlatform/knowledge-catalog
- Google Research announcement: https://x.com/GoogleResearch/status/2065475343205740911
- Marie Haynes on X, "could replace Notion or Obsidian" (2026-06-12): https://x.com/Marie_Haynes/status/2065531158356717721
- OKF v0.2 adds trust signals (Google Cloud blog, 2026-07-24): https://cloud.google.com/blog/products/data-analytics/okf-v0-2-adds-trust-signals
- OKF SPEC.md v0.2 (raw): https://raw.githubusercontent.com/GoogleCloudPlatform/knowledge-catalog/main/okf/SPEC.md
- OKF SPEC.md v0.1 (launch commit): https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/ee67a5ca27044ebe7c38385f5b6cffc2305a9c1a/okf/SPEC.md
- OKF dedicated repo: https://github.com/GoogleCloudPlatform/open-knowledge-format
- Publishing an OKF bundle to Knowledge Catalog (connector doc): https://github.com/GoogleCloudPlatform/open-knowledge-format/blob/main/connectors/gcp-knowledge-catalog.md
- Metadata as Code (
kcmd) README: https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/toolbox/mdcode - Knowledge Catalog issue tracker (OKF proposals): https://github.com/GoogleCloudPlatform/knowledge-catalog/issues
- Issue #286, log.md frontmatter: https://github.com/GoogleCloudPlatform/knowledge-catalog/issues/286
- Karpathy's LLM Wiki gist: https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
- HN, "Google proposes Open Knowledge Format based on Markdown": https://news.ycombinator.com/item?id=48517735
- HN, "Googles specification (and tooling) for the LLM wiki": https://news.ycombinator.com/item?id=48620068
- HN, "Open Knowledge format v0.2 tackles agentic trust": https://news.ycombinator.com/item?id=49059889
- HN, "Show HN: MCP Memory, fast agent memory using Google's OKF and SQLite FTS5": https://news.ycombinator.com/item?id=49286073
- HN search for other OKF tools (okf-lint, signed-okf, Kage, Kiso, okfctl): https://hn.algolia.com/?q=open%20knowledge%20format
Related
- Google Cloud Knowledge Catalog
- Model Context Protocol (MCP)
- AI Agents
- Large Language Models (LLMs)
- Obsidian
- Markdown
- Git
- Andrej Karpathy
- Open source
- TypedMark
- LLM Wiki
- AGENTS.md (File Convention)
- llms.txt convention
- Notion
- Obsidian Starter Kit
- Obsidian Properties
- File over app principle
- Personal Knowledge Management (PKM)
About Sébastien
Ready to get to the next level?
Found this valuable? Share it with someone who needs it.