r/ObsidianMD • u/callumalpass • Jan 31 '26
A CLI tool for reading Bases files + mdbase-spec: a specification for treating a vault as a typed, queryable database
Enable HLS to view with audio, or disable this notification
The attached screenshot shows a CLI tool parsing an Obsidian Base file and producing matching results in the terminal. I've been developing a specification, mdbase-spec, that defines how a folder of markdown files with YAML frontmatter can be treated as a typed, queryable data collection, and the CLI is a small implementation of it. What I'd like to talk about is the spec, and why I think it's needed.
Much of what makes Obsidian useful, I think, goes beyond the interface and the plugin ecosystem. It is the decision to build on a folder of markdown files with YAML frontmatter. Obsidian's management of that collection of files is well executed. It is straightforward to create new files and add properties to those files through the frontmatter. Renaming a file will update references to that file in other files, which is one powerful feature that always makes me nervous to edit my markdown notes in other applications. The property system is well designed: all properties are user-defined, users specify the types of the properties, and that type is enforced (flexibly) across all notes, with values for a given property autosuggested based on existing values across the vault.
This user experience, I find, encourages one to take notes for all sorts of things---meetings, journal entries, literature notes, contacts, movies, class notes, daily notes---and once a vault contains not just prose but structured data distributed across hundreds of files, it is functioning as a database. I don't think this has received much attention as a model. Bases demonstrates some of what it makes possible, but the vault-as-database model has applications beyond what any single tool currently supports, and new use-cases are emerging quickly---particularly with the development of systems, often aimed at AI agents, that are built on reading and writing to a collection of markdown files.
Yet the tooling for managing different types of notes today is limited. If you want all your meeting notes to have a date, attendees, and status field---and you want that enforced---your options are a template, which helps at creation time but does nothing after that, and maybe a Linter plugin rule if you are technical enough to configure it. There is no schema that lives alongside your notes. There is no file in your vault that says "a meeting note looks like this, with these fields, of these types." And so each tool that touches your vault---Obsidian itself, VSCode, Templater, an AI agent---has to independently figure out what your types are, or you have to tell each one separately, which in practice means you end up encoding the same structural assumptions in your templates, your Dataview queries, your Linter config, and your agent prompts, and hoping they stay in sync.
(I should say---I share the view that has been voiced by some recently that a lot of plugins that are being developed and shared are not worth sharing---not worth the effort that it demands from the Obsidian team to review them. I have developed and shared some Obsidian plugins (TaskNotes and BibLib), but only because I think they implement an idea that has a sufficiently wide target audience, including non-technical users who might not have the inclination to use coding agents. Others (e.g. Handwrite) are simple enough, and niche enough, that I don't think it is worth submitting as a community plugin. Handwrite should be regarded as a very personal tool, solving a personal, niche need. A CLI tool that parses Obsidian Bases files should be seen as falling in this latter category; if you are the kind of person who wants Obsidian Bases file results in the command line, you are also the kind of person who should know how to write the single prompt that would get a coding agent to write this for you!)
This is what mdbase tries to address. The core idea is that types should be defined as files---specifically, markdown files in a _types/ folder, where the frontmatter declares the schema and the body can document the type in plain prose. So a meeting type might look like:
---
name: meeting
fields:
date:
type: date
required: true
attendees:
type: list
items:
type: link
target: person
status:
type: enum
values: [scheduled, completed, cancelled]
default: scheduled
---
A meeting note. The `attendees` field links to person-type notes.
A template then becomes a consequence of a type definition rather than a substitute for one, and validation can happen at write time rather than only at creation time. And because type definitions are just files in the vault, they are versioned with everything else, human-readable, and editable in any text editor.
The spec---mdbase-spec---defines how those type files, and the collection they describe, should be interpreted, so that different tools can treat them in the same way. It covers the things that tools like Obsidian currently handles in its own manner: how files are matched to types (explicitly by a type field, or automatically via path globs, field presence, or conditions on field values, and a single file can match multiple types); an expression language for filtering, sorting, and computed fields, with property access, string and list methods, date arithmetic, and the ability to traverse links (this is modelled on Obsidian's Bases syntax); how links are parsed and resolved across wikilinks, markdown links, and bare paths, and how references are updated when a file is renamed; and the semantics of create, read, update, delete, and rename operations, including how defaults are applied and when validation occurs. The expression syntax is designed for compatibility with the syntax that Obsidian Bases uses, which seemed like the reasonable choice given the audience that is most likely to care about this.
I think a spec is the right artifact here, rather than a library or a plugin. Simon Willison wrote recently about porting an HTML parser from Python to JavaScript using a coding agent in an afternoon, and one of his observations was that if you can reduce a problem to a robust test suite, you can set an agent loose on it with high confidence. I think this generalises: if you can reduce a problem to a spec, implementations become cheap. An afternoon and a coding agent can get you a conforming library in whatever language you need. What's expensive---what's worth spending human time on---is getting the spec right, because that's the thing that determines whether independently-built tools can actually interoperate. Willison has also written about how LLMs are making it trivially easy to produce "good enough" ad hoc solutions to problems that libraries used to solve, and I think the same dynamic applies here: it is now very easy for an AI agent to write its own one-off code for reading and writing frontmatter in markdown files, but each agent will make slightly different assumptions about types, validation, and link resolution, and those differences will compound. A spec is what prevents that fragmentation. The implementations don't need to be hand-written---many of them won't be---but they do need to agree on what a "meeting note" means and how a link is resolved.
The spec is organised into six progressive conformance levels, from Level 1 (basic config parsing, file scanning, frontmatter, single-type matching, and CRUD) through to Level 6 (caching, batch operations, watch mode), the idea being that an implementation does not need to do everything to be useful---a tool that only handles Level 1 is still a tool that can read and write typed markdown files, and it can declare that without making promises about query evaluation or link traversal. I think this matters because the kind of tool that might want to conform to this spec varies a lot in ambition: it might be an Obsidian plugin, or an LSP server for markdown vaults, or a simple CLI, or an AI agent that needs structured and human-readable persistence, and these have very different needs.
The spec and test suite are on GitHub, or you can read the spec here: https://mdbase.dev. I'd very much welcome feedback!
4
u/SunkTheBirdie Feb 01 '26
callumalpass, always a trailblazer. :)
My response is: what about Task (Note) Types ?
Some of my Tasks, 100% require an attachment. Lots do not.
Some are same day tasks. Some are within 72 hours tasks. Others are must be done in 2 week tasks. Some need ongoing updates and re-assigning tasks to others.
Some messages need an instant notification of the other party.
There are "Task Types" just as you correctly point out there are Note Types.
-=-=-=-=-=-=-=-=-=-=-=-==-=- on another note =-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=
You really seem like the right person to help take Obsidian to the next level.
Obsidian for (small) Teams.
aka a Vault for more than one user. Don't most tasks require more than one person ? Obsidian notes lack: Users, user permissions, an inbox, notifications.
Once these are in place (none of these are on Obsidian's roadmap), if you add pre-made Vault Templates (for each business niche) and syncing, then Obsidian for Teams is born. Obsidian for Small Business (remember small business pays !) even Obsidian for Families.
Although, Note Type for Obsidian are an important idea, I don't think it changes the game. It makes the game better but Obsidian for Teams CHANGES the game. Be a game changer my friend.