Blog
Engineering

Modelling references without modelling yourself into a corner

Copying an author's name into every post they wrote is the fastest thing to build and the slowest thing to fix. A reference writes it once, and the cost is one extra step on every read.

Tomás Ferreira

Copying an author's name into every post they wrote is the fastest thing to build and the slowest thing to fix. A reference writes that information once and points at it, and the cost is one extra step on every read — a trade that is obviously right at the tenth post and obviously wrong to argue about at the first.

References are pointers, not joins

A reference is an id in a wrapper: `{ _ref: "author-maya-chen", _type: "reference" }`. It carries no target type, deliberately. A type inside the reference would mean renaming a type is a migration of every document that mentions it, and what a reference may point at is the schema's business anyway. A query's `->` is what turns the pointer into the document.

The uncomfortable part is deletion. Nothing stops a reference from pointing at a document that is gone, and nothing should: a strict foreign key over content makes one editor's save wait on another editor's write. Instead the read is written to survive a missing target, and a sweep finds the orphans on a schedule.

A reference that cannot dangle is a reference that blocks somebody else's save.

Two habits that keep it healthy

  • Project the fields you need across the reference, and nothing else.
  • Write the renderer so a missing target draws nothing rather than throwing.
  • Reach outwards with `^` when a nested projection needs a field from outside it.
  • Sweep for orphans on a schedule rather than on every write.

Done that way, a rename is one document, a deletion is a decision somebody makes on purpose, and the content stays readable while both happen.

*[_type == "post"]{ title, "author": author->{ name, role }, "categoryCount": count(categories) }

Counting a list of references without dereferencing any of them is worth noticing: `count(categories)` answers from the document in hand, and a page that only needed the number never reads the categories at all.

Tomás Ferreira

Principal engineer

Tomás wrote the query engine behind Achar and has strong opinions about what a query language should refuse to do. Previously he built a document store that outlived three rewrites of its frontend, which is where most of those opinions came from.

Related posts

Content operations
Engineering

Your content model is the product decision

Every content system fails the same way: a field that grew a second meaning, and a template that reads it both ways. The model is the one thing you will still be living with in three years.

Maya Chen
Engineering

GROQ in twenty minutes

There is no join, because there is no second table. A query says which documents and then what to return, and everything that looks like syntax is one of those two halves made more specific.

Tomás Ferreira
Engineering
Content operations

Draft, publish, and the two-row trick

Achar stores a draft as a second document whose id begins `drafts.` rather than a flag on one row. It looks like duplication, and it is the reason a headline can be rewritten for a week without touching what the site serves.

Tomás Ferreira