Blog
Design systems
Engineering

Portable Text, and why rich text should be data

Rich text stored as HTML is a string with structure in it that only a browser can read. Stored as an array of blocks it is data: queryable, transformable, and renderable by something that is not a browser.

Priya Raman

Rich text stored as HTML is a string with structure inside it that only a browser can read. Stored as an array of blocks, it is data: queryable, transformable, and renderable by something that is not a browser — which is the requirement that shows up the first time content has to reach an app, an email, or a feed.

Blocks, spans, and marks

A paragraph is a block with a style. The words in it are spans, and a span's `marks` array says what is applied to it — `strong`, `em`, `code`, or the key of a link definition held beside the block. Nothing in that structure implies a tag, which is what lets one document become a web page, a native view and a plain-text email without being parsed three times.

The thing teams miss is that marks are per span rather than per range. A sentence with a bold phrase in it is three spans, not one span with offsets — more data, and much less ambiguity, since there is no way to describe a range that has drifted out of step with the text it pointed into.

If your rich text cannot be counted, filtered or migrated, you have a blob with formatting in it.

What the shape buys you

  • A renderer per surface, and no parser anywhere.
  • Queries over structure — count the posts with an image in the body.
  • Validation of a document rather than of a string, with a path to point at.
  • Migration by transformation, run once, checked before it is kept.

The renderer is a switch over `_type` and `style`, which is about thirty lines, and it is the only place in a codebase that knows what a paragraph looks like. That is the whole return on the array.

for (const node of body) { if (node._type === "block" && node.style === "h2") renderHeading(node) }

The same body is what the studio edits, and the editor is a view of the array rather than a source of truth about it — which is why a block pasted in from somewhere else is validated on the way in, not on the way out.

Priya Raman

Design systems lead

Priya designs the studio and the interfaces content teams stare at for six hours a day. She came to content tooling from editorial design, which she says is the same job: deciding what a reader sees first, and defending that decision against everybody who wants to add one more thing.

Related posts

Content operations
Engineering

Your content model is the product decision

Every content system fails the same way: a field that grew a second meaning, and a template that reads it both ways. The model is the one thing you will still be living with in three years.

Maya Chen
Engineering

GROQ in twenty minutes

There is no join, because there is no second table. A query says which documents and then what to return, and everything that looks like syntax is one of those two halves made more specific.

Tomás Ferreira
Engineering
Content operations

Draft, publish, and the two-row trick

Achar stores a draft as a second document whose id begins `drafts.` rather than a flag on one row. It looks like duplication, and it is the reason a headline can be rewritten for a week without touching what the site serves.

Tomás Ferreira