Prerendering an SPA so its content routes are crawlable

A single-page app serves one empty HTML shell for every route, so 21 documentation pages and every blog post shared the same title and description. Writing real HTML per route at build time is a postbuild script, not a framework migration.

One blank document duplicated across many routes on the left, and distinct filled documents per route on the right

Everything a client-rendered app knows about a page is decided after JavaScript runs. Before that it is one HTML file with an empty root element, served identically for every URL. That is fine for an application behind a login and wrong for anything you want linked, shared or found.

The problem

The documentation site has 21 pages and the blog has fifteen posts. All of them were the same document as far as anything outside a browser was concerned: same <title>, same description, no content in the body.

Three things break, in increasing order of how much they matter.

Search engines do execute JavaScript, so the pages are indexable in principle, but the rendered content arrives in a second pass that is neither guaranteed nor prompt. Link previews do not execute JavaScript at all, so a post shared in Slack or on LinkedIn shows the site's generic description rather than the post's. And a reader on a slow connection gets nothing until the bundle arrives, on a page whose content is three paragraphs of text.

The framework answer is to move to something with server rendering. That is a large change to make for a set of pages that are static once built.

What prerendering at build time means

After the normal production build, a script walks the routes, and for each one writes a real HTML file with that route's metadata baked into the shell:

build/facility-siting/index.html
build/blog/control-flow-in-edges-not-prompts/index.html

Each file is the same JavaScript bundle with a different head. The metadata is applied by rewriting the built template rather than templating it from scratch:

html = replaceTag(html, /<title>[\s\S]*?<\/title>/, `<title>${escapeHtml(meta.title)}</title>`);
html = replaceTag(
  html,
  /<meta\s+property="og:description"[^>]*\/>/,
  `<meta property="og:description" content="${escapeHtml(meta.description)}" />`
);

Title, description, the Open Graph and Twitter tags, a canonical link, and JSON-LD where the page type warrants it. The same pass emits sitemap.xml and robots.txt, so the route list has exactly one definition rather than one for the router and another for the sitemap.

That regex rewriting deserves a caveat rather than a defence: it is string surgery on generated HTML, and it depends on the tags existing in the source template in a matching shape. It is tolerable because the input is a file in the same repository, not arbitrary HTML, and because replaceTag leaves the document untouched when a pattern does not match, so a missing tag is a missing tag rather than a corrupted file.

Serving it: static files beat the SPA rewrite

Prerendering only helps if the host serves those files rather than the shell. A single-page app needs a catch-all so deep links reach the router:

/*  /index.html  200

Read literally that would serve the shell for /facility-siting too, undoing the work. It does not, because the host checks for an existing file first and only falls through to the rewrite when there is none. Prerendered routes hit their own index.html; anything else hits the router. The precedence is doing all the work, and it is invisible in the config.

One related setting matters more than it should. Netlify's pretty URL handling redirects /blog/post to /blog/post/ by default, so every shared link cost a 301 before it rendered. Turning it off serves the prerendered file directly:

[build.processing.html]
  pretty_urls = false

What it costs

A second renderer. The blog stores content as a typed block tree, and prerendering needs that tree as HTML at build time, in Node, without React. So there is a blocksToHtml alongside the React components that render the same blocks in the app. Two renderers over one tree, and nothing forces them to agree. A block type added to one and not the other degrades silently in whichever was missed.

Build-time coupling to the route list. A page that exists in the router but not in the route manifest is not prerendered and reverts to the old behaviour without failing anything. The docs site guards this by asserting every slug in its page manifest produced a file, and failing the build otherwise. The blog does not have the equivalent check.

Staleness. The HTML is correct as of the build. Content published afterwards is not on the site until something rebuilds, which is a property of the whole build-time content approach rather than of prerendering specifically.

Where it breaks

Only the head is real. The body is still an empty root element. This fixes metadata, link previews and canonical URLs; it does not put the page's text in the initial HTML. A crawler that does not execute JavaScript still sees nothing to read. Full static rendering of the content into the body would fix that and is not implemented.

JSON-LD is applied per route by hand. It is correct where it exists and absent where nobody added it. Nothing validates the structured data, so an error there is silent until a search console notices it weeks later.

The sitemap asserts a change frequency it cannot know. Values like weekly and monthly are declared per route in the manifest and are a guess. They are hints, and search engines mostly ignore them, but they are stated with more confidence than they deserve.

The check only exists on one of the two sites. The docs build fails if a page is missing from the output. The product build does not, so a blog post that silently fails to prerender ships.

More on what the documentation covers and what the system is for. FRED itself is at biofred.us.

References

  1. Google Search Central. JavaScript SEO basics.

  2. Open Graph protocol. ogp.me.

  3. Netlify. Redirects and rewrites.