metafetch

Metafetch reads papers now — any URL property, and all six authors

A note about an arXiv paper keys its link as `arxiv:`, not `url:` — so metafetch couldn't see it at all. Now it shows you every URL in your frontmatter and lets you pick, comes back with the whole author list when the source is scholarly, and can stamp a stable code so the note is referenceable anywhere.

Why Care?

You want a note about a paper. Not a citation buried in a document — a document whose subject is the source, with the metadata in its frontmatter where Bases and Dataview can reach it.

So you make a note and key the link the way it makes sense to key it:

YAML
arxiv_url: https://arxiv.org/abs/2512.01939

Metafetch couldn't see that. It read exactly one property — url — and if that property wasn't there, it gave up. Every source note keyed arxiv:, ssrn:, nature:, doi:, or techcrunch: was invisible to a plugin whose entire job is fetching metadata for URLs.

And when you did get it to fetch an arXiv page, it came back with a title, an abstract, and the arXiv logo. No authors. No publication date. For a paper, those are most of the point.

Both fixed.

What's New?

  • A URL picker. New command, Fetch from a frontmatter URL…: it finds every http(s) value in the note's frontmatter — under any property name — and lets you choose which to fetch, with a provider selector.

  • Scholarly metadata. Academic publishers use a different meta-tag standard than the rest of the web. We read it now: all authors, publication date, title.

  • Your og_image won't be offered as a fetch target. Twice over — by field name and by file extension.

  • Fetching arxiv: no longer invents a url: key you didn't ask for.

  • Optional vault-wide identity codes, so a source note can be referenced by something stabler than its filename. Off by default.

Real output, same note, before and after:

YAML
# before
og_title: "An Empirical Study of Agent Developer Practices in AI Agent Frameworks"
og_site_name: "arXiv.org"
og_type: "website"

# after — the same fields, plus
authors:
  - Yanlin Wang
  - Xinyi Xu
  - Jiachi Chen
  - Tingting Bi
  - Wenchao Gu
  - Zibin Zheng
og_published: "2025-12-01"

Obsidian types both correctly — authors renders as a real list property, og_published as a date.

Why one property was never going to be enough

The old code, in full:

TS
const url = typeof fm.url === 'string' ? fm.url : null;
if (!url) { new Notice('Metafetch: no `url` field in frontmatter'); return; }

The temptation is to fix this with a precedence list — try url, then arxiv, then doi, then… But there's no correct order. A note can hold several URLs that are all legitimately fetchable: an abs page and a PDF, a DOI and a publisher landing page. Which one you meant is a judgement, and the person making it is sitting right there.

So: show what's there, let them choose.

The interesting part turned out to be what to leave out. A note that's been fetched before carries og_image and og_favicon — both URLs, neither something anyone wants to fetch metadata from. Two exclusion mechanisms, because one isn't enough: by configured field name (handles the defaults) and by image extension (catches the same values when you've renamed those fields). Array values get enumerated too — mirrors[0], mirrors[1] — since a paper commonly has both an abs link and a PDF.

Three reasons arXiv came back empty

Not one bug. Three, stacked.

1. Wrong standard. arXiv publishes scholarly metadata as Highwire Press citation_* tags. We only looked for article:author and name=author. Sitting on that page unread: citation_author ×6, citation_date, citation_title, citation_pdf_url, citation_arxiv_id.

2. First match only. getMeta() returned the first regex hit. arXiv emits one tag per author:

HTML
<meta name="citation_author" content="Wang, Yanlin" />
<meta name="citation_author" content="Xu, Xinyi" />
<meta name="citation_author" content="Chen, Jiachi" />

So even after teaching it the right tag name, a six-author paper would have come back with one author and five people silently dropped. getMetaAll() collects repeated tags; getMeta() is now a one-line wrapper over it.

3. Wrong shape. arXiv writes Wang, Yanlin and 2025/12/01. cite-wide's citation schema validates "FirstName LastName", and everything downstream wants ISO dates. Both normalized on the way in — with a guard, so Jane Doe, PhD doesn't become PhD Jane Doe.

Tiers are first-non-empty, not merged

Worth naming because the alternative looks reasonable until it isn't. A journal page can carry a generic author tag naming one person and a complete citation_author set. Merging them lists the same people twice under two spellings. So the tiers are tried in order and the first non-empty one wins.

citation_* leads, because it's complete where the generic tags are lossy. Pages without it fall straight through to what they do publish — which is why nothing changed for the rest of the web.

One small defensive touch: article:author is frequently a Facebook profile URL rather than a name, so URL-valued authors are dropped.

What checking first saved us from building

The plan for this called for a JSON-LD schema.org tier, on the reasoning that trade press would need it.

Before writing it, we pulled a live TechCrunch article and looked. It serves <meta name="author"> and <meta property="article:published_time">, and correctly reports og:type: article. It already worked. Verified through the real code path — Lucas Ropek, 2026-08-17T21:27:05+00:00, no changes required.

So the JSON-LD tier was dropped as speculative. The gap was never "articles", it was scholarly sources specifically — trade press is well served by the generic tags, and academic publishers are the ones on a separate standard. An hour of building the wrong abstraction, avoided by one curl.

An identity code, so the note can be pointed at

A source note is only worth making if you can reference it later. And the usual way — [[An Empirical Study of Agent Developer Practices]] — is more fragile than it looks. It breaks when you rename the file, and it goes ambiguous the moment two notes share a title.

So there's now an opt-in setting: stamp a short, vault-unique code onto notes as they're fetched.

YAML
hex_code: k4m2x9

That code is stable. Rename the file, move it between folders, roll the vault up into a published site — every reference still resolves.

It's off by default. Everything else metafetch writes is fetched — it came from the page. A code is different: it's a commitment the plugin invents on your behalf, and it isn't reversible in the way a re-fetch is. That should be a choice you make, not a surprise. Field name and length are configurable, like every other field metafetch writes.

These are not hexadecimal

The name is historical, and it's worth being explicit because the obvious reading of "hex_code" is wrong. We deliberately do not limit the alphabet to [0-9a-f]. Six characters of [a-z0-9]:

AlphabetSize6-character space
hex [0-9a-f]1616.7 million
ours [a-z0-9]362.18 billion

Identical six bytes on disk. Roughly 130× the room, and therefore ~130× less likely that two notes ever collide. There's no argument for paying the narrower alphabet's price when the wider one is free.

Generation uses rejection sampling rather than byte % 36. Modulo would make the first four letters of the alphabet very slightly more likely than the rest, because 256 doesn't divide evenly by 36. A tiny bias — and a silly one to accept in an identifier whose entire purpose is not colliding.

Two rules the implementation actually enforces

Write-once. An existing code is never overwritten or regenerated. This is the whole ballgame: a code that changes on the next fetch is worse than no code, because every reference pointing at the old one breaks silently and you won't find out until much later.

Uniqueness is checked, not assumed. Every mint reads the codes already in use across the vault. 2.18 billion makes collisions remote, but "remote" and "checked" aren't the same thing.

There's a subtlety we got wrong on the first pass. Obsidian's metadata cache lags actual file writes — so during a batch run, a code written to a file seconds ago isn't visible in the cache yet, and the next file could be handed the same one. The first version of this shipped with a comment confidently claiming that re-reading per file handled it. It didn't. Batch runs now thread a run-scoped set alongside the cache lookup, verified against a 200-file simulated batch.

Under the hood

UnitJob
frontmatterUrls.tsFind fetchable URLs; exclude metafetch's own output
SelectUrlModal.tsThe picker — built on Obsidian's native Setting
getMetaAll()Every value of a repeated meta tag
normalizeAuthorName()Last, First → First Last, credential-safe
normalizeDate()Highwire slashes → ISO, timestamps untouched
hexCode.tsMint, collision-check, and stamp vault identity codes

The modal is built from Obsidian's Setting component rather than custom markup, and that's deliberate: while building it we found metafetch's esbuild CSS entry point is src/styles/citations.css, which contains nothing but .cite-wide-* classes. The plugin has been shipping a sibling plugin's stylesheet, and its own .metafetch-modal / .opengraph-* classes are defined nowhere. Native components are styled by the app regardless of that. The CSS problem is real and gets its own fix.

The existing commands are untouched. runFetchScript takes an optional explicit URL; without it, the fm.url path behaves exactly as before — only the failure notice changed, to point at the picker.

Tests are up to 71 across four units, all wired into pnpm test: 16 for frontmatter round-tripping, 14 for URL discovery, 22 for the scholarly helpers, 19 for identity codes — including a guard that the full 36-character alphabet actually shows up in output, so we can't silently regress into real hex.

What's Next

The third phase is the source-note shape — and the schema for it already exists. cite-wide's Lossless-Citation-Standards.md defines 24 fields across 12 publisher_type values, and that taxonomy already names both cases by hand: Academic-Working-Paper lists arXiv, Industry-Media lists TechCrunch. No vocabulary needs inventing; it needs implementing.

Two things deliberately left alone:

  • og_type still says website on arXiv. That's arXiv's literal og:type value. Overriding it would mean writing something the source didn't say. Inferring "working paper" is publisher_type's job, in Phase 3.

  • Multi-URL notes still share one set of og_* fields, so fetching a second URL overwrites the first's. Prefixing per source property would fix it but breaks the configurable-field-name contract. Unresolved on purpose rather than settled by accident.

Full write-up, including the relationship to the open grab-reference decision: context-v/specs/Metafetch-Source-Notes-From-Any-Journal-URL.md in the parent content-farm repo.

Also shipped today: [[2026-08-17_01]] — the frontmatter quoting fix that's the reason the authors list above renders as clean bare items instead of six quoted strings.