Metafetch reads papers now — any URL property, and all six authors
A note about an arXiv paper keys its link as `arxiv:`, not `url:` — so metafetch couldn't see it at all. Now it shows you every URL in your frontmatter and lets you pick, comes back with the whole author list when the source is scholarly, and can stamp a stable code so the note is referenceable anywhere.
Why Care? — the note metafetch couldn't see
What's New? — the short list, with real before/after
Three reasons arXiv came back empty — the stacked failures
An identity code, so the note can be pointed at — and why it isn't hexadecimal
Under the hood — the units, the CSS problem, the tests
What's Next — Phase 3, and what we left alone
Why Care?
You want a note about a paper. Not a citation buried in a document — a document whose subject is the source, with the metadata in its frontmatter where Bases and Dataview can reach it.
So you make a note and key the link the way it makes sense to key it:
arxiv_url: https://arxiv.org/abs/2512.01939 Metafetch couldn't see that. It read exactly one property — url — and if that
property wasn't there, it gave up. Every source note keyed arxiv:, ssrn:,
nature:, doi:, or techcrunch: was invisible to a plugin whose entire job
is fetching metadata for URLs.
And when you did get it to fetch an arXiv page, it came back with a title, an abstract, and the arXiv logo. No authors. No publication date. For a paper, those are most of the point.
Both fixed.
What's New?
A URL picker. New command, Fetch from a frontmatter URL…: it finds every
http(s)value in the note's frontmatter — under any property name — and lets you choose which to fetch, with a provider selector.Scholarly metadata. Academic publishers use a different meta-tag standard than the rest of the web. We read it now: all authors, publication date, title.
Your
og_imagewon't be offered as a fetch target. Twice over — by field name and by file extension.Fetching
arxiv:no longer invents aurl:key you didn't ask for.Optional vault-wide identity codes, so a source note can be referenced by something stabler than its filename. Off by default.
Real output, same note, before and after:
# before
og_title: "An Empirical Study of Agent Developer Practices in AI Agent Frameworks"
og_site_name: "arXiv.org"
og_type: "website"
# after — the same fields, plus
authors:
- Yanlin Wang
- Xinyi Xu
- Jiachi Chen
- Tingting Bi
- Wenchao Gu
- Zibin Zheng
og_published: "2025-12-01" Obsidian types both correctly — authors renders as a real list property,
og_published as a date.
Why one property was never going to be enough
The old code, in full:
const url = typeof fm.url === 'string' ? fm.url : null;
if (!url) { new Notice('Metafetch: no `url` field in frontmatter'); return; } The temptation is to fix this with a precedence list — try url, then arxiv,
then doi, then… But there's no correct order. A note can hold several URLs
that are all legitimately fetchable: an abs page and a PDF, a DOI and a
publisher landing page. Which one you meant is a judgement, and the person
making it is sitting right there.
So: show what's there, let them choose.
The interesting part turned out to be what to leave out. A note that's been
fetched before carries og_image and og_favicon — both URLs, neither
something anyone wants to fetch metadata from. Two exclusion mechanisms,
because one isn't enough: by configured field name (handles the defaults) and by
image extension (catches the same values when you've renamed those fields).
Array values get enumerated too — mirrors[0], mirrors[1] — since a paper
commonly has both an abs link and a PDF.
Three reasons arXiv came back empty
Not one bug. Three, stacked.
1. Wrong standard. arXiv publishes scholarly metadata as Highwire Press
citation_* tags. We only looked for article:author and name=author. Sitting
on that page unread: citation_author ×6, citation_date, citation_title,
citation_pdf_url, citation_arxiv_id.
2. First match only. getMeta() returned the first regex hit. arXiv emits
one tag per author:
<meta name="citation_author" content="Wang, Yanlin" />
<meta name="citation_author" content="Xu, Xinyi" />
<meta name="citation_author" content="Chen, Jiachi" /> So even after teaching it the right tag name, a six-author paper would have come
back with one author and five people silently dropped. getMetaAll() collects
repeated tags; getMeta() is now a one-line wrapper over it.
3. Wrong shape. arXiv writes Wang, Yanlin and 2025/12/01. cite-wide's
citation schema validates "FirstName LastName", and everything downstream
wants ISO dates. Both normalized on the way in — with a guard, so
Jane Doe, PhD doesn't become PhD Jane Doe.
Tiers are first-non-empty, not merged
Worth naming because the alternative looks reasonable until it isn't. A journal
page can carry a generic author tag naming one person and a complete
citation_author set. Merging them lists the same people twice under two
spellings. So the tiers are tried in order and the first non-empty one wins.
citation_* leads, because it's complete where the generic tags are lossy. Pages
without it fall straight through to what they do publish — which is why nothing
changed for the rest of the web.
One small defensive touch: article:author is frequently a Facebook profile URL
rather than a name, so URL-valued authors are dropped.
What checking first saved us from building
The plan for this called for a JSON-LD schema.org tier, on the reasoning that trade press would need it.
Before writing it, we pulled a live TechCrunch article and looked. It serves
<meta name="author"> and <meta property="article:published_time">, and
correctly reports og:type: article. It already worked. Verified through
the real code path — Lucas Ropek, 2026-08-17T21:27:05+00:00, no changes
required.
So the JSON-LD tier was dropped as speculative. The gap was never "articles",
it was scholarly sources specifically — trade press is well served by the
generic tags, and academic publishers are the ones on a separate standard. An
hour of building the wrong abstraction, avoided by one curl.
An identity code, so the note can be pointed at
A source note is only worth making if you can reference it later. And the
usual way — [[An Empirical Study of Agent Developer Practices]] — is more
fragile than it looks. It breaks when you rename the file, and it goes ambiguous
the moment two notes share a title.
So there's now an opt-in setting: stamp a short, vault-unique code onto notes as they're fetched.
hex_code: k4m2x9 That code is stable. Rename the file, move it between folders, roll the vault up into a published site — every reference still resolves.
It's off by default. Everything else metafetch writes is fetched — it came from the page. A code is different: it's a commitment the plugin invents on your behalf, and it isn't reversible in the way a re-fetch is. That should be a choice you make, not a surprise. Field name and length are configurable, like every other field metafetch writes.
These are not hexadecimal
The name is historical, and it's worth being explicit because the obvious
reading of "hex_code" is wrong. We deliberately do not limit the alphabet to
[0-9a-f]. Six characters of [a-z0-9]:
| Alphabet | Size | 6-character space |
hex [0-9a-f] | 16 | 16.7 million |
ours [a-z0-9] | 36 | 2.18 billion |
Identical six bytes on disk. Roughly 130× the room, and therefore ~130× less likely that two notes ever collide. There's no argument for paying the narrower alphabet's price when the wider one is free.
Generation uses rejection sampling rather than byte % 36. Modulo would make
the first four letters of the alphabet very slightly more likely than the rest,
because 256 doesn't divide evenly by 36. A tiny bias — and a silly one to accept
in an identifier whose entire purpose is not colliding.
Two rules the implementation actually enforces
Write-once. An existing code is never overwritten or regenerated. This is the whole ballgame: a code that changes on the next fetch is worse than no code, because every reference pointing at the old one breaks silently and you won't find out until much later.
Uniqueness is checked, not assumed. Every mint reads the codes already in use across the vault. 2.18 billion makes collisions remote, but "remote" and "checked" aren't the same thing.
There's a subtlety we got wrong on the first pass. Obsidian's metadata cache lags actual file writes — so during a batch run, a code written to a file seconds ago isn't visible in the cache yet, and the next file could be handed the same one. The first version of this shipped with a comment confidently claiming that re-reading per file handled it. It didn't. Batch runs now thread a run-scoped set alongside the cache lookup, verified against a 200-file simulated batch.
Under the hood
| Unit | Job |
frontmatterUrls.ts | Find fetchable URLs; exclude metafetch's own output |
SelectUrlModal.ts | The picker — built on Obsidian's native Setting |
getMetaAll() | Every value of a repeated meta tag |
normalizeAuthorName() | Last, First → First Last, credential-safe |
normalizeDate() | Highwire slashes → ISO, timestamps untouched |
hexCode.ts | Mint, collision-check, and stamp vault identity codes |
The modal is built from Obsidian's Setting component rather than custom
markup, and that's deliberate: while building it we found metafetch's esbuild
CSS entry point is src/styles/citations.css, which contains nothing but
.cite-wide-* classes. The plugin has been shipping a sibling plugin's
stylesheet, and its own .metafetch-modal / .opengraph-* classes are
defined nowhere. Native components are styled by the app regardless of that.
The CSS problem is real and gets its own fix.
The existing commands are untouched. runFetchScript takes an optional
explicit URL; without it, the fm.url path behaves exactly as before — only the
failure notice changed, to point at the picker.
Tests are up to 71 across four units, all wired into pnpm test: 16 for
frontmatter round-tripping, 14 for URL discovery, 22 for the scholarly helpers,
19 for identity codes — including a guard that the full 36-character alphabet
actually shows up in output, so we can't silently regress into real hex.
What's Next
The third phase is the source-note shape — and the schema for it already
exists. cite-wide's Lossless-Citation-Standards.md defines 24 fields across 12
publisher_type values, and that taxonomy already names both cases by hand:
Academic-Working-Paper lists arXiv, Industry-Media lists TechCrunch. No
vocabulary needs inventing; it needs implementing.
Two things deliberately left alone:
og_typestill sayswebsiteon arXiv. That's arXiv's literalog:typevalue. Overriding it would mean writing something the source didn't say. Inferring "working paper" ispublisher_type's job, in Phase 3.Multi-URL notes still share one set of
og_*fields, so fetching a second URL overwrites the first's. Prefixing per source property would fix it but breaks the configurable-field-name contract. Unresolved on purpose rather than settled by accident.
Full write-up, including the relationship to the open grab-reference decision:
context-v/specs/Metafetch-Source-Notes-From-Any-Journal-URL.md in the parent
content-farm repo.
Also shipped today: [[2026-08-17_01]] — the frontmatter quoting fix that's the
reason the authors list above renders as clean bare items instead of six
quoted strings.