Who Cited It

Methodology

What is in this corpus, where every figure came from, and the three things this site refuses to do.

Sources

SourceLicenceUsed for
OpenAlexCC0works, authors, institutions, topics and citation edges

Nothing here is scraped from a publisher's website. Every record comes from an open API or bulk dump under a licence that permits reuse, and each page names the stored payload it was rendered from.

Collaboration weight is not a count of shared papers

Two people who wrote a two-author paper together last year worked together. Two people who appear in the middle of a 900-author collaboration paper in 1998 probably did not. Ranking collaborators by shared-paper count treats those as the same fact, so this site weights each shared work by three things:

This is a claim about what collaboration means, not a measurement. It is stated here so you can disagree with it.

Author identity is inferred, and we show our work

Author records in every open bibliographic database are produced by a disambiguation algorithm. It splits one researcher across several records and merges several researchers into one. Tools built on top of this usually render the output as fact.

Instead, every author page carries a confidence band and the signals behind it: whether a human-claimed ORCID is present, how many distinct name forms appear, how many institutions show up inside any five-year window, and how much of the record sits in one field. No record is ever merged, split or dropped here. A record we doubt says so and stays visible.

Currently 3,643 high, 890 medium and 2,373 low confidence.

Paper records get the same treatment

A work record can contradict itself, and in this corpus a surprising number do. We check four things that need no second source to verify: whether the record lists any authors at all, whether it records references behind a large citation count, whether the year embedded in its DOI agrees with its own publication year, and whether it has a title.

2,089 complete · 810 partial · 101 suspect.

Deliberately not checked: whether a citation count is implausible. This corpus is selected by citation count, so every work in it is an outlier against the wider population, and a threshold fitted here would be fitted to the selection rather than to the data. Calling a number wrong needs evidence this tier does not carry.

Nothing is deleted or corrected. A record flagged suspect stays on the site with its figures intact and a panel saying what is wrong with it — silently fixing a source is how a reader ends up trusting a number nobody can trace.

Why some abstracts are missing

An abstract present in a source record is not permission to republish it. Publishers have had abstracts removed from open indexes in bulk — Springer Nature in 2022 and Elsevier in 2024 — and the ones that remain carry the licence of the copy they came from. This site renders an abstract only when the work has an open-access copy under a licence that permits redistribution, and says which of the two reasons applies when it does not.

236 rendered · 1,747 withheld for licence · 1,017 absent from the source.

What this site will not do