What this tool does
Citracker pulls a researcher's complete publication record — including the full reference list of every paper — from OpenAlex, then measures four families of signal that the research-integrity literature associates with citation manipulation. Everything runs in your browser; no account, no API key, no server.
1 Self-citation
What share of the citations to their work comes from their own papers. Reported both as a share of citations received (the measure that tells you whether an h-index is inflated) and a share of citations given (the measure of citing habit). Benchmarked against the ~12.7% population median from Ioannidis et al., PLoS Biology 2019.
2 Citation rings
Closed, reciprocal citation loops — "citation cartels". The tool does not just count who someone cites; it fetches the reference lists of their 10 closest collaborators and checks whether the citations come back. Heavy one-way citation of a senior figure is ordinary academia. Heavy, balanced, two-way citation inside a small group is not.
3 Venue funnelling
Citations steered into one journal or conference — the pattern behind Impact Factor inflation. The sharp signal is when the cited papers were published: Impact Factor counts only the previous two years, so citations clustered in that window, in a venue the author also publishes in or edits, are the classic signature.
4 Citation velocity
Year-over-year jumps in citations received that deviate from the author's own growth trend by more than a modified z-score of 3.5. The tool names the specific papers that drove each spike, because that usually settles the question immediately.
Read this before you use the output
None of these signals prove misconduct, and this tool cannot detect fraud. It measures statistical patterns that are consistent with manipulation and equally consistent with legitimate behaviour. Every metric here has innocent explanations: a long-running research programme legitimately builds on its own prior work; a small subfield contains few people and few venues, so its citation graph looks closed; a paper can genuinely go viral.
It also cannot read citation context — it has no idea whether a citation is substantive, perfunctory, or critical — and it inherits every error in OpenAlex's author-name disambiguation, which is imperfect. A high score means "a human who knows this field should look at the details", never "this person did something wrong". Do not publish, share, or act on a score as if it were a finding.
1 Set up API access
Data source
Which database the analysis runs on. They are not equivalent — the choice changes what can be measured at all, not just how fast it runs.
Since February 2026 OpenAlex uses usage-based pricing. Requests without a key get about $0.10/day of budget — not enough to complete even one analysis. A free key raises that to $1/day, roughly five full analyses. Getting one takes about 30 seconds and costs nothing.
No key set. The key is stored only in this browser's local storage and is sent only to the OpenAlex API.
Running on Semantic Scholar. No key and no budget needed — analyses are free and unlimited, and references arrive with their authors and venue already attached, so nothing has to be resolved. Two things to know: citation velocity cannot be measured at all (Semantic Scholar publishes no per-year citation counts), and publishers withhold reference lists for many paywalled papers, so citation-direction coverage is often much thinner than OpenAlex's. The report states the coverage it achieved and warns loudly if it was too low to trust.
Analysis depth
Complete is the default and applies no sampling: every publication, every distinct cited work, and all 10 closest collaborators checked against their full reference lists. Its cost depends on how much the researcher cites, so it is measured and shown in the progress bar before it is spent. The other presets exist only to fit a smaller budget. If the budget runs out mid-run you still get a report, clearly marked as partial.
Cluster analysis
The checks above look at one person. These look at the group — the citation network among the target and the collaborators probed above — because a citation cartel is a property of a subgraph, not of any single pair. Each tier switches independently.
2 Find the researcher
3 Analysis
Self-citation rate
Reciprocal citation
Venue concentration
Citation velocity
Collaborator citation exchange
What this shows: for each of the researcher's closest collaborators, how many
times this researcher cited them ("given") set against how many times they cited this
researcher back. One row per collaborator, ordered by how much of the exchange
co-authorship fails to explain.
What to look for: the two directions together, not either alone. Heavy one-way
citation of a senior figure is ordinary academia; a large, balanced exchange inside a
small group is the citation-cartel pattern.
| Collaborator | Shared papers | Citations given | Citations back | Balance | Assessment |
|---|
Where their citations go
What this shows: the journals and conferences this researcher's own reference
lists point at, ranked by volume — i.e. where they direct the credit they hand out.
What to look for: a single venue taking a large share, especially in the last
column. Impact Factor counts only citations to papers from the previous two years, so citations
bunched into that window are the signature of Impact-Factor inflation.
| Venue cited | Citations | Share | Within 2-yr Impact Factor window |
|---|
Where they publish
What this shows: the venues this researcher publishes their own work in, with
how many papers and how many citations each has earned them.
What to look for: a venue appearing prominently here and in the table
above. Steering citations into a journal is only interesting when the researcher has a stake in
that journal — that overlap is what turns a statistic into a motive.
| Venue | Papers | Share | Citations received |
|---|
A venue appearing prominently in both tables is the situation worth a look — it means the author has a stake in the venue they are steering citations toward.
Their most self-cited papers
What this shows: which of the researcher's own papers they cite most often in
their later work.
What to look for: whether self-citation is concentrated or diffuse. Repeatedly
citing one genuine flagship result is normal; self-citation spread thinly across dozens of
papers is the pattern that inflates an h-index.
| Paper | Year | Times self-cited |
|---|
Citations received per year
What this shows: how many citations the researcher's whole body of work
received in each calendar year, summed across every paper.
What to look for: a year that breaks sharply from their own trend (shown in
red). Note the weakness of this signal in isolation — a genuinely viral paper produces exactly
the same shape as manipulation, which is why the breakdown underneath names the papers
responsible.
The group around this researcher
The citation network, drawn
What this shows: every member of the group as a circle, with an arrow
for each direction citations flow. Circles are sized by how many citations a member
receives from inside the group; lines are thickened by how many citations they
carry. Two people who cite each other appear as two curved arrows bowing apart, so you
can see whether the traffic is balanced or one-way. Position is meaningful: the layout
pulls heavily-citing members together, so a tight subgroup clumps visually.
What to look for: a knot of thick two-way arrows between people who are
not linked by a dashed supervisory line — that is exchange without an obvious
innocent explanation.
Who cites whom
What this shows: the complete citation traffic inside the group. Read a
row as "this person cites…" and a column as "…this person". Each cell is the
number of citations sent in that direction; darker means more. The diagonal is blank because
citing yourself is measured separately, further up the report. ★ marks the researcher the
report is about.
What to look for: dark cells that mirror each other across the diagonal —
that is citation flowing both ways between the same two people. Hovering any cell
highlights its border and the border of the opposite cell — the same pair of people,
the other direction — and prints both counts with their balance underneath, so you can
read a relationship without scanning across the grid.
Pairs exceeding a degree-preserving null model
What this shows: every pair in the group that exchanged enough citations to
test, with how many citations actually passed between them ("observed") against how many
chance alone predicts ("expected"). "Lift" is the ratio.
What to look for: high lift with few shared papers. The expectation
already accounts for both people being prolific and well cited, so a high lift means something
beyond volume is going on — but co-authors citing each other is simply what collaboration is,
which is why the discounted figure is the one that matters.
| From → To | Observed | Expected | Lift | Shared papers | q-value | Verdict |
|---|
Structure
What this shows: four shape measurements of the group's citation network —
how much of its citing stays internal, the largest set of people who all cite each other, the
number of A→B→C→A loops, and how often citation runs in both directions.
What to look for: several of these elevated at once. Any one of them alone is
also what a small, legitimate, tight-knit specialty looks like.
Where each member's citations come from
What this shows: for each member, what share of the citations they
receive comes from inside this group rather than from the wider literature. This is
the mirror of every other measure in the report, which look at citations a person gives.
What to look for: a high share. That is what makes a citation count
self-sustaining — a researcher with genuine reach is cited by strangers.
| Member | Citations from inside the group | Total received | Share from inside |
|---|
This is the mirror of everything else in Citracker, which otherwise only looks at citations a person gives. A high share here is what would make an h-index self-sustaining: a legitimate researcher is cited by strangers.
Behaviour
Neighbourhood
Coverage of this analysis
What this shows: how much of the record the numbers above actually rest on —
how many works were read, how many had reference lists, how many citations could be traced to a
specific author and venue, and what the run cost.
Why it matters: every percentage in this report is computed over the attributed
subset. If attribution is well below 100%, treat the ring and venue shares as approximate.
Export this report
Every number above traces to a count in the JSON. Export it if you are handing the case to someone else — a score without its underlying counts is not reviewable.
The PDF button opens your browser's print dialog — choose “Save as PDF” as the destination. The page restyles itself for paper: light background, expanded tables, page breaks between sections, and the caveats carried onto the front page. Interactive controls and the raw JSON are omitted.
Try these examples
Two calibration groups. The controls should score low; the documented cases are researchers whose extreme self-citation has been reported in the peer-reviewed bibliometrics literature. Those reports are cited below — they are not claims originating from this tool, and extreme self-citation is not in itself misconduct.
Controls — expected to score low
Documented high self-citation
Source: Ioannidis, Baas, Klavans & Boyack, "A standardized citation metrics author database annotated for scientific field", PLoS Biology 17(8), 2019, and the associated public dataset, which reports per-author self-citation rates. See also Nature's coverage, "Hundreds of extreme self-citing scientists revealed in new database" (2019).