literature-review-toolkit¶
Scripts that let an LLM agent run a literature review without fabricating references.
Version 2.0.0 · MIT license · AI disclosure
The agent decides what to search, what matters, how to group it, and optionally how to write it up. The scripts handle the API calls, verification and bookkeeping. That is the work LLMs do worst, and one invented author or wrong DOI can quietly corrupt a review.
A finished lineage figure. 190 verified papers on how the brain represents complexity, grouped into six theoretical families on a timeline from 1948 to 2026. Dot area shows citation count, and the labeled dots are landmarks the toolkit chose automatically. Every dot is a verified, canonically formatted reference in the accompanying spreadsheet.
Judgment versus ground truth¶
A review needs human judgment at only three points. Everything between them has a ground truth: a DOI either resolves to the cited paper or it does not. So the toolkit checks those steps by script instead of trusting the agent's memory.
1
Scope¶
You choose the topic and how far back to search, or, in lab mode, which lab's publications to start from. You also say how big a search you want.
2
Families and the timeline¶
The agent always offers the timeline and proposes the families it would use. You use them, change them, or skip the timeline.
3
The write-up (optional)¶
The agent writes the narrative review. This is the one judgment step the toolkit does not mechanize.
Between those points, every step runs automatically:
- Antecedents. A required second search finds the field's methodological, empirical and theoretical roots, which a forward search misses.
- Verification. Every citation is checked against PubMed, PMC, CrossRef, DataCite and arXiv. Without a duty to verify, search agents got about 1 in 4 citations wrong.
- Canonical references. Every reference is rebuilt from its verified DOI into APA-7. References with no DOI get a hand check.
- Summary checks. A checking agent with no web access compares every summary with the paper's abstract.
- Citation counts. Counts come from OpenAlex, checked against Semantic Scholar.
- Missed papers. The corpus's own reference lists, and the papers that cite its landmarks, surface papers the search missed.
- The audit. A failed check stops the build. The spreadsheet runs the full audit and refuses a table that fails it.
What you get¶
:material-table: Spreadsheet¶
The core deliverable. One row per paper: canonical reference, summary, tag, family and citation counts, colored by where the paper came from.
:material-chart-timeline-variant: Lineage timeline¶
Offered on every review. An interactive HTML figure that lays out the families
on a timeline, with an SVG copy, plus PNG and PDF when rsvg-convert or
Inkscape is installed. It is often the most useful thing a review produces.
:material-file-document-edit: Review article (optional)¶
An AI-authored narrative review (.docx and a web page), with its references
drawn from the verified corpus.
Two modes¶
Start from a question and search outward.
"Literature review on the anatomical connections between the visual system and the cerebellum — primate or human, any tractography method, back to the 1970s."
Start from a lab's publications, derive its research themes, then search outward to place that work in the field.
"Review the Gallant lab's human-imaging work in the context of the broader field."
From verification on, both modes run the same pipeline. See Choosing a front end.
Get started¶
You need Python 3, a contact email, and two free API keys. Without an OpenAlex key, every client on one IP address shares a single daily budget, and a campus network can spend it before you start. Keyless Semantic Scholar is heavily throttled. In the directory that will hold your reviews:
git clone https://github.com/gallantlab/literature-review-toolkit.git
cd literature-review-toolkit
pip install -r requirements.txt
export LITREVIEW_EMAIL=you@institution.edu # NCBI and CrossRef require one
export OPENALEX_API_KEY=... # https://help.openalex.org/api/authentication
export S2_API_KEY=... # https://www.semanticscholar.org/product/api#api-key-form
Then open Claude Code in the directory that holds the clone and describe the review. Say how big a search you want, in your own words: "a quick look at the key papers", "the core literature", "about 300 papers" or "everything". Say nothing and you get the standard search. Every check runs at every size. The scales are listed in Topic mode.
Before it searches, the agent runs the
preflight. It
checks for a newer toolkit, your keys and today's OpenAlex budget, and asks you
to choose when access is short. The agent then follows
PLAYBOOK.md.
A full build takes hours. Installation and
environment variables are covered in the
manual.
Next steps¶
- Operator manual: install, rules, every phase with its command and gate, how to read the outputs, and troubleshooting.
- Examples: a finished review in each mode and a gallery of lineage figures.
AI disclosure¶
The toolkit was built with AI. Most of its code and documentation were written by Claude, Anthropic's AI model, under the direction of Jack Gallant (Gallant Lab, UC Berkeley).
What it produces is AI-generated too. An LLM agent runs the searches, writes each paper's summary from its abstract (not the full text), proposes the family groupings, and writes the review articles. Every review article states its AI authorship.
Only facts are machine-checked. The scripts verify every citation against the literature databases and rebuild every reference from its DOI. Summaries, groupings and interpretation cannot be checked that way; read them as an AI's reading of the abstracts.
