A literature review pipeline

Feed The Saurus your papers.

From a folder of PDFs to a literature review, with every claim traced to its page.

Upload the corpus you already curated. A multi-agent pipeline analyzes each paper, extracts themes and claims, deduplicates themes across papers, and writes a cohesive review where every citation resolves to paper, page, and paragraph. You watch every stage. You keep the judgment.

Open source · Runs locally · 20 to 100 PDFs per job
A baby brontosaurus eating PDF papers, with a finished Review book floating nearby. Below: INPUT, DIGESTION, SYNTHESIS. fig. 1 — digestion, schematic
The bookkeeping barrier

Literature reviews take weeks because of bookkeeping, not reading.

You have already done the hard part. The folder has forty papers in it, chosen by you, for a reason. What remains is the part nobody enjoys: reading each one, noting which themes it touches, tracking which papers agree and which do not, finding the page and the paragraph for every claim, and then writing prose that holds all of it together without losing a single citation.

That is not thinking. It is accounting. It fails in three predictable ways.

The Context Cliff.

Working memory collapses around paper twelve. By paper thirty-eight you are no longer synthesizing a corpus; you are surviving a reading list, and the last paper is read in isolation from the first.

The Lexical Disconnect.

Authors describing the same mechanism use different words for it. One paper says "circadian disruption," another "chronobiological misalignment." A keyword search treats them as two unrelated inquiries, and so does a tired reader.

The Provenance Drift.

Notes detach from page numbers. Spreadsheets lose paragraph locations. By the time you write, an assertion has no verifiable address, and finding it again costs an afternoon.

Corpus, not document

Synthesis by theme, not summary by paper.

Document chatbots operate on one paper at a time. Handed a folder, they produce a stack of isolated summaries: paper A studied this, paper B studied that. That is not a literature review. A review is a horizontal cross-section of a field, organized by the concepts running through the literature, not by the bibliographies of the authors who wrote it.

The Summary Stack

Forty papers processed as forty silos.

Paper 01 · summary of methods, summary of results
Paper 02 · summary of methods, summary of results
Paper 03 · summary of methods, summary of results
… 37 more summaries

A list. No cross-paper agreement, no tension, no synthesis.

The Theme Matrix

One theme, every paper that touches it.

Circadian misalignment

Roenneberg [4] · Kalsbeek [11] · Vetter [23]

Contrasts persistent melatonin delay with partial re-entrainment between rotations.

Sleep debt and metabolic markers

Scheer [2] · Buxton [9] · Morris [17]

Maps agreement on glucose tolerance, disagreement on cortisol timing.

A draft woven around ideas, with the sources still attached.

Feed

Drop your PDFs. The Saurus accepts the corpus as it is: no reformatting, no reference manager export, no per-paper setup.

Digest

Each paper is analyzed independently and in parallel. Themes and claims are extracted, then themes are matched across the corpus by semantic similarity so that two papers naming the same idea differently end up in the same place.

Review

Each theme is reviewed against every paper that touches it. The pipeline then writes a cohesive literature review, structured by theme, with inline citations resolved to paper, page, and paragraph.

Explore

Browse per-paper findings, inspect the full trace of every stage, and ask the assistant questions about the corpus. It answers with the same citation discipline as the review.

Paper Analysis → Theme Dedup → Theme Review → Aggregation

Shared vocabulary does not settle a disagreement. The claims and their sources remain there for you to compare.

Any tool can summarize a paper. Here is what happens after that.

Four things The Saurus does that a chat window over your PDFs does not. Each one is shown, not asserted.

Stage 2 · Theme Dedup

Part thesaurus. It connects what authors said differently.

A thesaurus groups words that share a meaning. The Saurus groups themes that share a phenomenon.

Extraction is easy. The hard part is recognizing that paper 4 calls it "circadian disruption," paper 11 calls it "chronobiological misalignment," and paper 23 calls it "shift-work phase shift" while spending three pages on the same thing. The dedup stage matches themes across the corpus by embedding similarity and merges them, so the review is organized around what the literature discusses, not around the order you uploaded the files.

Themes are matched by embedding similarity, then confirmed pairwise before merging. Members that fail the check are kept separate.

Raw theme labels · extracted from 40 papers
[4]Roenneberg et al. 2023circadian disruption
[11]Kalsbeek et al. 2024chronobiological misalignment
[23]Vetter et al. 2022shift-work phase shift
Theme Dedup · embedding match → pairwise confirm
Merged theme · in the review
Theme 03

Circadian misalignment

3 papers · 11 claims · 1 observed tension

Merged from: circadian disruption · chronobiological misalignment · shift-work phase shift
Stage 3 · Citation

Every citation has an address.

Not a paper title. Not "according to the literature." A paper, a page, and a paragraph. Every claim in the review resolves to a location you can open and read, which means every sentence is one you can defend in a committee meeting, or delete because you disagree with it.

Generated review · §2.3 Recovery dynamics

Night-shift schedules are consistently associated with delayed melatonin onset relative to day workers [4](p.7,§2), an effect that persists for at least three consecutive rest days [11](p.12,§3). Vetter and colleagues report a smaller but measurable shift in rotating-shift cohorts [23](p.4,§1), which they attribute to partial re-entrainment between rotations [23](p.5,§2).

Grounded citation
[11] Kalsbeek et al. 2024 · p.12 · §3
"Dim-light melatonin onset remained delayed by a mean of 2.1 h on the third consecutive rest day, indicating that re-entrainment to a day-active schedule was incomplete within the recovery window studied."
Entailment status
Entailed
NLI score
0.94
Claim ID
kalsbeek-2024-c07
Cited in
§2.3, §4.1
Checked by NLI cross-encoder · citation guard passed

[n] resolves to the reference entry. (p.12,§3) resolves to the paragraph. Both are checked before the review is returned.

§2.3 · 142 claims verified · 0 ungrounded · 3 escalated to reask · specimen figures

Stage 4 · Grounding

The model asserts. Something else checks.

Language models produce citations that read correctly and are not. The Saurus does not ask you to trust that this did not happen. It checks, mechanically, at two points, and sends failures back rather than through.

  1. Entailment.

    NLI cross-encoder

    A natural-language-inference cross-encoder reads each cited claim against the cited passage and scores whether the passage actually entails the claim. It has one job, and it is not the model that wrote the review.

  2. Citation integrity.

    citation guard

    A guard verifies that every [n](p.x,§y) in the review resolves to a real claim extracted in stage one. A citation to a claim that does not exist is rejected before you see it.

  3. Reask, not shrug.

    reask

    When entailment is borderline or a citation fails to resolve, the offending passage is sent back to the writing agent with the failure attached. Borderline cases are escalated to a stricter check rather than waved through. Failures that survive are shown in the trace, not smoothed over.

    ↶ writing agent

⟳ Theme Review · batch 3 · reask · entailment 0.41 < 0.60 · retry 1/2

Illustrative trace line · threshold shown as specified in the brief.

This reduces fabrication risk by mechanism. The mechanism is the message.

The check tests support in a passage. It does not establish that a study is sound or that your corpus covers the field. You still assess the evidence and the draft's interpretation.

Stage 5 · Trace

Watch every stage. Nothing is hidden.

A single prompt over a context window is not a process you can inspect, resume, or reproduce. The Saurus runs as a durable pipeline: each paper analyzed independently, each stage recorded as it happens, each agent's events visible in real time. If a paper cannot be parsed, the review says so and continues. If the job crashes, it resumes from the last completed stage, and the trace shows the resume honestly.

Every event is persisted as it occurs. The trace you see during the run is the trace you can open afterward.

JOB circadian-2026-09 · Recorded sample run

Corpus: Circadian Misalignment in Shift Work (40 papers · 380 pages)

39/40
Parsed
187
Themes raw
41
Themes merged
312
Claims traced
Ingestion 40/40 converted12.4s
Paper Analysis 39/40 · 312 claims · 187 themes4m 02s
[7] Sørensen et al. · could not be parsed · skipped
Theme Dedup 187 → 41 themes38s
Theme Review batch 6/9 · 5 themes per batchrunning
batch 3 · reask · entailment 0.41 < 0.60 · retry 1/2
Aggregation pending

resumed from journal

Event log · append-only NDJSON
 paper_analysis/[7]   parse failed · skipped · review continues
 theme_dedup          merged "circadian disruption" + "chronobiological misalignment" → "Circadian misalignment"
 theme_review/batch-3 reask · entailment 0.41 < 0.60 · retry 1/2
 theme_review/batch-3 retry passed · entailment 0.87
 workflow             resumed from journal · 3 stages replayed, 0 re-executed

Downstream of the tools you already use.

Use Semantic Scholar or Research Rabbit to find the papers. Use Elicit or Consensus to pull claims from them. Use a writing assistant to polish your own argument. The Saurus is the step in between: it takes the folder and produces the review those tools do not write.

A place in your existing research workflow.
Your toolsThey give youThe Saurus gives you
Discovery toolsPapers to readSynthesis of papers you already have
Extraction toolsClaims per paper, side by sideThemes deduplicated across papers, written into a review
General LLMsAn answer, unverifiedA review, traced to page and paragraph, grounding-checked
Writing assistantsFaster prose for your argumentThe literature-review draft your argument sits on
Manual reviewTotal control, days of workSame control, minutes of bookkeeping
The division of labor

A draft to argue with, not a verdict to accept.

The Saurus produces a draft and the evidence behind it. It does not rank your papers by importance, does not decide what the field has concluded, and does not claim to have understood anything. It produces the map. Reading is still yours, and so is the argument.

The review is editable and exportable. The per-paper findings are browsable. And the assistant, when you ask it something, answers with citations rather than verdicts.

Export · Markdown / PDF

What The Saurus does

  • Extracts themes and claims in parallel across your PDFs
  • Reconciles diverging author terminology through semantic deduplication
  • Verifies every cited claim against its source passage with an NLI check
  • Delivers an editable review with citations resolved to page and paragraph
  • Keeps a persistent, inspectable trace for every stage

What The Saurus will not do

  • Assert that a review is complete without your reading it
  • Fabricate a citation or smooth over a broken reference
  • Claim to have understood the implications of your field
  • Produce your thesis contribution or original argument
  • Train on your unreleased manuscripts

Keep asking questions of the corpus.

Answers cite the same claim set as the review. If the corpus does not contain it, the assistant says so.

You

Which papers disagree about whether the melatonin delay recovers within a week?

Assistant

Two positions appear in the corpus. Kalsbeek et al. report persistence beyond three rest days [11](p.12,§3), while Vetter et al. observe partial re-entrainment between rotations [23](p.5,§2). Roenneberg et al. do not address recovery timing directly [4]. The corpus does not contain a study extending beyond seven days, so it cannot settle the question as posed. The review discusses this under §2.3, Recovery dynamics.

The Saurus does the digestion. You decide what the literature supports.

Feed The Saurus your papers.

Forty papers in. One review out. Every claim traced to a paragraph.

Hungry for papers since the Cretaceous!