Feed The Saurus your papers.
From a folder of PDFs to a literature review, with every claim traced to its page.
Upload the corpus you already curated. A multi-agent pipeline analyzes each paper, extracts themes and claims, deduplicates themes across papers, and writes a cohesive review where every citation resolves to paper, page, and paragraph. You watch every stage. You keep the judgment.
fig. 1 — digestion, schematic
Literature reviews take weeks because of bookkeeping, not reading.
You have already done the hard part. The folder has forty papers in it, chosen by you, for a reason. What remains is the part nobody enjoys: reading each one, noting which themes it touches, tracking which papers agree and which do not, finding the page and the paragraph for every claim, and then writing prose that holds all of it together without losing a single citation.
That is not thinking. It is accounting. It fails in three predictable ways.
The Context Cliff.
Working memory collapses around paper twelve. By paper thirty-eight you are no longer synthesizing a corpus; you are surviving a reading list, and the last paper is read in isolation from the first.
The Lexical Disconnect.
Authors describing the same mechanism use different words for it. One paper says "circadian disruption," another "chronobiological misalignment." A keyword search treats them as two unrelated inquiries, and so does a tired reader.
The Provenance Drift.
Notes detach from page numbers. Spreadsheets lose paragraph locations. By the time you write, an assertion has no verifiable address, and finding it again costs an afternoon.
Synthesis by theme, not summary by paper.
Document chatbots operate on one paper at a time. Handed a folder, they produce a stack of isolated summaries: paper A studied this, paper B studied that. That is not a literature review. A review is a horizontal cross-section of a field, organized by the concepts running through the literature, not by the bibliographies of the authors who wrote it.
Forty papers processed as forty silos.
A list. No cross-paper agreement, no tension, no synthesis.
One theme, every paper that touches it.
Circadian misalignment
Roenneberg [4] · Kalsbeek [11] · Vetter [23]
Contrasts persistent melatonin delay with partial re-entrainment between rotations.
Sleep debt and metabolic markers
Scheer [2] · Buxton [9] · Morris [17]
Maps agreement on glucose tolerance, disagreement on cortisol timing.
A draft woven around ideas, with the sources still attached.
Feed
Drop your PDFs. The Saurus accepts the corpus as it is: no reformatting, no reference manager export, no per-paper setup.
Digest
Each paper is analyzed independently and in parallel. Themes and claims are extracted, then themes are matched across the corpus by semantic similarity so that two papers naming the same idea differently end up in the same place.
Review
Each theme is reviewed against every paper that touches it. The pipeline then writes a cohesive literature review, structured by theme, with inline citations resolved to paper, page, and paragraph.
Explore
Browse per-paper findings, inspect the full trace of every stage, and ask the assistant questions about the corpus. It answers with the same citation discipline as the review.
Shared vocabulary does not settle a disagreement. The claims and their sources remain there for you to compare.
Any tool can summarize a paper. Here is what happens after that.
Four things The Saurus does that a chat window over your PDFs does not. Each one is shown, not asserted.
Part thesaurus. It connects what authors said differently.
A thesaurus groups words that share a meaning. The Saurus groups themes that share a phenomenon.
Extraction is easy. The hard part is recognizing that paper 4 calls it "circadian disruption," paper 11 calls it "chronobiological misalignment," and paper 23 calls it "shift-work phase shift" while spending three pages on the same thing. The dedup stage matches themes across the corpus by embedding similarity and merges them, so the review is organized around what the literature discusses, not around the order you uploaded the files.
Themes are matched by embedding similarity, then confirmed pairwise before merging. Members that fail the check are kept separate.
Circadian misalignment
Every citation has an address.
Not a paper title. Not "according to the literature." A paper, a page, and a paragraph. Every claim in the review resolves to a location you can open and read, which means every sentence is one you can defend in a committee meeting, or delete because you disagree with it.
Night-shift schedules are consistently associated with delayed melatonin onset relative to day workers [4](p.7,§2), an effect that persists for at least three consecutive rest days [11](p.12,§3). Vetter and colleagues report a smaller but measurable shift in rotating-shift cohorts [23](p.4,§1), which they attribute to partial re-entrainment between rotations [23](p.5,§2).
"Dim-light melatonin onset remained delayed by a mean of 2.1 h on the third consecutive rest day, indicating that re-entrainment to a day-active schedule was incomplete within the recovery window studied."
- Entailment status
- Entailed
- NLI score
- 0.94
- Claim ID
- kalsbeek-2024-c07
- Cited in
- §2.3, §4.1
[n] resolves to the reference entry. (p.12,§3) resolves to the paragraph. Both are checked before the review is returned.
§2.3 · 142 claims verified · 0 ungrounded · 3 escalated to reask · specimen figures
The model asserts. Something else checks.
Language models produce citations that read correctly and are not. The Saurus does not ask you to trust that this did not happen. It checks, mechanically, at two points, and sends failures back rather than through.
-
Entailment.
NLI cross-encoder
A natural-language-inference cross-encoder reads each cited claim against the cited passage and scores whether the passage actually entails the claim. It has one job, and it is not the model that wrote the review.
-
Citation integrity.
citation guard
A guard verifies that every [n](p.x,§y) in the review resolves to a real claim extracted in stage one. A citation to a claim that does not exist is rejected before you see it.
-
Reask, not shrug.
reask
When entailment is borderline or a citation fails to resolve, the offending passage is sent back to the writing agent with the failure attached. Borderline cases are escalated to a stricter check rather than waved through. Failures that survive are shown in the trace, not smoothed over.
↶ writing agent
⟳ Theme Review · batch 3 · reask · entailment 0.41 < 0.60 · retry 1/2
Illustrative trace line · threshold shown as specified in the brief.
This reduces fabrication risk by mechanism. The mechanism is the message.
The check tests support in a passage. It does not establish that a study is sound or that your corpus covers the field. You still assess the evidence and the draft's interpretation.
Watch every stage. Nothing is hidden.
A single prompt over a context window is not a process you can inspect, resume, or reproduce. The Saurus runs as a durable pipeline: each paper analyzed independently, each stage recorded as it happens, each agent's events visible in real time. If a paper cannot be parsed, the review says so and continues. If the job crashes, it resumes from the last completed stage, and the trace shows the resume honestly.
Every event is persisted as it occurs. The trace you see during the run is the trace you can open afterward.
Corpus: Circadian Misalignment in Shift Work (40 papers · 380 pages)
resumed from journal
paper_analysis/[7] parse failed · skipped · review continues theme_dedup merged "circadian disruption" + "chronobiological misalignment" → "Circadian misalignment" theme_review/batch-3 reask · entailment 0.41 < 0.60 · retry 1/2 theme_review/batch-3 retry passed · entailment 0.87 workflow resumed from journal · 3 stages replayed, 0 re-executed
Downstream of the tools you already use.
Use Semantic Scholar or Research Rabbit to find the papers. Use Elicit or Consensus to pull claims from them. Use a writing assistant to polish your own argument. The Saurus is the step in between: it takes the folder and produces the review those tools do not write.
| Your tools | They give you | The Saurus gives you |
|---|---|---|
| Discovery tools | Papers to read | Synthesis of papers you already have |
| Extraction tools | Claims per paper, side by side | Themes deduplicated across papers, written into a review |
| General LLMs | An answer, unverified | A review, traced to page and paragraph, grounding-checked |
| Writing assistants | Faster prose for your argument | The literature-review draft your argument sits on |
| Manual review | Total control, days of work | Same control, minutes of bookkeeping |
Feed The Saurus your papers.
Forty papers in. One review out. Every claim traced to a paragraph.
Or inspect a finished review right now:
Preview specimens only. Finished recordings are not bundled with this page.
Hungry for papers since the Cretaceous!