Book research fails at one specific point, and it is not collection. Research arrives organised by source; writing needs it organised by argument. The system's entire job is translating between those two shapes, and that translation is called extraction.

Category
Writing
Author

Justkay
Documentary Filmmaker & Founder at Storyflow
Topics
2026-09-06
•
19 min read
•
WritingBook research fails at one specific point, and it is not collection. Research arrives organised by source, a book, an interview, an archive folder, and writing needs it organised by argument, chapter by chapter. The system's entire job is translating between those two shapes, and that translation is called extraction: turning each source into individual claims written in your own words, each carrying its citation and a note about where it might belong. Writers who skip extraction end up with four hundred highlighted PDFs and no book, because a highlight is a decision deferred rather than a decision made. Zotero handles citations, Obsidian or Notion holds the claims, and Scrivener holds the draft.
Full disclosure: Storyflow is our product. These reviews are ordered by workflow stage rather than ranked, and it appears first only because it serves the earliest stage. It has no citation management, no bibliography generation, no PDF annotation, no reference metadata and no filtered database views over hundreds of claims. For a book with a bibliography Zotero is required and this does not replace it; the article recommends Zotero plus Obsidian or Notion, with Scrivener for the draft.
Four states per source, and the count in one of them predicts whether the book gets written.
| Tool | Best For | AI Features | Price |
|---|---|---|---|
| Captured | In the library with metadata | A reading list, harmless | State 1 |
| Read | You know what is in it | The dangerous pile | State 2 |
| Extracted | Claims in your own words | With page numbers | State 3 |
| Placed | Assigned to a probable chapter | Even a wrong guess | State 4 |
Ask a writer eighteen months into a book what went wrong and it is rarely that they did not gather enough. It is that they gathered a great deal and cannot use it.
The recognisable state: four hundred PDFs with yellow highlighting, thirty hours of interview audio, a folder of photographs, and a growing sense that starting to write means first re-reading everything. Which it does, because nothing in that pile is in a form the writing can consume.
The structural reason is a shape mismatch. Research arrives organised by source: this book, this interview, this archive box. Writing consumes material organised by argument: this claim supports chapter four, this anecdote opens chapter six.
Nothing converts between those two shapes automatically. A folder structure does not do it, tags do not do it, and no AI tool currently does it reliably, because the conversion requires deciding what each piece of material means to your argument, which is the writing itself in miniature.
The system below exists to make that conversion a small daily task rather than a wall you hit in year two.
Every source is in exactly one of four states, and knowing the counts is the health check.
Captured. It exists in your library with enough metadata to find and cite it. Nothing has been read.
Read. You have been through it and know roughly what is in it.
Extracted. Its useful content has been turned into individual claims in your own words, each with a page or timestamp reference.
Placed. Those claims have been assigned to a probable chapter or section.
The number that predicts whether a book will stall is the count of sources sitting in Read but not Extracted. Captured-but-unread is fine, that is a reading list. Read-but-unextracted is the dangerous pile, because you have spent the time and retained none of the value in a usable form, and the memory of having read something decays much faster than writers expect.
Keep extraction within about a week of reading. Beyond that you are re-reading.
Extraction is turning a source into atomic claims. One idea per note, written in your own words, with the citation attached.
Why your own words and not the quote. A highlighted passage records that something struck you and not why. Eighteen months later you have the sentence and not the thought, and you have to reconstruct your reaction by re-reading the surrounding pages. A paraphrase forces the thinking at the moment you are best placed to do it, which is immediately after reading.
Store both: your claim as the note, the verbatim quote as a field on it. The quote is for the page; the claim is for the system.
What an extracted claim looks like:
A book-length project typically produces somewhere between three hundred and two thousand of these. That sounds enormous and it accumulates at four or five a day without effort.
The disagreement note deserves emphasis. When a source says something you doubt, writing down why is how arguments get built. Writers who record only what they agree with produce books that are summaries of their reading, and the doubt is where the original contribution lives.
This is the structural decision that determines whether the system works, and most writers get it wrong initially.
Source-shaped notes mirror your library: one document per book, holding everything from it. This feels natural, matches how the material arrived, and is unusable for writing, because chapter four needs three sentences from six different books and there is no way to gather them.
Claim-shaped notes are one document per idea. Each carries its source rather than living inside it. Now a chapter is a filter or a collection of claims, and assembling one is a query rather than an archaeological dig.
A single source contributes to five different arguments, which is exactly why it cannot live in one place. That is the whole case, and it is the same insight behind Zettelkasten, though the full method is more apparatus than most book projects need.
The practical minimum: one note per claim, a source field, a chapter field, and a tag or two. That is enough. Elaborate note systems consume the time the writing needs.
Once claims exist, the chapter structure becomes a view over them rather than a separate document.
Assign a probable chapter at extraction time, even badly. A wrong guess is trivially fixable and an unassigned claim is invisible when you write. The point is not accuracy, it is that everything has a home.
Then read the chapters as lists. A chapter with four claims is thin and you know in month three rather than month fourteen. A chapter with ninety is really two chapters. This is the earliest reliable signal about structure available, and it costs nothing because the claims already exist.
Watch for the claim that will not sit anywhere. Material that resists placement is usually either irrelevant, in which case cut it, or the seed of a chapter you have not recognised yet. Both outcomes are useful and both are invisible without the placement step.
Write from the claims, not from the sources. When drafting, work from the chapter's claim list. If a claim needs its full context, go back to the source, but start from your own compressed version. Starting from the sources means re-reading, which is where months go.
For anything with a story, three additional artefacts earn their keep, and they are not notes.
A timeline. Every dated event, in order, with its source. Narrative nonfiction lives or dies on chronology, and contradictions between sources about dates are extremely common and nearly impossible to spot without a single ordered list.
A cast list. Every person, who they are, when they appear, and how their name is spelled. Spelling matters more than it sounds; a name rendered two ways across a manuscript is a copy-editing problem you will pay for.
A places index. Locations, with their state at the relevant time, since places change names, borders and functions.
All three are tables rather than prose, and all three are best built continuously during extraction. Built retrospectively, they take weeks.
A history book, roughly ninety sources, twelve interviews, two archive visits. The arithmetic is the point.
Months 1 to 4, reading and extracting in parallel. About twenty-five sources read, extracted within a week each, producing roughly four hundred claims at four a day. The discipline that made this work was not reading faster, it was refusing to start a new source while two sat unextracted.
Month 5, the first structural read. Filtering the claims by chapter field produces the first honest look at the book: chapter two has 90 claims, chapter six has 11, and thirty claims sit unassigned. Chapter six is not a chapter. The thirty unassigned ones cluster into a subject nobody had planned, which becomes chapter nine and is later the best chapter in the book.
That finding arrived in month five and would otherwise have arrived in month fourteen, during drafting, when restructuring costs an entire pass.
Months 6 to 11, interviews and archives. Interviews transcribed and extracted the same week, which matters more than for books because spoken material decays fastest in memory. Archive visits extracted on the train home, badly, then cleaned up the following day. The timeline catches two sources disagreeing about a date by eleven months, which turns out to be the more interesting fact.
Months 12 to 16, drafting. Each chapter drafted from its claim list rather than from sources. Roughly one source in five needed to be reopened for full context, which is the number that indicates extraction was good enough. Had that been one in two, the extraction was too shallow.
Month 17, the fact-check pass. Every factual assertion checked against the source at the recorded page. Nine days. Four errors found, and the worst was a claim attributed to an author who was in fact summarising an argument they went on to reject. That error had been in the notes for fourteen months and was invisible in every reading of the manuscript.
Month 18, bibliography. Generated from Zotero in an afternoon, in the publisher's style, without anybody retyping anything.
The system above assumes eighteen months. Under a compressed schedule the correct compromises are specific, and they are not the ones people make.
Do not skip extraction, reduce coverage. The instinct under pressure is to read everything and extract nothing, which produces the year-two failure in six months instead. Read fewer sources and extract all of them.
Extract more coarsely. A claim per page rather than per idea, with the page number, is meaningfully worse than proper extraction and enormously better than highlighting. It still converts source-shaped material into something filterable.
Keep the page numbers no matter what. This is the compromise never worth making, because it converts a one-day fact-check into a two-week one, and the fact-check is the pass most likely to be cut when time runs out.
Drop the linking, keep the chapter field. The connections between claims are valuable and they are a luxury. The chapter assignment is what makes drafting possible at all.
Accept a thinner book rather than an unverified one. The failure mode of a rushed research book is not that it is short; it is that it contains four errors nobody caught, and those are permanent in a way a missing chapter is not.
| Zotero | Obsidian | Notion | Scrivener | DEVONthink | |
|---|---|---|---|---|---|
Citation management | Yes, strongest | Via plugin | Manual | Basic | Basic |
PDF annotation | Yes | Via plugin | No | Adequate | Yes, strongest |
Claim notes with links | Adequate | Yes, strongest | Yes | Adequate | Yes |
Chapter assembly and filtering | No | Yes | Yes, strongest | Yes, strongest | Adequate |
Long-form drafting | No | Adequate | Adequate | Yes, strongest | No |
Full-text search across sources | Adequate | Yes | Adequate | Yes, strongest | Yes, strongest |
Works offline | Yes | Yes, strongest | Limited | Yes | Yes |
Cost | Free, paid storage | Free personal | Free tier | One-off | One-off, higher |
Best for | Citations, always | The claim layer | The claim layer, collaborative | The draft | Large archives |

Book research with sources, claims and chapter structure held together
Turning extracted claims into a chapter structure is arrangement rather than filtering: clusters form, sequences get tried, and the outline emerges from moving things around.

These are ordered by where they enter the work, not by overall quality. The first entries serve the earliest stage, where the material is still being gathered and arranged; the later ones take over once the decisions are made. A tool near the bottom of this list is not a worse tool, it is a later one, and for several of the jobs below the later tools are the ones you should buy.


The verdict: useful for the structural stage where claims are being arranged into an argument, and not a research management system.
Best for: working out the shape of the book, with the material visible while you do it.
Why it comes first. The step where extracted claims become a chapter structure is spatial: clusters form, sequences get tried, and an outline emerges from arrangement rather than from a list. Storyflow's canvas holds text and images together and its AI reads everything on the current board plus up to one Tactic and up to three documents brought in with an @-mention.
Where it loses, and on this article's subject it loses on the essentials: no citation management of any kind, no bibliography generation, no PDF annotation, no reference metadata, and no filtered database views over hundreds of claims. For a book with a bibliography, Zotero is required and this does not replace it. The article recommends Zotero plus Obsidian or Notion, with Scrivener for the draft. Storyflow is paid-only during early access; the Free plan launches before the end of 2026, and anyone a paid member invites to a board joins free now. Plus is $7.99/mo annual, Pro $14/mo annual, Max $39/mo annual.
The verdict: use it, and use it for citations specifically. This is the least negotiable recommendation in the article.
Best for: every book project with a bibliography.
Why it is here. Citation management is a solved problem and Zotero solves it: capture a source with one click including metadata, store the PDF, generate a bibliography in any style, and insert citations into a manuscript that renumber themselves.
Trying to do citations in a note tool is a false economy that surfaces at the worst possible moment, which is the fortnight before delivery when a publisher wants a formatted bibliography and your references live as inconsistent hand-typed strings across four hundred notes.
Strengths: free and open source, browser capture, excellent metadata handling, word processor integration, group libraries for collaboration.
Limitations: it is a reference manager, not a thinking tool. Its notes are adequate and not where claims should live. The interface is functional rather than pleasant.
The verdict: the strongest home for the claim layer, especially for projects running over years.
Best for: writers who want local files, dense linking between claims, and no dependency on a company.
Why it ranks here. One markdown file per claim, linked to related claims, tagged by chapter, searchable instantly. Local plain text means the notes will open in twenty years, which for a project measured in years is a real consideration rather than a philosophical one.
The linking is the differentiator. Connecting a claim to a related claim builds the argument structure as a side effect of extraction, and the graph occasionally surfaces a connection you did not consciously make.
Limitations: it does nothing out of the box and the plugin ecosystem is a genuine time sink. Citation handling needs a plugin bridging to Zotero. Collaboration is poor. The customisation is the trap: hours spent configuring a note system are hours not spent extracting.
The verdict: the better claim layer for collaborative projects, and the weaker one for longevity.
Best for: co-authored books, projects with a researcher, and writers who want filtered views without configuration.
Why it ranks here. A claims database with source, chapter, tag and status fields gives you exactly the filtered views the method needs, with no setup beyond creating the fields. Filtering to a chapter is a click, and the chapter-as-a-list check is immediate.
Limitations: citation management is manual and inadequate, so Zotero remains necessary. Performance degrades on large databases and a book's claims are a large database. Offline access is limited, which matters in archives and on trains. Everything lives in someone else's service.
The verdict: the best drafting environment, and a mediocre research system that many writers use as both.
Best for: the draft, and for keeping reference material beside it.
Why it ranks here. Scrivener's binder holds the manuscript and the research in one project, its corkboard allows spatial arrangement of sections, and its compile handles the formats publishers want. For writing a book, it remains the strongest tool available.
The common mistake is using it as the research system too. Scrivener's research folder is source-shaped: PDFs and documents in a tree. It does not do claim-level notes with filtering well, so the shape mismatch this article describes goes unaddressed and reappears at draft two.
Limitations: weak citation handling, no linking between notes, and the research folder encourages the source-shaped organisation that causes the problem.
The verdict: the specialist for very large archives, and more than most books need.
Best for: projects with thousands of documents, particularly scanned archival material.
Why it ranks here. Its search and classification across enormous document sets is the strongest here, it OCRs scans, and its suggestion of related documents genuinely surfaces material you had forgotten.
Limitations: Mac and iOS only, a steep learning curve, a one-off cost at the higher end, and it is a document manager rather than a claim system. Below about a thousand sources it is more apparatus than the project needs.
Zotero for sources and citations. Obsidian or Notion for claims. Scrivener for the draft.
Three tools, each doing what it is best at, with two handoffs:
Writers regularly try to collapse this to one tool. Two tools is achievable if you accept Notion's weak citations for a book with a light bibliography. One tool is not achievable for anything academic, because no note tool does citation management adequately.
Before delivery, verify every factual claim against the source rather than against your notes.
This is not pessimism about your note-taking. It is that notes drift: a paraphrase written quickly loses a qualifier, a date gets transcribed with two digits swapped, and a claim attributed to the author was actually them summarising someone they disagreed with. That last error is the most common and the most damaging, and it is invisible in a note that says only what was claimed.
The practical method: go through the manuscript, and for every factual assertion open the source at the recorded page. If your extraction included exact locations this takes a day for a book. If it did not, it takes a fortnight and you will be tempted to skip it.
That single consideration justifies recording page numbers at extraction time even when it feels like unnecessary precision.
Worth addressing directly, because summarising sources is the most obvious application and the one with a specific hazard.
Genuinely useful: transcription of interviews, which has become good enough to change what is practical for oral-history work; full-text search across a large corpus in natural language; and drafting a first-pass summary of a long document to decide whether it is worth reading properly.
Useful with care: proposing which chapter a claim might belong to, which is a classification task with no factual stakes. If it guesses wrong you move a note.
The hazard is extraction itself. An AI summary of a source produces claims in the model's words rather than yours, which removes exactly the step that makes the system work. The thinking is not a by-product of extraction, it is the point of it, and a note you did not write is a note you will not remember reaching.
Worse, AI-extracted claims carry an attribution risk that is hard to detect: a model summarising a text will sometimes present an argument the author was describing as one the author held. That is the same error the fact-check pass exists to catch, generated at scale rather than occasionally.
The workable rule: use AI to decide what to read and to find things you have already read, and write the claims yourself. The reading is where the time goes; the extraction is where the book comes from.
Zotero is the most trusted for citation management, and among academics it is close to a default alongside its commercial alternatives. Obsidian has strong trust among writers wanting local files and longevity. Scrivener remains the most trusted long-form drafting tool. DEVONthink is trusted by researchers with very large archives, and Notion by collaborative and less citation-heavy projects.
Zotero is free and open source with paid storage that is modestly priced, which is the fairest arrangement in the category. Obsidian is free for personal use with paid sync. Scrivener's one-off licence is honest and includes major updates. DEVONthink's one-off cost is at the higher end. Notion's free tier is usable and its costs rise with collaborators.
Scrivener and Zotero both have long records and, importantly for multi-year projects, reliable file formats and backups. Obsidian's plain markdown files are the safest long-term storage here because they do not depend on the application continuing to exist. Notion is reliable in service and its large-database performance is the weakest for this specific use.
Zotero first and always, because it is free and citation management is the one thing you cannot improvise. Then Obsidian or Notion for claims depending on whether longevity or collaboration matters more. Scrivener for the draft if you are writing something book-length. DEVONthink only above roughly a thousand sources.
Zotero is underrated by non-academic writers who assume it is only for scholars, when any book with endnotes benefits equally. Plain markdown files in any editor are underrated as a claim layer, since the method matters far more than the software. And a spreadsheet is underrated for the timeline and cast list that narrative nonfiction needs.
The problem is a shape mismatch and the solution is extraction. Research arrives organised by source; writing needs it organised by argument; converting between the two is a small daily task or an insurmountable one in year two.
Write claims in your own words, one idea per note, with the page number and a guessed chapter. Record what you doubt as well as what you accept. Keep the read-but-unextracted pile small, because that number is the one that predicts whether the book gets written.
Use Zotero for citations regardless of what else you use. It is free, and it is the one part of this you cannot improvise later.
Zotero for sources and citations, Obsidian or Notion for the claim layer, and Scrivener for the draft. No single tool does all three well: reference managers do not think, note tools do citation management badly, and drafting tools encourage source-shaped research organisation that fails at draft two. The method matters more than the tools, and the method is extraction.
Because research arrives organised by source and writing consumes material organised by argument, and nothing converts between the two automatically. Four hundred highlighted PDFs are organised by book; chapter four needs three sentences from six different books. Without extraction, starting to write means re-reading everything, which is where multi-year projects stall.
Turning a source into individual claims, one idea per note, written in your own words, each with the source and exact page or timestamp attached, plus a guess at which chapter it belongs to. It is the step that translates source-shaped material into argument-shaped material, and it is the one writers skip because highlighting feels like progress and is a decision deferred.
On ideas, one note per claim rather than one per source. A single source contributes to five different arguments, so material that lives inside a source document cannot be gathered into a chapter. Claim-shaped notes each carry their source as a field, which makes assembling a chapter a filter rather than an archaeological dig through documents.
Because a highlight records that something struck you and not why, and eighteen months later you have the sentence without the thought. A paraphrase forces the thinking at the moment you are best placed to do it, immediately after reading. Store both: your claim as the note, and the verbatim quote as a field, because the quote is for the page and the claim is for the system.
Typically three hundred to two thousand extracted claims, depending on how research-heavy the book is. That number sounds alarming and accumulates at four or five a day without effort. The count that actually matters is not the total but how many sources sit in the read-but-not-extracted pile, which is the number that predicts whether a project will stall.
For anything with endnotes or a bibliography, yes. Citation management is a solved problem and doing it manually surfaces at the worst moment, which is the fortnight before delivery when a publisher wants a formatted bibliography and your references exist as inconsistent hand-typed strings across hundreds of notes. Zotero is free, so the only cost is learning it.
Obsidian for solo projects running over years, because local plain-text files will open in twenty years and the linking between claims builds argument structure as a side effect. Notion for co-authored projects and anything with a researcher, because filtered database views require no configuration and collaboration works. Both need Zotero alongside them for citations.
For the draft, Scrivener is the best tool available. As a research system it is mediocre, because its research folder is source-shaped: PDFs and documents in a tree, without claim-level notes and filtering. Writers who use it for both hit the shape-mismatch problem at draft two, having felt organised throughout. Keep claims outside it and draft inside it.
Assign a probable chapter at extraction time, even as a guess, because a wrong guess is trivially fixable and an unassigned claim is invisible when you write. Then read each chapter as a list of its claims: four claims means the chapter is thin, ninety means it is really two, and a claim that refuses to sit anywhere is either cuttable or the seed of a chapter you have not recognised.
A timeline of every dated event with its source, a cast list of every person with spellings, and a places index with each location's state at the relevant time. All three are tables, all three should be built continuously during extraction, and built retrospectively they take weeks. The timeline in particular is the only reliable way to catch the date contradictions between sources that are extremely common.
You do not check against the notes, you check against the sources. Notes drift: a quick paraphrase loses a qualifier, a date gets transposed, and most damagingly a claim gets attributed to an author who was actually summarising someone they disagreed with. Open the source at the recorded page for every factual assertion, which takes a day if you recorded page numbers and a fortnight if you did not.
Treating collection as progress. A growing library of highlighted PDFs feels productive and produces no usable material, because highlighting defers every decision to a future self who will have forgotten the context. The second most common is organising notes by source, which mirrors how material arrived and cannot be assembled into chapters.
Within about a week of reading, because the memory of what a source contained and why it mattered decays much faster than writers expect. Beyond a week you are re-reading rather than extracting, which doubles the cost. Extracting four or five claims a day keeps pace with normal reading and never accumulates into a wall.
Yes, and it is the most valuable annotation you will write. When a source claims something you doubt, writing down why is where original arguments come from. Writers who record only what they agree with produce books that read as summaries of their reading, because the doubt is where the contribution lives and it is the first thing forgotten.
It can produce them and doing so removes the step that makes the system work, because the thinking is not a by-product of extraction but the point of it. A note written in a model's words is a note you will not remember reaching. There is also a specific attribution risk: models summarising a text will sometimes present an argument the author was describing as one the author held, which is the exact error the fact-check pass exists to catch, generated at scale.
Deciding what to read and finding things you have already read. Transcription of interviews is genuinely transformative for oral-history work, natural-language search across a large corpus is useful, and a first-pass summary of a long document is a reasonable way to decide whether it deserves proper reading. Chapter classification of a claim is also safe, since a wrong guess costs a drag.
Read fewer sources and extract all of them, rather than reading everything and extracting nothing, which produces the year-two failure in six months. Extract more coarsely, a claim per page rather than per idea, and drop the linking between claims. Keep the page numbers regardless, because that single field is what turns a two-week fact-check into a one-day one, and the fact-check is what gets cut when time runs out.
Table of Contents
Gather sources, personas, and findings on one canvas, then let the AI read across all of it. Open any of these research boards to start.
A visual AI workspace where every feature lives inside one canvas. No tab-switching, no context lost.
Build your entire board from a single message
Type what you need in the AI chat at the bottom of your canvas. The AI adds cards, headings, and structure directly onto your board.
Use expert frameworks as AI context
Type @ in the AI chat and choose any Tactic. The AI tailors every response to that framework instead of giving generic advice.
Turn your board into a mind map in seconds
Ask the AI to restructure your canvas as a mindmap. It connects your ideas into a visual hierarchy so you can see how everything relates.
Storyflow actually began as a personal tool while working on creative and research projects.
We kept running into the same problem: ideas were scattered everywhere: notes, documents, and whiteboards.
Nothing helped us see how everything connected.
So we started building a workspace designed around how ideas actually grow.
→ Read how Storyflow was created
Justkay
Documentary Filmmaker & Founder at Storyflow
Published: 2026-09-06
Transform your creative workflow with AI-powered tools. Generate ideas, create content, and boost your productivity in minutes instead of hours.