Chunkbook.

Cut a long document into labelled parts that fit in a model's context — along its own chapter boundaries, or along ranges you set by hand.

runs in this browser
nothing is uploaded
pdf · epub · txt · md
by Soroush Najjar
One — Source
WHY

A long book won't fit in one chat. Sending it in random slices breaks sentences and loses the thread.

WHAT THIS DOES

Cuts the book at its own chapter boundaries, and writes on each piece which part it is and which pages it covers.

WHAT YOU GET

A folder of numbered files. Hand them to a model one at a time — each one says where it sits in the book.

Drop a PDF, EPUB, TXT or Markdown file

Or choose one. Large files are fine — the whole thing stays on your machine.

Reading…
File
Pages
Words
Est. tokens
Outline entries
This book has no text layer. Every page is a picture, so there is nothing to read or cut. Chunkbook doesn't do OCR — but one command fixes it, and then everything here works.


        

Paste both lines into Terminal. The first moves to your Downloads folder — change it if the book is somewhere else. A few minutes for a few hundred pages, then drop the new file back here. The pages look the same; the words just become readable.

First time? Install it

          

command not found means it isn't installed yet — run the line above for your system first. On macOS you need Homebrew before brew works. Windows: install WSL, then use the Ubuntu line.

No terminal?

Acrobat Pro — Scan & OCR → Recognise Text. Paid, but it keeps the layout.
Google Docs — free, but throws away page numbers and structure. Poor for books.

Or use Manual ranges below to cut this PDF into smaller page-range PDFs as it is.
The printed page numbers don't match the PDF.
Ranges and labels below follow whichever you pick.
This is an EPUB, so there are no fixed pages. An ebook reflows to whatever screen it is shown on, so the book has no page 412 of its own. It has been divided into numbered sections of roughly a page's worth of text each, which is what the ranges and labels below count in.

The good news is that the chapter structure is exact rather than inferred: it was read straight out of the book's own contents file.
Pictures don't survive a text export. Diagrams, treatment algorithms, tables set as images, scans — none of it comes through, and the text around it will keep referring to figures that aren't there.
Takes about as long as reading the file did.

are contents, index or reference lists — . Found by their shape, and by what the book's own outline calls them. These are lists of page numbers rather than prose, so they cost a great deal and tell a model nothing.
Reading the bar below. It is the whole file, left to right. Red ticks are the chapter headings that were found. Blue is covered by a part, gold is covered twice, grey is not covered at all.
A token is roughly three quarters of an English word — a 6,000-token part is about 4,500 words.
part heading overlap not covered contents / index
Two — Cutting

Cut along the book's own structure

Read from the book's own outline when it has one, worked out from type size and numbering when it doesn't. Any language.

What this book is made of

Pick the unit to cut at. Anything still too big is broken down further on its own.

Fine tuning
Detected structure
This book has no outline and no headings could be worked out from the text, so it has been cut into even parts instead. Switch to manual ranges if you want to place the cuts yourself.

Cut at page ranges you choose

Name each part, give it a first and last page. Ranges may overlap and needn't cover the whole file.

#Part nameFirst pageLast pageEst. tokens
No ranges yet. Add one, or fill the file evenly.
Three — Parts

Parts

Four — Something wrong?

Tell us what broke

This tool guesses a great deal — where chapters begin, which pages are contents, how the numbering lines up. When it guesses wrong, that is worth knowing.

What gets included

Nothing is sent until you press a button. No page content, no file name — only numbers about how the tool behaved. Edit anything you like.