Your book is too large for the chat. Drop in a 900-page PDF and get back numbered files sized to fit the context window — cut at the real chapter breaks rather than every 5,000 characters, each one labelled with the pages it came from and the parts on either side of it.
One book → six labelled parts
chapter marks in red
Works with textbooks, manuals, reports, novels — any PDF, plain text or Markdown file that carries actual text. There is nothing to learn: every setting has a sensible default and a sentence next to it explaining what it does.
PDF, plain text or Markdown. It reads the pages, finds the chapter headings, and shows you what it found before touching anything.
Follow the book's own chapters, or type your own page ranges by hand and name each one. Mix the two if you like — start from the chapters, then correct them.
A ZIP of numbered files plus a contents page. Open one, hand it to a model, and it knows which part it is and what came before.
The claim on this page is a small one, so here is the whole of what sits behind it.
A scan is a picture of a book, not a book. There is no text to cut. Chunkbook does not do OCR and isn't going to — turning pictures back into text is a separate job, and doing it badly is worse than not doing it at all. Run the file through an OCR tool first. Compressing it doesn't help either, since a model would still be looking at pictures. Meanwhile the splitter can still cut the PDF itself into smaller page-range PDFs.
Part sizes are worked out from the mix of characters, not a real tokenizer. Expect them to be right within about 15 percent, and leave a little headroom.
When a book carries a real contents list inside the file, it uses that and gets it right. When it doesn't, it infers headings from type size and numbering — and on a busy or badly typeset book it invents headings that aren't there. The splitter says which of the two it's doing, and when it's guessing it tells you so and suggests cutting by hand instead. Every cut can be corrected either way.
Nothing leaves your machine, which is the point — and also the catch. There is no account and no saved work. Close the tab and the ranges you typed are gone.
Free, and likely to stay that way. There is no server to pay for.
Open the splitter