PDF splitter for ChatGPT, Claude and other LLMs

It actually
works.

Your book is too large for the chat. Drop in a 900-page PDF and get back numbered files sized to fit the context window — cut at the real chapter breaks rather than every 5,000 characters, each one labelled with the pages it came from and the parts on either side of it.

Open the splitter no account · no upload · one HTML file

One book → six labelled parts

chapter marks in red

01 · pp. 1–44 02 · pp. 45–88 03 · pp. 89–137 04 · pp. 138–186 05 · pp. 187–226 06 · pp. 227–262 start of the book end
Using it

Three steps, about a minute

Works with textbooks, manuals, reports, novels — any PDF, plain text or Markdown file that carries actual text. There is nothing to learn: every setting has a sensible default and a sentence next to it explaining what it does.

STEP ONE

Drop the file in

PDF, plain text or Markdown. It reads the pages, finds the chapter headings, and shows you what it found before touching anything.

STEP TWO

Choose the cuts

Follow the book's own chapters, or type your own page ranges by hand and name each one. Mix the two if you like — start from the chapters, then correct them.

STEP THREE

Download the parts

A ZIP of numbered files plus a contents page. Open one, hand it to a model, and it knows which part it is and what came before.

Before you start

Four things it won't do

The claim on this page is a small one, so here is the whole of what sits behind it.

×

It can't read scanned pages

A scan is a picture of a book, not a book. There is no text to cut. Chunkbook does not do OCR and isn't going to — turning pictures back into text is a separate job, and doing it badly is worse than not doing it at all. Run the file through an OCR tool first. Compressing it doesn't help either, since a model would still be looking at pictures. Meanwhile the splitter can still cut the PDF itself into smaller page-range PDFs.

×

The size numbers are estimates

Part sizes are worked out from the mix of characters, not a real tokenizer. Expect them to be right within about 15 percent, and leave a little headroom.

×

Chapter detection guesses, sometimes wrong

When a book carries a real contents list inside the file, it uses that and gets it right. When it doesn't, it infers headings from type size and numbering — and on a busy or badly typeset book it invents headings that aren't there. The splitter says which of the two it's doing, and when it's guessing it tells you so and suggests cutting by hand instead. Every cut can be corrected either way.

×

It doesn't remember anything

Nothing leaves your machine, which is the point — and also the catch. There is no account and no saved work. Close the tab and the ranges you typed are gone.

Free, and likely to stay that way. There is no server to pay for.

Open the splitter