Skip to content
WorldofPDFs

Why convert a PDF to Markdown before giving it to an AI model?

A language model reads structure as well as words. Raw text copied out of a PDF throws the structure away; Markdown keeps it in a form the model already understands.

The short answer

To convert a PDF to Markdown for ChatGPT or another language model, extract the text with its structure intact: headings marked with #, lists marked with -, paragraphs separated by blank lines. That gives the model the document's outline as well as its words, which helps it find sections, summarise accurately and quote the right part.

Copying from a PDF viewer and pasting does the opposite. It hands over the words with the outline removed and a scatter of line breaks in the middle of sentences.

Why pasted PDF text comes out broken

A PDF does not store paragraphs, headings or lists. It stores instructions to draw particular glyphs at particular coordinates. A line break in the pasted text is simply where the typesetter ended a line on the page; a heading is text that happens to be larger and bolder; a two-column layout is two blocks of text that happen to sit side by side.

Select-and-copy in a viewer reproduces the geometry rather than the meaning. Lines break mid-sentence, columns can interleave, running headers and page numbers land in the middle of paragraphs, and nothing distinguishes a section title from the sentence below it. A model will still read it, but it is reading a flattened document and inferring the structure, and inference is where summaries go wrong.

What Markdown preserves

Markdown is a good intermediate format for a model for the plain reason that models have read an enormous amount of it. A line beginning with ## is unambiguous: a section starts here. A dash at the start of a line is a list item. A blank line ends a paragraph. None of that needs explaining in your prompt.

  • Headings. A converter infers them from relative size: it measures the size the bulk of the document is set in, and treats blocks clearly larger than that as headings, one level per step up. Measuring against the document itself rather than fixed point sizes matters, because a 9pt paper and a 13pt report disagree about what counts as big.
  • Lists. Lines starting with a bullet character or a number followed by a full stop or bracket become list items rather than a run of short sentences.
  • Paragraphs. Lines are joined where the spacing says the sentence continues, so the model sees whole sentences.
  • Page boundaries. Marking where each original page ended keeps references like "see page 12" meaningful without inventing headings that were never in the document.

Tables are the hard part

A PDF has no table object. It has text runs that line up into a grid because the typesetter put them there. Turning that back into rows and columns means clustering text by position and guessing where one cell ends, and the guess fails on merged cells, wrapped text and tables without ruling lines.

The World of PDF Markdown converter does not attempt it. Table contents come through as text in reading order rather than as a Markdown table, because a confidently wrong grid is worse than honest text: a model will happily answer from a figure that landed in the wrong column. When the tables are the point, PDF to Excel rebuilds them as a spreadsheet using column detection, which you can check by eye before pasting the rows into the conversation.

When to OCR first

If the PDF came from a scanner or a phone camera, every page is a picture and there is no text to convert. The Markdown converter detects that and says so rather than producing an empty file. Run OCR PDF over it first to add a real text layer, then convert.

Two caveats are worth carrying into the conversation. OCR is a best guess, so figures, names and reference numbers deserve a proofread before a model treats them as fact; 0 and O, 1 and l are the classic confusions. And many AI tools will read an uploaded scan using their own vision models, which can work, but you then have no copy of what text the model actually saw. Converting locally gives you that text to inspect.

A workflow that works

Convert only what you need. A page range keeps a long report down to the section you are asking about, which leaves the model less irrelevant material to wade through and keeps you from sending pages nobody needed to see.

Skim the Markdown before pasting it. Heading levels come from font sizes, so a document that uses size inconsistently can produce an uneven hierarchy; fixing a few # marks takes a minute. Then paste it with a short instruction that refers to the structure — "answer from the section headed Termination" works far better when that heading exists in what you sent.

Where plain words are all you want — a quote, a search index, a word count — Extract Text gives you the flat text without markup. For anything a model has to navigate, keep the structure.

Both run in the browser on World of PDF, so the conversion itself sends nothing anywhere. Pasting the result into a chat is a separate decision, and worth thinking about for sensitive documents.

Frequently asked questions

Is it better to upload a PDF or paste Markdown into ChatGPT?

Pasting Markdown gives you control over exactly what the model sees, and keeps headings and lists explicit. Uploading the PDF leaves the extraction to the service, and sends the whole file rather than just the parts you chose.

Does PDF to Markdown keep tables?

Not as Markdown tables. Table contents come through as text in reading order. To rebuild rows and columns, convert the table with PDF to Excel instead.

Why does my PDF convert to empty Markdown?

It is almost certainly a scan, so the pages are images with no text layer. Run it through OCR first and then convert.

Does Markdown use fewer tokens than the raw PDF text?

Markdown adds only a few characters per heading or list item, so it costs about the same as plain text and far less than HTML. The gain is in clarity rather than length.

Is my PDF uploaded when I convert it to Markdown?

No. On World of PDF the conversion runs entirely in your browser. What you then paste into an AI service is up to you.