Skip to content
WorldofPDFs

Getting a table out of a PDF and into Excel intact

The table you can see on the page is not in the file, which explains nearly every way the conversion goes wrong.

The short answer

To convert a PDF table to Excel, use a converter that reads the PDF's text positions and rebuilds the grid from them, then check the columns before you trust the figures. It works well on PDFs exported from software. It fails on scans until OCR adds text, and it struggles with merged cells and columns with no gap between them.

Why those particular failures, and not others, comes down to one fact about the format.

A PDF has no table in it

A spreadsheet stores cells. A Word document stores a table object with rows and columns. A PDF stores neither. What it records is a list of text runs, each one a string drawn at a particular x and y position in a particular font, plus a separate set of lines and rectangles that may or may not be the ruling around those runs.

When software exports a table to PDF, it draws each cell's text at the right spot and draws the borders, then throws the table structure away. The grid your eye sees is the result of things being lined up. Nothing in the file says that "£1,240" belongs to the row labelled "March" and the column headed "Spend".

So a converter cannot read a table out of a PDF. It has to infer one. The usual method is geometric: group text runs into rows by the height of their baselines, then look for vertical channels of empty space that run down through those rows and treat each channel as a column boundary. Every cell is then placed in whichever column its position falls into.

When it works

The inference is reliable when the PDF was produced by software from real text, the columns have clear whitespace between them, and each cell fits on one line. That covers most bank statements, invoices, exported reports and accounting summaries, which is fortunate, because those are most of what people convert.

A good converter also does something about the values themselves. A figure that lands in Excel as text looks right and refuses to sum. World of PDF's PDF to Excel tool writes unambiguous numbers as real numbers, keeping currency and percentage formatting and reading accounting negatives in parentheses as negatives, while leaving anything ambiguous as text rather than guessing. It runs in the browser, so a bank statement is never uploaded to be converted.

When it fails, and why

Each common failure is the geometric method meeting a layout that breaks its assumptions.

  • Scans. A scanned table is a picture. There are no text runs to position, so there is nothing to infer from. OCR has to recognise the characters first, and any OCR misreads, a 5 read as an S or a dropped decimal point, travel straight into the spreadsheet.
  • Columns with no gap. When right-aligned figures sit hard against the label to their left, there is no whitespace channel between them, and two columns arrive as one.
  • Merged cells. A heading that spans three columns, or a label that spans two rows, has no single position that belongs to one column. Converters either put it in the first column, split it, or treat its row as prose and drop it.
  • Wrapped cell text. A description that wraps onto a second line inside its cell is, to the geometry, a second row with text in only one column. You get an extra row instead of a longer cell.
  • Tables across pages. Each page is reconstructed separately, so a table that continues onto the next page can come out as separate blocks with a repeated header row in the middle. Look for an option that puts every page on one sheet.

How to check the result before you rely on it

A converted table that is slightly wrong is more dangerous than one that is obviously broken, because it gets used. A few checks catch most problems.

Compare the row count with the PDF. Extra rows usually mean wrapped text; missing rows usually mean merged headings were dropped. Then check that numbers are right-aligned in Excel, which is how it shows a cell it reads as a number; left-aligned figures are text and will not total. Finally, sum a column and compare it with the total printed in the PDF. If the document has a total row, this is the single most useful check there is, because it catches a dropped row, a misread digit and a figure shifted into the wrong column all at once.

If the columns came out merged, try taking all the text on the page rather than detected tables only, then split the merged column in Excel. If the tool returned nothing at all, select some text in the PDF reader. If you cannot, the file is a scan, and the order of operations is OCR first, then convert.

If there is no table to find

Sometimes what you want is the text, not the grid: a list of names, a paragraph of figures, a schedule laid out as prose. Forcing that into columns produces a mess. Extracting the plain text and pasting it into the spreadsheet, then using Excel's own Text to Columns with a delimiter you choose, is often quicker than fighting a converter that is looking for a table that is not there.

Frequently asked questions

Why does my PDF table come out as one column in Excel?

The converter found no whitespace gap between the columns, usually because the text is tightly set. Try converting all text on the page and splitting the column in Excel with Text to Columns.

Can I convert a scanned table to Excel?

Only after OCR, which recognises the characters and adds a text layer the converter can read. Check the figures against the scan afterwards, because recognition errors carry straight into the cells.

Why won't the numbers in my converted sheet add up?

They probably arrived as text rather than numbers. Text figures are left-aligned in Excel by default, and a converter that types values as numbers avoids the problem.

Why do merged cells break PDF to Excel conversion?

A merged cell spans several columns or rows, so it has no single position to place it by. Converters have to guess where it belongs, and the guess is often wrong.

Is it safe to convert a bank statement online?

That depends on whether the tool uploads it. World of PDF's converter builds the spreadsheet in your browser, so the statement never leaves your device.