Why PDF to Excel loses your columns, and how to get clean ones.

You converted the PDF, opened the sheet, and the columns are wrong: a description split across three, a total sitting under the wrong header, a blank column every second row. It is not you and it is not the file. It is what a converter is.

  • Last updated2026-09-13

The short answer

A PDF has no columns. It has characters placed at coordinates on a page. A converter draws a grid over those coordinates and puts each character in whichever cell it falls into. When a line of text wraps, the second half falls into a new row. When a total is printed slightly to the right of the amounts above it, it falls into a new column. The converter did exactly what it does; the layout simply did not survive being treated as a table.

The way to get clean columns is to stop asking the tool to guess the table from the page, and tell it which columns you want. Supplier, invoice number, date, VAT, total. Then the tool's job is to find those values, wherever the layout put them, and the columns come out clean because you defined them. That is what Saff does, and it is the whole difference.

The four ways columns go wrong, and what each one looks like

Wrapped text becomes extra rows. A line-item description that ran to two lines on the page arrives as two rows, the second with nothing in the other columns. Filter for blanks in the quantity column and you will find them. Merged cells become extra columns. A field printed slightly right of its neighbour, the total under the amounts or a right-aligned balance, lands in a column of its own, so the sheet has more columns than the page had.

Different pages, different grids. The converter fits a grid per page. Page two of a long invoice has a different header and different column positions, so the grid changes and the columns shift mid-sheet. Different suppliers, different tables. Convert ten invoices from ten suppliers and you get ten sheets with ten column orders, which is why the copy-and-paste into a master sheet takes longer than typing would have.

A converter can be tuned for one of these at a time, and you will find settings for column detection and for keeping the layout. None of them can know that the number printed in the shaded box is the VAT and that the one below it is the total, because that is not a layout fact. It is a meaning fact, and a grid does not carry meaning.

What breaks, and what a column-first tool does instead
What you see in the sheetWhy the converter did itWhat Saff does instead
A description split across two or three rowsThe text wrapped on the page; each line became a row.The description is one field in one row, however many lines it took.
An extra column with only totals in itTotals were printed right of the amounts, so they got their own grid column.Total is a column you named; the value goes there.
Columns shift halfway down the sheetA new page, a new grid.Every row is read against the same table, whatever the page.
Ten invoices, ten column ordersEach supplier's layout produced its own grid.One table, one column order, every supplier.
A number in the wrong columnThe value sat closer to the neighbouring header than its own.The value is matched to a named field, not to a position.
Nothing at all, or the picture in cell A1The PDF was a scan with no text layer.The page is read as an image; unclear values are flagged.

How Saff keeps the columns you asked for

You start with the table, not the file. Name the columns, give each a short description if the name is ambiguous, and choose a type: text, date, number, currency. Then upload the PDFs. Saff reads each document for those fields and puts each value in the column you named. A supplier who prints the total top-right and one who prints it bottom-left produce identical rows, because the column was never about position.

Where a value is unclear, the cell is flagged and the page shown beside it, so the check is a glance rather than a hunt. Export gives you one workbook with the same columns for every document you uploaded, ready to sort, filter and total without a single merged cell to undo.

When a converter is still the right tool

If the PDF was exported from a spreadsheet in the first place, a single sheet with a clean grid and no wrapped text, a converter will reproduce it well and there is no reason to define columns. The problems above appear on documents that were designed to be read by people: invoices, statements, delivery notes, forms.

And if you need the page reproduced, layout and all, rather than the data pulled out of it, a converter is the tool for that. Saff gives you a table of the values you asked for, not a copy of the page.

How do I convert a PDF to Excel and keep the columns?

Decide what the columns are before you convert, and use a tool that fills named columns rather than drawing a grid. In Saff, create the table with your column names, upload the PDF, and each value goes to the column you named regardless of where it sat on the page.

Why is my PDF to Excel converter not working?

If the sheet comes back blank, the PDF is probably a scan with no text layer, and a converter that copies text finds nothing. If the sheet comes back with the wrong columns, the layout did not fit a grid. Both are what a converter does, not a fault you can fix in its settings.

Can I convert a PDF to Excel column by column?

Yes, that is exactly how Saff works: you name the columns and it fills them. You can add or remove a column later and re-run the batch; the other columns are unaffected.

Does this work for scanned PDFs too?

Yes. A scan is read as an image, the values are extracted into your columns, and anything unclear is flagged for a person to check.

Get the columns you asked for

Name your columns, upload the PDF that would not convert, and see the row come back clean. Fifteen documents a month are free.