Skip to content

Convert PDF to Excel

Pull the tables out of a PDF into a spreadsheet you can edit.

Drop your file here

Processed securely and deleted within the hour.

Up to 100 MB · free · no signup

    What is actually possible here

    A PDF has no idea it contains a table. It contains characters at coordinates and, sometimes, lines drawn near them. "Extracting a table" means inferring the rows and columns from those positions — so the result depends on how the PDF was made, and being clear about that up front saves a lot of disappointment.

    Where it works well: PDFs exported from a spreadsheet, a database report, an accounting package or a bank. The characters are real text, the columns line up to the pixel, and the recovered grid is usually exact.

    Where it works less well: tables with merged cells, multi-line cells, or no ruling lines and irregular spacing. You get the data with the grid slightly wrong, which is still much faster than retyping.

    Where it does not work at all: a scanned page. A scan is a photograph, and there is no text in it to find — every cell would come back empty. Run OCR over it first to add a text layer, then come back.

    How the columns are found

    Detection uses pdfplumber, which works from the ruling lines the document draws and, where there are none, from the vertical alignment of the characters themselves. That second case is the interesting one: a table with no borders is held together only by the fact that every value in a column starts at the same x position, and that is enough to recover it.

    It is also why a column of right-aligned numbers next to a column of left-aligned text sometimes merges — there is no gap of whitespace running the full height of the table to split on.

    One sheet or many

    A document with several tables can put each on its own sheet, or all of them on one sheet separated by a blank row. Separate sheets are easier to work with when the tables have different columns; one sheet is easier when they are the same table continued across pages, because you can delete the repeated headers and have a single range.

    Page ranges save time and mistakes

    Naming the pages you want — 4-9, or 2,5,11 — is worth doing on a long report. Detection over a hundred pages of prose finds spurious tables in indented paragraphs and address blocks, and you then have to pick your real data out of them.

    Check the numbers before you trust them

    One habit worth having: total a column in the spreadsheet and compare it with the total printed in the PDF. Extraction errors are almost always structural — a merged cell, a missed row, a split column — and a total that disagrees finds them immediately. A total that agrees is good evidence the whole table came through.

    01

    Common questions

    Will it work on a scanned PDF?

    No. A scan is a photograph with no text in it, so every cell comes back empty. Run OCR over it first to add a text layer.

    Why are two columns merged into one?

    There was no column of whitespace running the full height of the table to split on — common where right-aligned numbers sit beside left-aligned text.

    Which PDFs give the best results?

    Ones exported from a spreadsheet, database or accounting system. The characters are real text and the columns align exactly, so the grid is usually recovered perfectly.

    How do I check the result is right?

    Total a column and compare it with the total printed in the PDF. Extraction errors are structural, so a disagreeing total finds them straight away.

    Should I use one sheet or one per table?

    One per table when the tables have different columns. One sheet when it is the same table continued across pages, so you can delete repeated headers and have a single range.

    Can I convert only some pages?

    Yes, and it is worth it on a long document — detection over pages of prose finds tables in indented paragraphs that you then have to sort through.