Getting a table out of a PDF and into Excel
There is a table in a report PDF and you need to calculate with those figures. You select it, copy, paste into Excel, and every value lands in a single cell in column A. So you retype it by hand.
Drop a PDF and the values that stand in table-like columns are split into cells. Each page becomes its own sheet, so you never lose track of where a number came from.
No upload
PDF to turn back
The PDF needs real text in it. A scanned PDF is a picture of a page, so there is nothing to pull out. Files open inside this browser and are never uploaded.
Reading order, paragraphs and tables come across. The exact placement on the page does not. A PDF never records what was a paragraph or a table, so both are reconstructed from where the characters sit.
Images, stamps and signatures printed on the page come across as well. Floating over the cells would get in the way of calculating, so they are gathered below the data on each sheet.
A scanned PDF holds photographs rather than text, so there is nothing to convert. Drop one in and the tool will say so.
Whole pages are not pasted in as pictures. What you download opens as an editable document in Word, Excel or Hangul.
The more table-like the document, the better the result. Several at once is fine.
Where the characters sit decides where each table starts and ends.
It opens in Excel or Google Sheets ready to calculate with.
Tables in PDFs very often have no lines around them at all. Where lines do exist they are drawn on a different layer from the text, invisible to anything pulling the words out. That is why approaches built on finding the lines fail so often.
So position is what gets read. Within a line, a gap much wider than a word space means a cell boundary, and when those gaps stand at the same horizontal position across several lines in a row, that block is a table. Three lines is the minimum. Two-line blocks are usually a title with a date pushed to the far right, and turning that into a table makes it harder to read, not easier.
An empty cell leaves no trace in the file, because there were never any characters there. Joining cells up in order would shift everything along by one, so each cell is dropped into the column nearest to it and the rest are left blank.
When tables are scattered over several pages, piling them into one sheet loses track of which page a figure came from. So each page becomes its own sheet, named by page number.
Lines that are not part of a table are kept too, one per row. Dropping the heading above a table, or the note saying which unit the figures are in, leaves you with numbers and no way to read them. Images and stamps printed on the page come across as well, gathered below the data on each sheet so they do not float over the cells and get in the way.
A PDF made by photographing paper or running it through an office scanner is a single picture per page. Your eyes see letters, but the file contains none. There is nothing to pull out, so there is nothing to rebuild.
The tool says so plainly instead of handing you an empty file. If only some pages are scans, it names which page numbers came back empty. Reading text out of pictures is not something this tool does.
The whole conversion happens inside this browser. Reading the PDF and writing the new document are both done by code running on your device. There is no upload step at all.
You can check the claim. Open the network tab in developer tools (F12) and run a conversion: an upload request the size of your file either appears or it does not.
No. Reading the PDF and writing the spreadsheet both happen inside this browser.
Yes, one line per row in the first column. Throwing away the heading above a table, or the unit the figures are quoted in, would leave numbers nobody can interpret.
That happens because an empty cell leaves no trace in the PDF. Dropping each cell into the nearest column reduces it a lot, but tables with very narrow column gaps can still come out misaligned.
Merged cells are not carried across. The value from a merged block lands in whichever single column is closest.
A PDF does not record whether a value is a number or a label, so it is stored exactly as it appears. Thousands separators and currency symbols make Excel read it as text. Changing the cell format in Excel fixes it.
No. A scan is a picture of a page and holds no text to extract.
There is a table in a report PDF and you need to calculate with those figures. You select it, copy, paste into Excel, and every value lands in a single cell in column A. So you retype it by hand.
You convert a PDF to Word and the table comes out crooked. You assume the converter is bad, try another one, and get much the same. Some problems do not go away by changing tools.