Getting a table out of a PDF and into Excel
There is a table in a report PDF and you need to calculate with those figures. You select it, copy, paste into Excel, and every value lands in a single cell in column A. So you retype it by hand.
Knowing why that happens usually saves you from the retyping.
Why copy and paste dumps everything into one cell
Because there is no table inside the PDF. The file holds instructions like print 1,204 at x 220, y 540, and nothing records that this value was the third row of the second column. That structure is discarded the moment a document is exported to PDF.
So the copy function simply strings characters together in the order it sees them. The gap between one cell and the next is not a tab or a comma, just a difference in coordinates, which gives the receiving application nothing to split on.
This one fact explains the rest. Rebuilding a table means reading where values sit rather than reading lines or tags, and reading position is always an inference.
When numbers arrive as text
Even when the cells split correctly, Excel often refuses to treat the values as numbers. A PDF does not record whether something is a number or a label, so it is carried across exactly as it appears. Most cases clear up with a couple of steps in Excel.
| If it arrived like this | Cause and fix |
|---|---|
| 1,204 sitting left aligned | The thousands separator makes Excel read it as text. Change the cell format to number, or strip the commas with find and replace |
| £18,400,000 will not add up | A currency symbol leads the value. Remove the symbol and set the cell format to currency instead |
| (1,204) not treated as negative | Accounting style puts negatives in brackets. If Excel does not recognise it, drop the brackets and prefix a minus sign |
| 38% is not 0.38 | It is text with a percent sign attached. Remove the sign and set the cell format to percentage |
| 1204 in unusually wide digits | Full-width numerals, common in documents from East Asia. Convert them to half-width with a formula |
| A reference number turned into a date | The opposite problem: Excel guessed too eagerly. Format the column as text before pasting |
When cells shift by one column
Empty cells cause this. An empty cell leaves no trace in a PDF at all: there are no characters, so there is no instruction to print any. That row therefore looks like it has one cell fewer, and everything after it slides across by one.
A converter reduces the problem by dropping each value into the column nearest to it, but tables with very narrow column gaps can still come out misaligned. Merged cells behave the same way. Your eye sees two cells joined; the file contains one piece of text sitting slightly further left than its neighbours.
Which PDFs work well
How the original was produced decides the outcome.
- Works well. A table built in Excel or Word and exported to PDF. The columns stand square, so position alone is enough to split the cells.
- Mixed results. Reports laid out in publishing software, where centred text and indents make the left edge of a column wander.
- Does not work. A scanned PDF. Every page is a photograph and holds no text at all to extract. This is not a shortcoming of the tool.
- Watch out for. Long tables spanning several pages. The header repeats on each page, so delete the duplicate header rows after converting.
Always check the numbers afterwards
Moving a table involves inference, which means it can fail quietly. A missing value or a one-column shift still looks perfectly tidy on screen.
If the original has a totals row, that is your best check. Sum the same column in Excel and compare it against the printed total. With no totals row, at least count the rows and columns. That single check costs far less than writing a report on the wrong figures.
Common questions
- Is there a way to pull out only the table?
- Splitting the PDF down to the pages that contain the table before converting gives a cleaner result. Mixing body paragraphs in leaves more for the tool to judge about where the table begins.
- All my numbers arrived as text.
- A PDF does not record whether a value is a number or a label, so it is stored as it appears. Selecting the column in Excel and running Text to Columns once from the Data tab usually converts the whole lot at a stroke.
- What happens to tables with merged cells?
- Merged cells are not carried across. The value from a merged block lands in whichever single column is closest. Re-merging in Excel afterwards is quicker than fighting it.
- What about a table spanning several pages?
- Each page becomes its own sheet. To join them, copy one sheet below the other and delete the repeated header rows.
- Is a scanned PDF really impossible?
- By extraction, yes. A scan is a photograph of a page and holds no text. Asking whoever issued the document for the original file is usually faster.
Last updated September 17, 2026