Scrape data from PDF — into a clean .xlsx

Upload the PDF on this page and get an editable spreadsheet with the table’s rows, columns, and numbers preserved.

AI extraction can contain errors. Verify the output against the source document before using it for accounting, reporting or any decision.

🤖
AI Table Extraction

Text PDFs are parsed on our server; scans are read by GPT-5 vision

📊
Excel, CSV, Word, PowerPoint

One result, any format from the Download menu: .xlsx, .csv, .docx, .pptx

🧾
Statements → QuickBooks, Xero

Bank and card statements export to .qbo, .ofx and import-ready CSV for QuickBooks and Xero

🔒
No Registration

3 files a day free, files deleted right after processing. Pro: $9/mo, no daily limit

Short answer: Use PDF2XLS to scrape tables from PDFs into a real .xlsx. It’s free for 3 files/day with no signup. Upload up to 10 MB, including scans and photos—GPT‑5 vision handles OCR automatically. You’ll download a spreadsheet with rows, columns, and numeric values preserved.

  • Free tier: 3 files/day, no signup
  • Size cap: 10 MB per file
  • Output: .xlsx for Excel, Google Sheets, LibreOffice, Numbers
  • Scans/photos: OCR via GPT‑5 vision, no toggle required
  • Privacy: Files are processed on the server and deleted immediately
  • Works on: Any browser — Windows, macOS, Linux, iPhone, iPad, Android

How do I scrape data from a PDF into Excel?

You can do it right here: upload the file, and get an .xlsx with the table laid out as rows and columns. Numbers come back as numbers so you can filter, sort, and total immediately.

Step-by-step: extract a table from your PDF

  1. Upload the PDF. Use the widget above. One file at a time, up to 10 MB.
  2. Let the AI read the pages. Native PDFs and scans go through the same pipeline; you don’t need to switch anything on.
  3. Download the .xlsx. The result opens in Excel, Google Sheets, LibreOffice Calc, and Apple Numbers.
  4. Open and scan the sheet. Check header rows, column boundaries, and that numeric columns are right-aligned and calculate.
  5. Save and name it. Keep the original PDF nearby for spot checks while you work.

If your PDF is a scan or a photo

Scanned documents are handled automatically with GPT‑5 vision (OCR included). If text is faint, skewed, or blurred, accuracy may drop. If possible, use a flatter photo, higher contrast, or a cleaner scan. For more on scans, see OCR online for tables.

Can I do this without PDF2XLS?

Yes, and sometimes that’s better for a specific file.

  • Excel’s Data > Get Data > From PDF: Works on text-based PDFs, lets you pick detected tables. It won’t read image-only scans and can miss irregular layouts.
  • Copy–paste from a PDF viewer: Fine for simple, single-page tables; breaks on multi-line cells and complex grids.
  • Google Drive + Google Docs: Upload the PDF, Open with Google Docs, then copy the table. Docs focuses on text, not table structure, and often loses columns. If you live in Sheets, converting here first and then importing the .xlsx is usually cleaner; see PDF to Google Sheets for the exact steps.

What to verify before you trust the spreadsheet

Spot-check a few rows against the PDF, then run this quick audit:

  • Columns/headers: Each header should sit above its column, no shifts left/right.
  • Numbers are numeric: Try SUM or sort; if it fails, retype one cell to nudge the column to numeric.
  • Dates: Confirm the format (MM/DD/YYYY vs DD/MM/YYYY) and that sorting by date behaves correctly.
  • Multi-line cells: Descriptions that wrap should remain in a single cell, not spill into the next column.
  • Totals: If the PDF shows a total, recalc it in Excel to confirm the same result.

Known limits and edge cases

PDF2XLS keeps the grid and values, not visual design. Fonts, colors, logos, images, charts, headers/footers, and formulas do not carry over. Max file size is 10 MB. One file at a time. We don’t support password‑protected PDFs, bulk conversion, or an API. Processing happens on the server through the OpenAI API; we delete files right after conversion.

Need CSV or to work in Sheets?

The output is .xlsx. To get a text file, open it in Excel or Sheets and export to CSV with your preferred delimiter and UTF‑8. If CSV is the goal, see PDF to CSV for exact export settings. To stay in Google’s ecosystem, follow the steps in PDF to Google Sheets.

Ready to go? Use the widget above to convert your PDF now.

FAQ

Yes. You can convert 3 files per day for free with no signup and no credit card. If you hit the limit, wait for the daily reset or split work across days.

Yes. Scans and photos go through GPT‑5 vision, so OCR is automatic. Very low‑quality, skewed, or blurry images may reduce accuracy; a clearer scan improves results.

Up to 10 MB per file. Larger PDFs need to be trimmed or split before uploading. The tool processes one file at a time.

No. Files are processed on the server and deleted immediately after conversion. Table recognition uses the OpenAI API; documents are not stored.

No. The converter extracts the table’s rows, columns, and cell values only. Visual styling, images, logos, charts, and formulas do not carry over.

Not here. This tool converts one file at a time and doesn’t offer an API. For automation, consider desktop scripts or libraries; for Google Sheets, import the .xlsx after conversion.