AI PDF Data Extraction
Pull the tables, line items and fields out of supplier documents, statements, reports and any other business PDF, scanned or digital, into Excel, CSV or JSON. Say what you need in plain words, and every file in the batch comes back in the same columns, this month and next.
What would you like to extract?
- No templates or zone setup
- Combines batches into one sheet
- Uploaded documents deleted within 48 hours
The output
What you get: one clean sheet from a folder of PDFs
Your columns, described in your prompt. Source File gives the file and page each row came from; Review Needed marks a value or a row the panel of AI agents cannot agree on, with what to check.
| Item Code | Description | Qty | Unit Price | Amount | Source File | Review Needed | |
|---|---|---|---|---|---|---|---|
| 2 | CB-1104 | Cable tray, galvanized 100mm | 40 | 18.60 | 744.00 | [Page 1] pricelist_q3.pdf | |
| 3 | CB-1108 | Cable tray, galvanized 200mm | 25 | 27.90 | 697.50 | [Page 1] pricelist_q3.pdf | |
| 4 | FX-0221 | Fixing kit, M8 assorted | 120 | 3.15 | 378.00 | [Page 2] pricelist_q3.pdf | |
| 5 | LT-3300 | LED batten 1500mm 4000K | 60 | 22.40 | 1,344.00 | [Page 3] pricelist_q3.pdf | |
| 6 | PV-8812 | Isolator switch 32A | 35 | 11.85 | 414.75 | scanned_rates.pdf | Verify unclear price |
| 7 | PV-8814 | Isolator switch 63A | 18 | 16.20 | 291.60 | scanned_rates.pdf | |
| 8 | CN-4410 | Conduit, PVC 25mm × 3m | 200 | 2.48 | 496.00 | [Page 1] supplier_b_rates.pdf | |
| 9 | CN-4415 | Conduit bends 25mm | 150 | 0.86 | 129.00 | [Page 1] supplier_b_rates.pdf | |
| 10 | GL-0034 | Gland pack, brass 20mm | 80 | 4.35 | 348.00 | [Page 2] supplier_b_rates.pdf | |
| 11 | TB-9920 | Terminal block strip 12-way | 90 | 1.95 | 175.50 | [Page 2] supplier_b_rates.pdf |
Three steps
How to extract data from PDFs
Upload your PDFs
Digital PDFs, scans and phone photos, exactly as you received them, with every supplier's layout in one batch. A pack that mixes document types can go in as one file: tell our AI which pages to ignore, cover sheets or summary pages for instance, and it will.
Describe the data you want
The table, the fields, one row per line or per document, in plain words. One description covers every layout in the batch, and where your instructions or the documents leave something genuinely open, our AI can stop and ask you rather than decide on its own.
Download, and know what to verify
Excel, CSV or JSON. In the results viewer, one click opens any row beside the page it came from, so you can check a Review Needed value, or spot-check a sample of a big job, against the PDF itself.
Reliability
Why not just upload the PDFs to a chat app?
A chat app or an agent reading PDFs on its own can hallucinate a value, skip a page of a long PDF, or report success over a failure, and nobody finds out. Here:
A panel of AI agents has to agree on every value
Every cell of every table, so you are not trusting a single pass over a page.
What the panel cannot agree on is flagged as Review Needed
With what to check and where to look, so you check only the less certain values, not the whole spreadsheet.
A page or document that fails is reported, never skipped in silence
You are told which file and which page, and you are not charged for it.
A single PDF up to 5,000 pages, a batch up to 6,000 files
Every page and every file read and checked the same way as the first, so a price list that runs to hundreds of pages, or a month of supplier PDFs, goes in as one job, and several jobs can run at once.
The same columns and formats, every time
Every supplier's layout, in any language, comes out in the same columns, date and number formats, and again next month, so it imports into your accounting software or feeds your own workflow without hand-fixing.
The same extraction from your code or your agent
Node.js and Python SDKs, and a REST API from any language. Or let your own agent run the extraction itself, through the MCP server or the skill it installs from this site.
Solutions
Extracting a specific document type?
The most common jobs each have a page of their own: invoices to Excel, bank statements to Excel or CSV, receipt OCR, payroll data extraction and utility bill data extraction.
Questions
PDF data extraction FAQ
How do I extract data from a PDF to Excel?
Upload your PDFs, digital or scanned; images work too. Say what you need in plain words: the table on each page, particular fields, one row per line item. Download the result as Excel, CSV or JSON. There are no templates or capture zones to set up; the prompt is the whole setup.
How is this different from a PDF-to-Excel converter?
A converter copies each page into a spreadsheet as it looks, headers, footers and page breaks included, and leaves the tidying to you. Extraction gives you only the data you describe, in the columns you name, one row per record, in the same shape for every file in the batch.
Can it extract tables that span multiple pages?
Yes. A table that runs across pages, repeats its header on each one, or changes layout from one document to the next comes out as one set of rows in your columns. A single PDF can run to 5,000 pages, every page read and checked the same way as the first, and every row carries the file and page it came from.
Does it work on scanned PDFs?
Yes. Scanned PDFs, lower-quality scans included, and phone photos (JPG, PNG) go in the same batch as digital PDFs, under the same prompt. A value or a row the panel of AI agents cannot agree on is flagged as Review Needed, with what to check.
Can it combine many PDFs into one spreadsheet?
Yes. A batch of up to 6,000 files comes back as one sheet, with the same columns and formats for every document and the file and page each row came from. A file that fails is reported, never skipped in silence, and you are not charged for it.
What kinds of documents does it handle best?
Business documents, in any language or script: supplier statements, with each invoice on the statement as a row; invoices, purchase orders and delivery notes, with the numbers, quantities and amounts for matching; price and product lists, however long; reports, order confirmations, and packs that mix them.
Is it free?
Every account gets 50 free pages every month, with full functionality and no credit card to start. Above that, credits start at $12 for 100 pages. You pay for pages processed and nothing else. There is no plan to size and no monthly charge: a quiet month costs nothing, and credits bought for a busy month keep for 18 months.
Is there a PDF data extraction API?
Yes. Node.js and Python SDKs, and a REST API from any language, run the same extraction as the web app and return the same Excel, CSV or JSON, Review Needed included, with documentation a coding assistant can build a working integration from. Your own agent can run it too, through the MCP server or the skill it installs from this site. Same account, same balance, same results in your dashboard.
Extract your first PDFs free
50 free pages every month: no subscription, no credit card. Start with a few of this month's PDFs, and save the prompt so next month's come back in the same columns.