AI PDF Data Extraction

Pull the tables, line items and fields out of supplier documents, statements, reports and any other business PDF, scanned or digital, into Excel, CSV or JSON. Say what you need in plain words, and every file in the batch comes back in the same columns, this month and next.

What would you like to extract?

Add filesPDF, JPG, PNG

50 pages free every month ·No subscription ·No credit card

Extract
  • No templates or zone setup
  • Combines batches into one sheet
  • Uploaded documents deleted within 48 hours

The output

What you get: one clean sheet from a folder of PDFs

Your columns, described in your prompt. Source File gives the file and page each row came from; Review Needed marks a value or a row the panel of AI agents cannot agree on, with what to check.

extracted_tables.xlsx
.xlsx .csv .json
Item CodeDescriptionQtyUnit PriceAmountSource FileReview Needed
2CB-1104Cable tray, galvanized 100mm4018.60744.00[Page 1] pricelist_q3.pdf
3CB-1108Cable tray, galvanized 200mm2527.90697.50[Page 1] pricelist_q3.pdf
4FX-0221Fixing kit, M8 assorted1203.15378.00[Page 2] pricelist_q3.pdf
5LT-3300LED batten 1500mm 4000K6022.401,344.00[Page 3] pricelist_q3.pdf
6PV-8812Isolator switch 32A3511.85414.75scanned_rates.pdfVerify unclear price
7PV-8814Isolator switch 63A1816.20291.60scanned_rates.pdf
8CN-4410Conduit, PVC 25mm × 3m2002.48496.00[Page 1] supplier_b_rates.pdf
9CN-4415Conduit bends 25mm1500.86129.00[Page 1] supplier_b_rates.pdf
10GL-0034Gland pack, brass 20mm804.35348.00[Page 2] supplier_b_rates.pdf
11TB-9920Terminal block strip 12-way901.95175.50[Page 2] supplier_b_rates.pdf

Three steps

How to extract data from PDFs

  1. Upload your PDFs

    Digital PDFs, scans and phone photos, exactly as you received them, with every supplier's layout in one batch. A pack that mixes document types can go in as one file: tell our AI which pages to ignore, cover sheets or summary pages for instance, and it will.

  2. Describe the data you want

    The table, the fields, one row per line or per document, in plain words. One description covers every layout in the batch, and where your instructions or the documents leave something genuinely open, our AI can stop and ask you rather than decide on its own.

  3. Download, and know what to verify

    Excel, CSV or JSON. In the results viewer, one click opens any row beside the page it came from, so you can check a Review Needed value, or spot-check a sample of a big job, against the PDF itself.

Reliability

Why not just upload the PDFs to a chat app?

A chat app or an agent reading PDFs on its own can hallucinate a value, skip a page of a long PDF, or report success over a failure, and nobody finds out. Here:

  • A panel of AI agents has to agree on every value

    Every cell of every table, so you are not trusting a single pass over a page.

  • What the panel cannot agree on is flagged as Review Needed

    With what to check and where to look, so you check only the less certain values, not the whole spreadsheet.

  • A page or document that fails is reported, never skipped in silence

    You are told which file and which page, and you are not charged for it.

  • A single PDF up to 5,000 pages, a batch up to 6,000 files

    Every page and every file read and checked the same way as the first, so a price list that runs to hundreds of pages, or a month of supplier PDFs, goes in as one job, and several jobs can run at once.

  • The same columns and formats, every time

    Every supplier's layout, in any language, comes out in the same columns, date and number formats, and again next month, so it imports into your accounting software or feeds your own workflow without hand-fixing.

  • The same extraction from your code or your agent

    Node.js and Python SDKs, and a REST API from any language. Or let your own agent run the extraction itself, through the MCP server or the skill it installs from this site.

Solutions

Extracting a specific document type?

The most common jobs each have a page of their own: invoices to Excel, bank statements to Excel or CSV, receipt OCR, payroll data extraction and utility bill data extraction.

Questions

PDF data extraction FAQ

How do I extract data from a PDF to Excel?

Upload your PDFs, digital or scanned; images work too. Say what you need in plain words: the table on each page, particular fields, one row per line item. Download the result as Excel, CSV or JSON. There are no templates or capture zones to set up; the prompt is the whole setup.

How is this different from a PDF-to-Excel converter?

A converter copies each page into a spreadsheet as it looks, headers, footers and page breaks included, and leaves the tidying to you. Extraction gives you only the data you describe, in the columns you name, one row per record, in the same shape for every file in the batch.

Can it extract tables that span multiple pages?

Yes. A table that runs across pages, repeats its header on each one, or changes layout from one document to the next comes out as one set of rows in your columns. A single PDF can run to 5,000 pages, every page read and checked the same way as the first, and every row carries the file and page it came from.

Does it work on scanned PDFs?

Yes. Scanned PDFs, lower-quality scans included, and phone photos (JPG, PNG) go in the same batch as digital PDFs, under the same prompt. A value or a row the panel of AI agents cannot agree on is flagged as Review Needed, with what to check.

Can it combine many PDFs into one spreadsheet?

Yes. A batch of up to 6,000 files comes back as one sheet, with the same columns and formats for every document and the file and page each row came from. A file that fails is reported, never skipped in silence, and you are not charged for it.

What kinds of documents does it handle best?

Business documents, in any language or script: supplier statements, with each invoice on the statement as a row; invoices, purchase orders and delivery notes, with the numbers, quantities and amounts for matching; price and product lists, however long; reports, order confirmations, and packs that mix them.

Is it free?

Every account gets 50 free pages every month, with full functionality and no credit card to start. Above that, credits start at $12 for 100 pages. You pay for pages processed and nothing else. There is no plan to size and no monthly charge: a quiet month costs nothing, and credits bought for a busy month keep for 18 months.

Is there a PDF data extraction API?

Yes. Node.js and Python SDKs, and a REST API from any language, run the same extraction as the web app and return the same Excel, CSV or JSON, Review Needed included, with documentation a coding assistant can build a working integration from. Your own agent can run it too, through the MCP server or the skill it installs from this site. Same account, same balance, same results in your dashboard.

Extract your first PDFs free

50 free pages every month: no subscription, no credit card. Start with a few of this month's PDFs, and save the prompt so next month's come back in the same columns.