Invoice Data Extraction Logo
Invoice Data Extraction
Start Extraction
Pricing
Extraction Guide
API
Sign inCreate account
Sign inCreate account
Start Extraction
Pricing
Extraction Guide
API
  1. Home
  2. Articles & Analysis
  3. Tax & Compliance
  4. EU Country-of-Origin Reconciliation From Supplier Invoices

EU Country-of-Origin Reconciliation From Supplier Invoices

How EU importers assemble line-level country-of-origin data from supplier invoices for Intrastat, customs entry, FTA claims, and audit-ready defence records.

Published
Apr 25, 2026
Updated
Aug 18, 2026
Reading Time
31 min
Author
David Harding
Topics:
Tax & ComplianceEUWholesale Distributionimport compliancecountry of originpreferential originREXIntrastatcustomscountry breakdowncross-border APcontrol totalssupplier invoicesExcel

On this page

Country-of-origin reconciliation from supplier invoices means capturing origin evidence at supplier-invoice-line level, then using that dataset for Intrastat reporting, customs entries, preferential-origin claims, and audit defence. Four fields drive the work: the per-line country of origin, the verbatim origin-statement text printed on the invoice, the supplier's Registered Exporter (REX) number where applicable, and the HS chapter that governs the rule of origin. Under the EU REX system, statements on origin attached to shipments under €6,000 in value can be made out without REX registration; above that threshold the supplier must hold a REX number for the statement to be valid.

This is an importer-side problem, and it sits in a different place from the export-side commercial-invoice question that dominates the search results for this topic. The input is a backlog of incoming supplier PDFs, often from a mix of EU and non-EU manufacturers. The output is a portfolio spreadsheet keyed at the supplier-invoice-line level — not a single declaration, not a one-off filing, but a continuous dataset that has to stay current as new invoices arrive every week. The narrower question — how to break one mixed-origin invoice down by country without losing the tie-out to its total — is the same dataset filtered to a single document, so the work below answers both.

What makes the reconciliation worth doing is that the same field on the same invoice serves four different obligations on four different cadences. Intrastat is monthly statistical reporting on intra-EU goods movements. Customs entry is filed per shipment at the border. Preferential-rate claims under EU FTAs are made at entry but evidenced in continuous documentation that has to survive a verification request months or years later. Audit defence happens on demand, when a customs authority challenges a duty rate and asks the importer to substantiate the origin claim against the original invoice and supporting documents. One dataset, four uses, four different patterns of access.

The Four Fields You Need From Each Supplier Invoice

Reconciliation begins at the document. Every supplier invoice that lands in the AP inbox carries — or should carry — four pieces of origin information, each in a different part of the page, each required for a different downstream use. Capturing all four at the line level is what lets the same dataset answer Intrastat, customs, FTA, and audit questions later.

Per-line country-of-origin marking. This is the country code or country name printed against each individual goods line, often in a column adjacent to the description and HS code. On a clean supplier invoice the marking is a two-letter ISO country code (CN, VN, BD, TR) sitting next to a description and quantity. On a multi-line invoice covering goods manufactured in different countries, this varies line by line — apparel orders from a single supplier routinely mix garments cut in one country and finished in another, and each line has to record the country attributed to that specific item. The reconciliation captures origin at the line level, not the invoice level. Recording one country for a fifteen-line invoice, when the lines actually came from three, is the data-quality failure the reconciliation is meant to prevent.

The verbatim origin-statement text. Where the supplier is asserting preferential origin under an EU FTA, an origin statement appears somewhere on the document — usually as a footer paragraph below the totals, sometimes on a continuation page, occasionally on a separate cover sheet attached to the invoice. The wording is not free-form. Under the EU Registered Exporter system, Annex 22-07 of the Union Customs Code Implementing Regulation (EU) 2015/2447 sets the prescribed text the supplier must use when making out a statement on origin: a single paragraph naming the country of preferential origin and, where relevant, identifying the products and the supplier's REX number. The reconciliation captures that paragraph verbatim — exact characters, no paraphrase, no summary. When a customs authority later asks to see the origin proof, the question is whether the wording on file matches the prescribed form. A summary of the gist is not evidence.

The supplier's REX number authenticating the statement. Where the statement on origin appears, it normally cites the supplier's REX registration number — the link between the supplier's claim on the invoice and the European Commission's REX database, which any customs authority can query to confirm the supplier's registration is current. The number has a fixed shape (country code plus alphanumeric identifier) and almost always appears inside or immediately adjacent to the statement-on-origin paragraph. Above the €6,000 shipment threshold, a statement on origin without a valid REX number is not valid for preferential treatment, so capturing the number is a different kind of mandatory than capturing the statement text — both have to be present, and they have to agree.

The HS chapter that governs the origin rule. The supplier's declared HS code on each line — at minimum the four-digit heading, ideally the full eight-digit Combined Nomenclature subheading or the ten-digit TARIC code where the supplier provides it — determines which product-specific rule of origin in the relevant FTA applies. A garment of woven cotton fabric falling under HS 6203 carries a different origin test from a knitted shirt under 6105, and both differ from a leather handbag under 4202. The chapter is what ties a given line to the correct rule, so it belongs on the input row in its own column rather than buried inside a free-text description. Extracting the HS chapter and country of origin from the supplier invoice as paired fields, line by line, is what makes the rule lookup possible later.

The four fields belong on every line, not collapsed into a single per-invoice "origin field." Row granularity is the supplier-invoice line; an invoice with twelve goods lines produces twelve rows in the dataset, each carrying its own country, its own HS code, and where applicable its own pointer back to the same statement-on-origin text and REX number that govern the whole invoice.

The practical reality is that not every inbound invoice will carry all four cleanly. Origin is often missing on intra-EU supplier invoices that do not need it — a French wholesaler invoicing into the Netherlands has no commercial reason to print country of origin against goods of EU origin. The origin-statement text only appears when the supplier is asserting preferential origin; for non-preferential entries it is absent by design. The REX number is missing when the supplier has not registered, when the shipment falls below the €6,000 threshold, or when the supplier is operating under approved-exporter status with a different reference instead. The HS code is sometimes missing, sometimes given at four-digit heading level only, sometimes wrong. Reconciliation has to record absence as a value — null, not skipped — because the absence itself is information that drives a follow-up query to the supplier.

Which Origin Signal Wins When an Invoice Carries Several

One invoice often carries several country signals, and they do not all mean the same thing. Rank them before a value lands on a row, and record which tier the value came from, because the question at audit is never whether a country looks plausible — it is where that country came from.

Per-line origin marking is the top tier. A country code printed against the goods line, in the supplier's own origin column, is what the entry and the Intrastat row are built on. Where it is present, nothing below it overrides it.

The statement on origin and the documents behind it come next. A statement on origin scoped to specific products, a manufacturer's declaration, or a mill or processing certificate can establish origin for a line whose own marking is missing or contested. This tier is document evidence rather than invoice evidence, so the row needs the supporting-document reference alongside the country.

Shipping and customs paperwork travelling with the consignment is the third tier. Packing lists, EUR.1 certificates, and the export declaration frequently carry origin at a granularity the invoice page does not. Read them before concluding a line has no origin signal; a summarised invoice line is often supported by detail that lives on an attachment rather than on the invoice itself.

Ship-from and supplier address blocks are a weak fourth tier. Where goods were dispatched from is not where they were made. A consolidation warehouse in the Netherlands dispatching goods manufactured in Vietnam produces a Dutch ship-from address on a shipment of Vietnamese origin, and reading dispatch as origin is one of the more expensive errors available at entry.

A partner VAT or GST prefix is not origin evidence at all. The country prefix on a registration number identifies the supplier's country of registration — supplier identity, not the origin of the goods and not the country a charge pertains to. A French-registered supplier can invoice goods manufactured in Turkey and services performed in Spain from the same VAT number. Read the prefix as an identity cross-check when a supplier group holds several registrations, and never let it override a line-level marking.

Currency is not evidence. CHF or JPY on a line may narrow the field, but EUR carries no country information inside the Eurozone. Currency is a tie-breaker between two candidate readings, never a source of a country value.

The tier is a value on the row, not a judgement kept in an operator's head. Default to the highest tier the line actually supports, fall back only when nothing above it exists, and label the fallback so the row can be found again when the supplier finally answers the query.

Extraction Versus Allocation: Missing and Ambiguous Country Data

Two operations produce country-tagged rows that look identical in the spreadsheet and are not interchangeable.

Extraction pulls a country the document already carries: the per-line marking, the statement on origin, the manufacturer declaration behind the line. The evidence text points back to the PDF, and a reviewer can confirm the value against the document in seconds.

Allocation derives a country the document does not carry by applying a buyer-side rule: apportioning a freight, duty, or handling charge across the lines of a mixed-origin consignment by value, weight, or units, or splitting a supplier-level charge across markets on an approved percentage. The evidence is the rule name, and the supporting paper is the buyer's written allocation policy rather than the supplier's invoice.

The distinction carries more weight here than in a general cost breakdown, because the two kinds of field have different tolerances. A landed-cost row can be allocated: it is the buyer's own arithmetic, and a documented rule applied consistently is the standard it has to meet. A declared country of origin cannot be allocated at all — it is a property of the goods and their manufacturing history, and no buyer-side rule can produce one. Where the origin field is missing, the answer is a supplier query, never an apportionment. Conflating the two is how a country dataset loses its defensibility: the rows look the same, so an allocated value gets read later as an extracted one.

Four ambiguity cases cover almost everything that needs a decision.

Missing country. The document carries no usable signal for the line. The origin field stays empty and the row is flagged for the supplier query rather than filled with the nearest plausible country. On the cost side of the same row, allocate only where a written rule the controller has approved actually applies, and name that rule in the evidence column. Do not infer a country because one feels likely; that is the failure mode the unallocated bucket exists to prevent.

Mixed-origin lines. One line covers goods from several countries. Where a sub-breakdown exists — a packing list, a manufacturer declaration, a per-market schedule behind a summarised line — extract one sub-row per origin from it, each carrying the parent line ID so the split stays traceable to the line it came from. Where no sub-breakdown exists, the line cannot be split by inference: cost components can be apportioned under the written rule, but the origin field stays open and any value no extracted sub-row covers sits in the unallocated bucket until the supplier answers.

Supplier country versus goods origin. The supplier's registered country, its VAT prefix, and its ship-from address answer who issued the invoice and where the consignment left from, not where the goods were made. Where nothing line-level exists and a supplier-level value is the most defensible working assumption for a cost row, record it and label the evidence tier "supplier-level fallback", so every row resting on that assumption can be pulled back in one filter when better evidence arrives. A supplier-level fallback is never adequate for a declared origin on an entry.

Rounding residuals. Per-country apportionment leaves sub-cent and single-cent differences behind. Keep them on a named rounding line in the control-total carry-through, separate from the country rows, rather than smearing them across countries to make a total look tidy.

The unallocated bucket is a labelled row, not an absence. It carries what the dataset cannot yet place: missing evidence, ambiguous matches, mixed-origin lines awaiting a supplier answer, supplier-level charges with no rule applied. Surfacing an unallocated amount is more defensible than spreading it across country rows on a guess, and it is the row a reviewer should read first.

Allocation rules belong in writing. A buyer-side policy naming the rules, the rates, what they apply to, and who approved them is the supporting paper behind every allocated row. Without it, allocation drifts between operators and between periods, and the drift stays invisible until someone reconciles two quarters against each other.

Where extraction tooling populates the dataset, the unallocated bucket and the review flags should be visible outputs rather than silent omissions. A line with no origin evidence should come back marked unallocated with a note on what was missing, so the reviewer knows which decision is theirs to make.


Driver 1 — Intrastat Statistical Declarations

Intrastat is the EU's monthly statistical declaration on intra-EU goods movements, separated into dispatches (goods leaving the member state to another EU member state) and arrivals (goods entering from another EU member state). The country-of-origin field is one of the modernised statistical variables and is captured at the commodity-line level — the same granularity the reconciliation dataset already holds. That statistical reporting trail is separate from the buyer-side VAT treatment covered in reverse-charge VAT accounting for intra-EU supplier invoices.

Under Regulation (EU) 2019/2152, country of origin and partner VAT ID became mandatory Intrastat variables for dispatches. Arrivals are less uniform: many member states require origin as a national addition, but it is not a bloc-wide arrivals requirement. The reconciliation should therefore capture origin continuously, while the filing logic should check the national statistical authority rules for each member state.

For an EU importer, this means origin should be captured when the supplier invoice is processed, not only when a filing calendar says it is due. Thresholds and reporting frequency differ by member state; they determine when to file, not whether the line-level origin field should be retained. When a row is generated, it draws supplier ID, invoice number, line ID, the eight-digit Combined Nomenclature code, country of origin, partner VAT ID, net mass, and statistical value from the dataset. For German filers, that same row becomes the source for preparing Intrastat declarations from supplier invoices.

Driver 2 — Customs Entry Filings

On every import declaration, country of origin is a mandatory data element that determines the third-country duty rate applied to the goods. The commercial invoice from the supplier is the primary documentary source for that field. Whether the importer files the declaration directly or hands the paperwork to a customs broker invoice and entry data extraction workflow, the origin per line is drawn off the invoice and the supporting documents behind it. The reconciliation dataset is the importer's authoritative copy of that origin per line — the same source the broker would otherwise pull each time, held centrally, with a known provenance.

Two distinct origin questions sit on the same shipment, and the reader has to keep them separate. Non-preferential origin is the default origin used for tariff classification, trade-defence measures (anti-dumping duties, safeguard measures), and statistical purposes. The rules for determining non-preferential origin are set in the EU Union Customs Code and apply regardless of any FTA. Preferential origin is the qualification used to claim a reduced or zero duty rate under a specific FTA, governed by the product-specific rules in that agreement. Both are recorded against the same shipment but answer different questions: where the goods are deemed to come from for general tariff and trade-policy purposes versus whether they qualify for preferential treatment under a particular agreement. The next section covers preferential origin in detail; this section focuses on the non-preferential field that lands on every entry.

The Incoterm on the supplier invoice influences which party files the entry and bears the import-side costs but does not alter the origin determination. Origin is a fact about the goods, established by the manufacturing process and the rules that classify it; the contract between buyer and seller does not change that fact. The Incoterms wording on the commercial invoice matters for who is responsible for the entry, not for what origin is declared on it.

The reconciliation problem at customs entry is what to do when the supplier invoice is incomplete. Two patterns recur. First, the invoice carries no per-line origin marking at all — common with intra-EU suppliers and with non-EU suppliers who have not yet been pushed by their importer customers to print origin per line. The importer has to go back to the supplier for the missing data before the entry can be filed correctly. Second, the invoice gives a single origin at invoice level for a shipment that actually contains goods of mixed origin, and the importer has to determine line-by-line origin from the supplier's underlying records. Both of these are recurring queries that consume time on every shipment unless the dataset records the answer once, with a pointer to where it came from, and reuses it. Capturing the result on the master row — origin, source of the answer, date confirmed, who confirmed it — is what prevents the same gap being chased twice for the same supplier.

Customs entry data also feeds the wider import-cost picture. The duty rate determined by the origin field lands on the freight or customs invoice the AP team receives a week or two after the goods themselves, and the same origin reconciliation row is what lets the cost team allocate that duty correctly back to the inventory line. For UK-importing readers handling UK postponed VAT accounting on imports, the origin field is also what reconciles the C79 import VAT certificate against the original supplier invoice and the broker's entry — same field, queried in a third place. Norwegian AP teams face a similar month-end evidence trail when reconciling TVINN declarations to supplier invoices and import VAT.

Driver 3 — Preferential-Origin Claims Under EU FTAs

A preferential rate of duty under an EU FTA requires proof of preferential origin. The proof can take several forms depending on the agreement and the supplier's status, and the reader's portfolio dataset has to record which form attaches to each line and where it lives. When an authority later asks why the importer paid zero duty instead of the third-country rate, the answer is a specific document type, a specific supplier reference, and a specific piece of text — not a general assertion that the goods qualify.

The REX statement on origin is the most common form of proof under modern EU FTAs, including the agreements with Canada (CETA), Japan (EU-Japan EPA), Vietnam (EU-Vietnam FTA), the United Kingdom (the Trade and Cooperation Agreement), and others. The supplier makes a statement on origin on the commercial invoice or on another commercial document related to the shipment, using the prescribed wording set out in Annex 22-07 of Implementing Regulation (EU) 2015/2447. According to European Commission Access2Markets guidance on the REX system, under the EU Registered Exporter system, for shipments with a value of less than 6,000 euro the statement of origin can be made out without the obligation to register; above that threshold the exporter must hold a REX number. The threshold is per shipment, not per supplier or per year, and it is the dividing line between an invoice that is self-evidencing for preferential treatment and one that needs the supplier's REX registration to back it.

EUR.1 movement certificates are the older paper-form proof, still in use under some agreements — notably under Pan-Euro-Mediterranean Convention rules with countries that have not transitioned to REX. The EUR.1 is issued by the customs authority of the export country, not by the supplier, and it travels separately from the commercial invoice as a stamped document. This is the fundamental contrast for the importer's reconciliation. A statement on origin is supplier-issued, sits on the invoice, and authenticates itself through the REX number where one is required. EUR.1 is government-issued, is a separate document, and authenticates itself through the customs authority's stamp and serial number. The dataset has to record which one applies per shipment and where the document is filed, because at audit they will be retrieved from different places. EUR.1 vs invoice declaration is the live distinction every importer working under a mix of FTAs has to keep straight.

The long-form invoice declaration available to approved exporters is the third form, used under some FTAs and pre-dating the REX system. It is restricted to suppliers who hold approved-exporter status with their own customs authority — a status granted on application and subject to ongoing audit. The declaration uses different prescribed wording from a REX statement on origin and cites the supplier's approved-exporter authorisation reference in place of (or, depending on the agreement, alongside) a REX number. From the importer's side, the practical difference is what appears on the invoice: a different paragraph of text, a different reference number, a different validity question to test if the claim is challenged. The reconciliation dataset records the proof type as a categorical field — REX statement on origin, EUR.1, approved-exporter invoice declaration, none — because the four are not interchangeable at audit.

The apparel and textile case is where the abstract framework hits the ground hardest, and it is worth working through because it is the case that makes the discipline worth running. HS Chapters 50 to 63 — covering everything from raw silk through woven and knitted fabrics to finished garments — carry product-specific rules of origin under the Pan-Euro-Mediterranean Convention used by EU FTAs across the Euro-Med region. The rules typically require manufacturing operations beyond simple assembly to confer originating status. Depending on the chapter, the test is fabric-forward (origin is conferred where fabric is produced from yarn) or yarn-forward (origin is conferred where yarn is produced from fibre and then woven or knitted into fabric in the same originating country). What this means for the importer is that a supplier in a PEM partner country can only claim preferential origin if the manufacturing pattern actually meets the test — fabric or yarn sourced from within the cumulation zone, processed in qualifying countries, with the finishing operations performed in the country claiming origin.

That sets the verification problem. An importer of apparel from a non-EU supplier cannot rely on the supplier's statement alone if the pattern of origin claims looks inconsistent with the manufacturing economics — for example, if a supplier in one country is claiming origin on garments made from fabrics that the country does not produce, or claims a rule-compliant cumulation that the supplier cannot evidence. Verifying the claim means holding manufacturer-level documentation correlating fabric or yarn origin to the finished goods on the invoice. The reconciliation dataset is what makes that verification structured rather than ad hoc. Every line carrying a textile HS code from a non-EU supplier sits in a queryable view, and the supporting documentation column either holds a reference or flags the line for follow-up before the next claim period. US-bound shippers working the same goods face a parallel exercise on the export side, where pulling fiber composition, knit-versus-woven, and Chapter 61/62 fields off the apparel commercial invoice is what feeds CBP entry classification. The USMCA equivalent of this verification problem — pulling yarn-forward proof off Mexican and Canadian apparel manufacturer invoices — runs the same fabric-and-yarn correlation question against a different agreement's rule set.

Driver 4 — Audit Defence When a Customs Authority Challenges the Claim

When a customs authority challenges an origin claim, the importer is the party on the hook. The supplier's statement on origin is the starting point, not the end of the matter. If the authority is not satisfied with the supporting evidence, the duty differential is recoverable from the importer along with interest and, depending on jurisdiction and the nature of the finding, potential penalties. The reconciliation dataset's value is that every line points back to the supplier invoice, the verbatim statement on origin, the REX number current at the time of the claim, and any external documentation supporting the position.

Retention obligations attach to the evidence, not just to the spreadsheet summary, and they outlive the claim. HMRC requires importers to retain records supporting an origin claim for four years from the date of the claim under the UK's preferential-rate framework. EU member states under the Union Customs Code apply equivalent retention windows — commonly three years from the end of the year in which the customs declaration was accepted, with national variations that extend the period in some member states. Confirm the specific retention period with the relevant authority for the jurisdictions the importer operates in, and keep the underlying supplier PDFs for the full period.

The hard evidence question is material-to-finished-good correlation: can the importer trace fabric, yarn, components, or other inputs through to the finished goods in a way that matches the rule of origin claimed? If the rule is yarn-forward and the importer cannot show where the yarn came from, the claim fails regardless of which customs authority is auditing it.

The operational implication is concrete: each row needs a source-file reference, the verbatim statement-on-origin text, the REX number, and any supporting-document reference for manufacturer declarations, bills of materials, supplier audits, or certificates of analysis. A dataset that records the claim but not the path back to the evidence is a list of assertions without a paper trail. At audit, that distinction is the difference between a defended claim and a recoverable duty.


The Spreadsheet Shape: One Row Per Supplier Invoice Line

Before the columns, decide what a row represents. The grain sets the row count, determines what still points back at the PDF, and decides which downstream questions the file can answer at all. Three grains are available.

One row per country. The portfolio rolls up to one row per origin country, with value, mass, and duty exposure summed across every line that carried it. This suits management reporting, supplier-mix analysis, and a quick read on how much of the book depends on a single origin. It is the wrong grain for anything that has to point back at a document, because the aggregation drops the line reference.

One row per invoice line, keyed by country. Each supplier-invoice line produces one row per origin it carries: a single-origin line produces one row, a three-origin line produces three, each keeping the parent line ID and its own evidence. Row counts inflate, but aggregating up to a country total is a pivot away, and every row still holds a thread back to the page it came from.

A country summary with an exceptions list. A summary table for the reporting reader, plus a line-grain list of the rows that need a supplier query or an allocation decision. This works while most lines carry a clean marking and only a minority need handling; it stops working when the exceptions stop being a minority.

Intrastat, customs entry, FTA claims, and audit defence all ask document-level questions, so line grain is the only one of the three that serves all four. Pick the grain that survives the hardest downstream question and aggregate upward from it. Starting aggregated and disaggregating later loses the line reference and the evidence text, which are exactly what the hardest question asks for.

The four drivers come together in a single row schema at that grain. One row per supplier-invoice line, sixteen columns, designed so that any of the four downstream uses can be served by selecting a subset of the same columns. The schema below is what the reconciliation dataset has to carry to satisfy Intrastat, customs entry, FTA claims, and audit defence from one source.

  • Supplier ID
  • Supplier invoice number
  • Invoice date
  • Line ID (sequential within the invoice)
  • HS code (eight-digit Combined Nomenclature where the supplier provides it)
  • HS chapter (derived from the HS code; carried as its own column for origin-rule lookup)
  • Declared country of origin (ISO two-letter code, per line)
  • Preferential-origin claim flag (yes / no)
  • Proof type (REX statement on origin, EUR.1, approved-exporter invoice declaration, or none)
  • REX number (where applicable)
  • Statement-on-origin verbatim text (captured exactly as it appears on the invoice)
  • Origin evidence text (the snippet the country value was read from — the origin-column entry, an address line, a schedule row)
  • Evidence tier (line marking, statement on origin, supporting document, supplier-level fallback, or allocated)
  • Source file reference (supplier PDF filename and page number)
  • Supporting-document reference (manufacturer declaration filename, BOM reference, or null)
  • Retention-until date

The two evidence columns are what make the rest of the dataset cheap to check. A reviewer holding the snippet the value was read from can confirm a country tag in seconds; without it, every challenged row costs a fresh read of the PDF. Keep the snippet short and recognisable — enough to find on the page, not the whole line — and keep the tier beside it, because the snippet records what was read and the tier records how much weight it carries.

Each driver queries a different subset. Intrastat draws supplier ID, invoice number, line ID, HS code, country of origin, plus quantities and values held alongside the row. Customs entry draws HS code, country of origin, and the proof-type field that determines whether the line is being entered at the third-country rate or under a preferential claim. The FTA claim itself draws proof type, REX number, statement-on-origin verbatim text, and HS chapter — the chapter being what ties the line to the correct product-specific rule. Audit defence draws source file reference, supporting-document reference, retention-until date, and the two evidence columns — the snippet the value was read from and the tier it came from — the columns that turn the dataset back into the underlying paper trail.

The honest part of the spreadsheet question is the population problem. Capturing sixteen fields per line across a backlog of supplier PDFs, then keeping the dataset current as new invoices arrive every week, is the bottleneck nearly every importer hits when they sit down to actually build this. Manual entry is the obvious answer and the worst answer. The verbatim origin-statement text and the REX number are both fields where transcription error is invisible until audit — a missing accent, a wrong digit, a paraphrased clause — and they are also exactly the fields the audit goes after first. A dataset that is 90% accurate on the easy columns and unreliable on the hard ones is worse than no dataset, because it generates confidence in claims that will not survive verification.

This is where the tool that builds the row directly from the PDF earns its place. It is built to extract line-level country-of-origin and origin-statement text from supplier invoices in one pass — taking a batch of up to 6,000 files in a single job — into the column shape this section has just defined. The interaction model is a single prompt over the file batch, the same prompt for ten invoices or ten thousand, returning a structured Excel, CSV, or JSON file with one row per line item. The four extraction targets — per-line country of origin, the verbatim footer text of the statement on origin, the REX number, and the HS code — are exactly the kind of fields the prompt should request explicitly, with the source-file reference and page number emitted alongside each row for the evidence chain audit defence requires. For an importer who is already maintaining the reconciliation logic in a spreadsheet, the population step is the part that does not have to be manual.

The same supplier-invoice-line dataset also supports adjacent workflows: landed cost accounting for wholesale distributors and freight and duty allocation in manufacturing. A country-keyed Excel breakdown of one mixed-origin invoice is the same rows filtered to a single document, so it costs nothing extra once the dataset exists. The reconciliation work is done once when the supplier invoice is processed; downstream teams query the same columns instead of rebuilding them.

Control Totals: Tying the Rows Back to the Invoice

A country dataset that does not reconcile to the invoices behind it is a working file that has not finished its work. The rule is simple to state and load-bearing for everything above it: for each invoice, the grand total equals the sum of the country-keyed row values plus the unallocated bucket. The same rule applies separately to the net subtotal, tax, freight, duty, discount, and rounding. Each control element has its own tie-out, and each one either holds or does not.

Every amount on the invoice lands in exactly one place — a country row, an allocation row, the unallocated bucket, or a named control line for freight, discount, or rounding. Anything unplaced is a reconciliation gap; anything placed twice is a double count. The tie-out surfaces both before the dataset feeds a declaration, which is the only point at which finding them is cheap. Line counts tie out on the same principle: an invoice with twelve goods lines produces twelve parent line IDs, and a split into sub-rows per origin has to reconcile back to the parent line it came from.

Tax deserves a separate pass, because it is the component most likely to break. A multi-country invoice commonly carries several tax types and rates, and the tax total is the sum across all of them. Each row's tax has to tie back not only to the invoice's tax total but to the sub-totals by tax type where the invoice carries them: a line at German 19% VAT does not reconcile against an Italian 22% IVA line, and collapsing the two into a single figure hides an error that only appears when the return is filed. Import VAT and duty run on their own timetable — determined by the origin field, but landing on the broker or freight invoice weeks after the goods — so tie them out against the entry rather than treating them as a component of the supplier invoice total.

Without the tie-out, the dataset is a set of plausible rows. With it, the same rows are an auditable working paper, which is what the four drivers actually need from it.

Why Country Is a Dimension You Can Defend

Country is a defensible reporting dimension only where the rows trace back to records: supplier invoices, statements on origin, manufacturer declarations, entry documents, contracts, or an approved allocation policy. Keep the country value, the evidence snippet, the evidence tier, the allocation rule where one applies, and the control-total tie-out together on the row, so the dataset can be reviewed months later without rebuilding the invoice from scratch.

IRS instructions for Form 8975 show the same principle operating at group-reporting scale: country-by-country reporting rests on tax-jurisdiction data the business can support from its underlying records. The obligation is a different one — a US group's country-by-country report is not an EU importer's origin claim — but the standard the data has to meet is identical. A country figure is defensible when the path back to the document that produced it is still on the row.

Extract invoice data to Excel with natural language prompts

Upload your invoices, describe what you need in plain language, and download clean, structured spreadsheets. No templates, no complex configuration.

Exceptional accuracy on financial documents
Parallel processing — large batches complete in minutes
50 free pages every month — no subscription
Any document layout, language, or scan quality
Native Excel types — numbers, dates, currencies
Files encrypted and auto-deleted within 48 hours
Start Extracting FreeView Pricing
Continue Reading

Related Articles

Explore adjacent guides and reference articles on this topic.

Norway TVINN Import VAT Reconciliation for AP

Reconcile Norwegian supplier invoices, forwarder bills, and TVINN declarations to import VAT, MVA-meldingen codes, and a month-end Excel workpaper.

Prepare the Irish Intrastat RPF CSV from Invoices

Build the Irish Intrastat working paper from invoices: fields lifted from invoices, enriched from product and shipment data, ready for RPF CSV upload.

Cyprus VAT Return From Supplier Invoices & Credit Notes

Prepare a Cyprus VAT return from supplier invoices and credit notes: purchase-register fields, 1-11B box mapping, and credit-note workflow before TFA filing.

Back to Articles & Analysis

Invoice Data Extraction

The AI-native automation platform for high-accuracy invoice extraction

Platform

  • Start Extraction
  • Home
  • Pricing
  • API
  • Python SDK
  • Node.js SDK

Solutions

  • Invoice to Excel
  • Invoice OCR Software
  • Bank Statement Converter
  • Receipt OCR
  • Utility Bill Extraction
  • Payroll Data Extraction
  • PDF Data Extraction

Resources

  • Articles
  • Contact

Trust & Security

  • Security
  • Subprocessors
  • AI Data Use

Legal

  • Terms of Service
  • Data Processing Addendum
  • Privacy Policy
  • Refund Policy
  • US State Privacy Rights
  • EEA/UK Privacy Rights
English
Sign inCreate account

© 2026 Invoice Data Extraction — DEH Technologies LLC

Secure by Design. Your data is never used for AI training.