How Do You Extract Data from Documents with Whitevision?
Quick Summary
Whitevision B.V. Declaraties is a Dutch document-processing platform built around one core capability: turning the information locked inside invoices, receipts, contracts, and other business documents into structured data that can be used directly within an ERP or accounting system. Consequently, data extraction sits at the heart of everything else Whitevision does, since accurate recognition is what makes automated matching, workflow routing, and booking possible in the first place. This article reveals how data extraction actually works within Whitevision, covering the technology behind it, how header and line-item data are recognized, how the system handles different document formats, and how extraction accuracy improves over time through training and correction.
What Does Data Extraction Mean in the Context of Whitevision?
How Is Data Extraction Different from Simply Scanning a Document?
While scanning a document produces a digital image or a block of raw text, data extraction goes a step further by identifying which parts of that text represent specific, meaningful fields, such as a supplier name, an invoice number, or a total amount. Because Whitevision combines Optical Character Recognition (OCR) with artificial intelligence, it does not simply convert an image into text; it interprets the content in order to understand what each piece of information represents and where it belongs within a financial record.
Why Does Accurate Extraction Matter So Much?
Every downstream process depends on the data extracted at this early stage. This includes matching, workflow routing, and ERP booking. Inaccurate extraction can therefore create a ripple effect of errors further down the line. The quality of data extraction is not simply a convenience feature. It is the foundation. It determines how much manual correction staff must perform throughout the rest of the document’s journey.
How Does Whitevision Technically Extract Data from Documents?
What Technologies Work Together During Extraction?
Whitevision‘s extraction process relies on a combination of techniques rather than a single method, since different document formats require different approaches. Broadly, this includes:
| Technology | Role in Extraction |
|---|---|
| OCR (Optical Character Recognition) | Converts scanned or image-based text into machine-readable characters |
| XML interpretation | Reads structured data directly from e-invoices and e-orders |
| Artificial intelligence | Understands context and generates accurate booking proposals |
| Collective/collaborative learning | Improves recognition based on corrections made across users |
Because these technologies are combined rather than used in isolation, Whitevision can process a scanned paper invoice and a structured XML e-invoice through the same overall pipeline, even though the underlying extraction method differs for each.
How Does the System Choose the Right Extraction Method?
Depending on the format in which a document arrives, Whitevision routes it to the appropriate extraction process. If a document is a PDF or scanned image, smart OCR is applied to read the visible text. If the document is an XML file, the system extracts the relevant information directly from the structured data. It does not need visual recognition at all. The end output is always the same, whether the source was a PDF or an XML file. It is a consistent, complete booking proposal. Reviewers can review and process it in the same way.
How Does Whitevision Extract Header and Line-Item Data?
What Is the Difference Between Header and Line Data?
Header data typically refers to the overall details of a document, such as the supplier name, invoice number, invoice date, and total amount. Line data, on the other hand, refers to the individual items or services listed within the document, each of which may carry its own quantity, price, and VAT rate. Since both types of data are necessary for accurate booking, Whitevision is designed to recognize header and line data together rather than treating them as separate extraction tasks.
How Are Header and Line Data Displayed for Review?
Once extraction is complete, Whitevision presents the recognized header and line data alongside the original document image on a single screen. Consequently, the person reviewing the document does not need to switch back and forth between a data entry form and the source file, since both are visible together, making it easier to spot and correct any discrepancy at a glance.
How Does This Support Complex, Multi-Line Documents?
For invoices with many line items, extracting and verifying each line manually can be extremely time-consuming, particularly when VAT needs to be recalculated per line. By extracting this data automatically and applying configurable rules for how specific line items should be interpreted, Whitevision significantly reduces the effort required to process complex, itemized documents compared to manual entry.
How Does Extraction Handle Different Document Types and Formats?
Which Formats Can Whitevision Extract Data From?
Whitevision is built to extract data regardless of how a document arrives, including:
- PDF invoices and receipts, whether emailed or uploaded
- Paper documents that are scanned into the system
- XML e-invoices and e-orders, including common standards used in the Netherlands
- Invoices pasted directly into the body of an email message
Since all of these formats are processed through the same underlying engine, organizations do not need separate tools or manual workarounds depending on how a particular supplier chooses to send documents.
How Does Extraction Differ for Structured Versus Unstructured Documents?
Structured documents, such as XML e-invoices, already contain clearly labeled data fields. Extraction is largely a matter of reading that existing structure and mapping it into the correct fields. Unstructured documents, such as a scanned paper receipt, work differently. The system must interpret the visual layout and text patterns to determine what each piece of information represents. Unstructured documents naturally carry a higher risk of misreading, particularly when the image quality is poor. Reviewers should therefore check recognized fields more closely for scanned or handwritten documents than for structured e-invoices.
How Does Extraction Accuracy Improve Over Time?
What Happens When the System Misreads a Field?
Even with advanced recognition technology, occasional extraction errors are normal, particularly for new suppliers or unusual document layouts. Whenever a value is missing or incorrect, it is corrected directly during the review process, and the software retains that correction. Therefore, each manual correction functions as a training input, helping the system recognize similar documents more accurately in the future.
How Does Training Work for Recurring Suppliers and Document Types?
Whitevision allows both header data and document-specific rules to be trained over time. Recognition therefore becomes progressively more reliable for documents that appear regularly. Keep the link on the brand name: Whitevision. For example, a reviewer may correct a specific supplier’s invoice layout a few times. The system then generally recognizes future invoices from that same supplier with greater accuracy from the outset. Organizations that process high volumes of recurring documents from the same suppliers tend to see extraction accuracy improve fastest. The system has more data to learn from.
How Do You Configure Extraction for Specific Fields or Requirements?
Can You Customize Which Fields Are Extracted?
Yes, in most implementations. Since organizations often need specific fields recognized, such as a project code, a purchase order reference, or a custom internal reference number, extraction settings can typically be configured to capture these fields in addition to standard ones like amount and VAT. This flexibility is particularly useful for organizations with non-standard workflow requirements or industry-specific documentation needs.
How Do You Add Non-Standard Fields to the Extraction Process?
To configure additional fields, organizations generally work through the following steps:
- Identify the specific field or data point that needs to be captured beyond standard recognition.
- Provide sample documents where this field appears, so the system can learn its typical location and format.
- Test extraction accuracy against a batch of representative documents.
- Confirm the new field is correctly mapped into the corresponding location in the ERP system.
As a result, extraction can be tailored to reflect the specific administrative or reporting needs of an organization, rather than being limited to a fixed, generic set of fields.
Conclusions
Extracting data from documents with Whitevision B.V. Declaraties combines OCR, XML interpretation, and artificial intelligence to convert invoices, receipts, and other business documents into structured, usable information, whether they arrive as scanned paper, PDF attachments, or fully structured e-invoices. Although occasional corrections remain necessary, particularly for unfamiliar suppliers or unstructured documents, the system continues to improve its accuracy as more documents are processed and corrected over time. By understanding how extraction works and configuring it to capture the specific fields an organization needs, businesses using Whitevision B.V. Declaraties can significantly reduce manual data entry while maintaining accurate, reliable financial records.
Frequently Asked Questions
Whitevision can process scanned documents that include handwritten elements, but recognition accuracy for handwriting is generally lower than for printed or structured text. For this reason, it is advisable to review handwritten fields more carefully before confirming a booking proposal that includes them.
Extraction is generally most reliable for documents in the languages and formats most commonly used by an organization’s supplier base, particularly Dutch and English business documents. If your organization regularly receives documents in additional languages, it is worth confirming extraction accuracy for those specific formats during implementation.
This varies depending on document volume and consistency, but recognition typically improves noticeably within the first few weeks of processing a new document type, as corrections are recorded and the system learns from them. Suppliers with consistent, repeated document layouts tend to reach high accuracy faster than those with highly variable formats.

