Data Processor transformation processes
unstructured and semi-structured file formats in a mapping. We can configure it
to process HTML pages, XML, JSON, and PDF documents. We can also convert
structured formats such as ACORD, HIPAA, HL7, EDI-X12, EDIFACT, AFP, and SWIFT.
For example, if we have customer invoices in Microsoft Word files, we can
configure a Data Processor transformation to parse the data from each word file
and extract the customer data to a Customer table and order information Orders
table.
The Data Processor
Transformation has the following
options:
Parser converts source documents to XML. The
output of a Parser is always XML. The input can have any format, such as text,
HTML, Word, PDF, or HL7.
Serializer converts an XML file to an output
document of any format. The output of a Serializer can be any format, such as a
text document, an HTML document, or a PDF.
Mapper converts an XML source document to
another XML structure or schema. You can convert the same XML documents as in
an XMap.
Transformer modifies the data in any format.
Adds, removes, converts, or changes text. Use Transformers with a Parser,
Mapper, or Serializer. You can also run a Transformer as stand-alone component.
Streamer splits large input documents, such as
multi-gigabyte data streams, into segments. The Streamer processes documents
that have multiple messages or records in them, such as HIPAA or EDI files.
In this blog, we
will see how to extract data from a PDF Document
and create a XML file using Data Processor Transformation. The source documents
that have fixed page layout like bills, invoices and account statements can be
parsed using positional format to find the data fields. An anchor is a signpost that you place in a
document, indicating the position of the data. The most commonly used anchors
are called Marker and Content anchors. These anchors are often used as a pair:
Marker anchor labels a location in a document.
Content anchor retrieves text from the
location. It stores the text that it extracts from a source document in a data
holder.
I have a PDF document with employee data as shown below:
FirstName: Chris
Lastname: Boyd
Department: HR
StartDate:
2009-10-11
You will also need
an XML Schema Definition (XSD) file which contains target XML schema. XSD file
looks like this:
<?xml
version="1.0" encoding="Windows-1252"?>
<xs:schema
attributeFormDefault="unqualified"
elementFormDefault="unqualified"
xmlns:xs="http://www.w3.org/2001/XMLSchema">
<xs:element
name="EmpPdf">
<xs:complexType>
<xs:sequence>
<xs:element
name="FirstName" type="xs:string" />
<xs:element
name="LastName" type="xs:string" />
<xs:element
name="Department" type="xs:string" />
<xs:element
name="StartDate" type="xs:date"
/>
</xs:sequence>
</xs:complexType>
</xs:element>
</xs:schema>
Using the above
data, I have explained in my video how to create a schema object and Data
Processor Transformation and use it in a mapping.
* In case of any
questions, feel free to leave comments on this page and I would get back as
soon as I can.