It is said that the world knowledge is in the Internet. Scientific knowledge is inbooks, journals and conference proceedings. To cope with the huge amount ofinformation clever algorithms are needed. They are filtering, sorting and ultima-tely mining the information, improving as they get more data.Acommon technique is to mine the text from the publications. But publicationsinclude more information than the their text. The position of a word gives cluesabout its meaning. Additional images either supplement the text or offer proof toa proposition. Tables only form semantic units when read in rows and columns.To deal with the additional information, classic text mining techniques have to becoupled with spatial data and image data.For this thesis a framework was developed that allows the analysis of layoutinformation in scientific documents. This framework has been used for threecase studies. The first one allows the automatic extraction of images and theirannotation in the paper. The second one refines that approach as images arefurther classified into semantic categories based on their content. The third casestudy examines the use of tables in this context. They all discover knowledgethat would not have been visible through classical text mining and give hard evi-dence to the hypothesis that using layout does indeed improve the possibilitiesof text mining.
Brigitte Mathiak
Bildsuche Layoutanalyse PDF Tabellenerkennung