Yayın:
An Integrated Architecture for Processing Business Documents in Turkish

dc.contributor.authorAdali, Serif
dc.contributor.authorSonmez, A. Coskun
dc.contributor.authorGokturk, Mehmet
dc.date.accessioned2026-06-27T13:09:29Z
dc.date.issued2009
dc.description.abstractThis paper covers the first research activity in the field of automatic processing of business documents in Turkish. In contrast to traditional information extraction systems which process input text as a linear sequence of words and locus on semantic aspects, proposed approach doesn't ignore document layout information and benefits hints provided by layout analysis. In addition, approach not only checks relations of entities across document for verifying its integrity, but also verifies extracted information against real word data (e.g. customer database). This rule-based approach uses a morphological analyzer for Turkish, a lexicon integrated domain ontology, a document layout analyzer, an extraction ontology and a template mining module. Based on extraction ontology, conceptual sentence analysis increases portability which requires only domain concepts when compared to information extraction systems that rely on large set of linguistic patterns.en
dc.identifier.endpage+
dc.identifier.isbn978-3-642-00381-3
dc.identifier.issn0302-9743
dc.identifier.startpage394
dc.identifier.urihttps://hdl.handle.net/20.500.14981/50523
dc.identifier.volume5449
dc.identifier.wos000265681200032
dc.language.isoeng
dc.publisherSPRINGER-VERLAG BERLIN
dc.relation.conference10th International Conference on Intelligent Text Processing and Computational Linguistics
dc.relation.ispartofCOMPUTATIONAL LINGUISTICS AND INTELLIGENT TEXT PROCESSING
dc.subjectINFORMATION EXTRACTION
dc.subjectComputer Science
dc.titleAn Integrated Architecture for Processing Business Documents in Turkish
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar