Publication:
Automatic Turkish text categorization in terms of author, genre and gender

Loading...
Thumbnail Image

Date

Institution Authors

Item type:Person,

Advisor

item.page.editor

Editor

Department

Journal Title

Journal ISSN

Volume Title

Publisher

SPRINGER-VERLAG BERLIN

DOI

Research Projects

Organizational Units

Journal Issue

Abstract

In this study, a first comprehensive text classification using n-gram model has been realized for Turkish. We worked in 3 different areas such as determining the identification of a Turkish document's author, classifying documents according to text's genre and identifying a gender of an author, automatically. Naive Bayes, Support Vector Machine, C 4.5 and Random Forest were used as classification methods and the results were given comparatively. The success in determining the author of the text, genre of the text and gender of the author was obtained as 83%, 93% and 96%, respectively.

Description

Journal or Series

NATURAL LANGUAGE PROCESSING AND INFORMATION SYSTEMS, PROCEEDINGS

ISSN

0302-9743

ISBN

3-540-34616-3

Rights

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By

Related Patent

Related Goal

0

Views

0

Downloads