Stylometric Profiling of Turkish Texts: Joint Estimation of Author, Region, Age and Genre

relationships.isProjectOf

relationships.isJournalIssueOf

Abstract

Authorship identification seeks to determine the writer of a text by analyzing distinctive linguistic and stylistic features. These characteristics may vary across dimensions such as region, age, and genre. Identifying an author’s stylistic fingerprint is essential in plagiarism detection, digital forensics, and computational linguistics. In this study, the authorship features of Turkish columnists were analyzed using Artificial Neural Networks (ANN), Support Vector Machines (SVM), and decision tree algorithms (J48 and Random Forest). Sixteen stylometric indicators were selected through the Zemberek natural language processing library and evaluated across six distinct datasets. The proposed system allows flexible parameter adjustment through a graphical interface and exports results in ARFF format for reproducibility. Experimental results demonstrated that Random Forest achieved the highest overall accuracy, particularly in regional and age-based datasets, with F-measures reaching up to 0.91. The accuracy rates were 73% for regional classification, 55% for genre classification, and 62.5% for age-based classification. The findings confirm that combining statistical learning with stylometric analysis provides a robust framework for Turkish authorship attribution, paving the way for future studies employing deep learning and transformer-based models.

Description

Keywords

Random Forest, Computer Science, Natural Language Processing, Turkish, Artificial Intelligence

Fields of Science

Citation

WoS Q

Scopus Q

Volume

14

Issue

1

Start Page

288

End Page

298
Google Scholar Logo
Google Scholar™
OpenAlex Logo
OpenAlex FWCI
0.00