Açık Kaynak İstihbaratı (OSINT) İçin Türkçe İçerik Temelli Kişilik Özellikleri Tahmini
Thesis Type: Doctorate
Institution Of The Thesis: Gazi University, Bilişim Enstitüsü, Bilgisayar Eğitimi, Turkey
Approval Date: 2023
Thesis Language: Turkish
Student: MUHAMMED ALİ KOŞAN
Principal Supervisor (For Co-Supervisor Theses): Hacer Karacan
Co-Supervisor: Burcu Ayşen Ürgen
Open Archive Collection: AVESIS Open Access Collection
Abstract:This study was conducted on the prediction of personality traits of users through the content obtained from social media platforms using open source intelligence. To create the basic framework of the research; the life cycle of open source intelligence has been adapted to predict the personality traits of users on social media platforms. Two simultaneous studies were conducted using this lifecycle. First, personality traits were predicted from English contents on Twitter obtained using 13 negative-positive words. In the study, a dataset of English personality traits was created and then a balanced and high-performance personality trait prediction model was created. The content was analyzed based on the social media platform by using semantic data analysis in the system established for the prediction model. By using semantic structures on the analyzed data, a balanced and generalizable personality trait prediction model was developed with data-specific preprocessing steps. Secondly, a gamified web application software was developed with certain personality traits tests to create a personality traits dataset with Turkish content. Along with the promotional activities of this software, data collection process was started, and both the personality traits and social media contents of the users were collected. When the collected data is examined, it is seen that the BFI-10 personality trait test results and the available Twitter account content have a significant proportion. Therefore, the final dataset was created from the personality traits results of the BFI-10 test and the content obtained from Twitter accounts. In the obtained dataset, firstly, the model obtained on the English dataset was adapted. Afterwards, a comprehensive analysis was made with semantic preprocessing methods, vectorization methods and deep learning models. As a result, it has been observed that the results obtained despite the complex structure of the Turkish language offer balanced, generalizable, and remarkably good results.