Hybrid feature selection approach to identify optimal features of profile metadata to detect social bots in Twitter

Eiman Alothali, Kadhim Hayawi, Hany Alashwal

Research output: Contribution to journalArticlepeer-review

12 Citations (Scopus)

Abstract

The last few years have revealed that social bots in social networks have become more sophisticated in design as they adapt their features to avoid detection systems. The deceptive nature of bots to mimic human users is due to the advancement of artificial intelligence and chatbots, where these bots learn and adjust very quickly. Therefore, finding the optimal features needed to detect them is an area for further investigation. In this paper, we propose a hybrid feature selection (FS) method to evaluate profile metadata features to find these optimal features, which are evaluated using random forest, naïve Bayes, support vector machines, and neural networks. We found that the cross-validation attribute evaluation performance was the best when compared to other FS methods. Our results show that the random forest classifier with six optimal features achieved the best score of 94.3% for the area under the curve. The results maintained overall 89% accuracy, 83.8% precision, and 83.3% recall for the bot class. We found that using four features: favorites_count, verified, statuses_count, and average_tweets_per_day, achieves good performance metrics for bot detection (84.1% precision, 81.2% recall).

Original languageEnglish
Article number84
JournalSocial Network Analysis and Mining
Volume11
Issue number1
DOIs
Publication statusPublished - Dec 2021

Keywords

  • Bot detection
  • Feature selection
  • Supervised learning
  • Twitter

ASJC Scopus subject areas

  • Information Systems
  • Communication
  • Media Technology
  • Human-Computer Interaction
  • Computer Science Applications

Fingerprint

Dive into the research topics of 'Hybrid feature selection approach to identify optimal features of profile metadata to detect social bots in Twitter'. Together they form a unique fingerprint.

Cite this