TY - GEN
T1 - GAIS
T2 - 15th IEEE International Conference on Knowledge Graphs, ICKG 2024, Co-located with 24th IEEE International Conference on Data Mining, ICDM 2024
AU - Rustamov, Zahiriddin
AU - Zaitouny, Ayham
AU - Damseh, Rafat
AU - Zaki, Nazar
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Instance selection (IS) is a crucial technique in machine learning that aims to reduce dataset size while maintaining model performance. This paper introduces a novel method called Graph Attention-based Instance Selection (GAIS), which leverages Graph Attention Networks (GATs) to identify the most informative instances in a dataset. GAIS represents the data as a graph and uses GATs to learn node representations, enabling it to capture complex relationships between instances. The method processes data in chunks, applies random masking and similarity thresholding during graph construction, and selects instances based on confidence scores from the trained GAT model. Experiments on 13 diverse datasets demonstrate that GAIS consistently outperforms traditional IS methods in terms of effectiveness, achieving high reduction rates (average 96%) while maintaining or improving model performance. Although GAIS exhibits slightly higher computational costs, its superior performance in maintaining accuracy with significantly reduced training data makes it a promising approach for graph-based data selection. Code is available at https://github.com/zahiriddin-rustamov/gais.
AB - Instance selection (IS) is a crucial technique in machine learning that aims to reduce dataset size while maintaining model performance. This paper introduces a novel method called Graph Attention-based Instance Selection (GAIS), which leverages Graph Attention Networks (GATs) to identify the most informative instances in a dataset. GAIS represents the data as a graph and uses GATs to learn node representations, enabling it to capture complex relationships between instances. The method processes data in chunks, applies random masking and similarity thresholding during graph construction, and selects instances based on confidence scores from the trained GAT model. Experiments on 13 diverse datasets demonstrate that GAIS consistently outperforms traditional IS methods in terms of effectiveness, achieving high reduction rates (average 96%) while maintaining or improving model performance. Although GAIS exhibits slightly higher computational costs, its superior performance in maintaining accuracy with significantly reduced training data makes it a promising approach for graph-based data selection. Code is available at https://github.com/zahiriddin-rustamov/gais.
KW - data reduction
KW - graph attention networks
KW - instance selection
KW - machine learning
UR - https://www.scopus.com/pages/publications/86000216324
UR - https://www.scopus.com/pages/publications/86000216324#tab=citedBy
U2 - 10.1109/ICKG63256.2024.00046
DO - 10.1109/ICKG63256.2024.00046
M3 - Conference contribution
AN - SCOPUS:86000216324
T3 - Proceedings - 2024 IEEE International Conference on Knowledge Graph, ICKG 2024
SP - 309
EP - 316
BT - Proceedings - 2024 IEEE International Conference on Knowledge Graph, ICKG 2024
A2 - Chen, Huajun
A2 - Fensel, Anna
A2 - Zhu, Xingquan
A2 - Wattenhofer, Roger
A2 - Wu, Xindong
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 11 December 2024 through 12 December 2024
ER -