Enhancing EEG-based Gaze Prediction with Transformers on EEGEyeNet
Open Access DepositedElectroencephalography (EEG) provides a non-invasive method for capturing brain activity, but its complex, high-dimensional nature poses significant challenges for accurate gaze estimation. Deep learning, particularly Vision Transformer (ViT) architectures, offers promising capabilities for gaze prediction using Electroencephalography (EEG) signals. We build upon the EEGViT model, adapting it for EEG-based regression tasks by optimizing its architecture and training process to improve gaze position predictions. Our experiments were conducted on the Absolute Position Task from the EEGEyeNet dataset, which aims to predict the spatial coordinates of a subject's gaze based solely on EEG data. The study investigates the impact of architectural modifications, including kernel size adjustments and dropout rate tuning, leading to an RMSE of 50.77 mm—improving model performance by ~2% compared to the original version of EEGViT-TCN. These results position the optimized model within 3% of the state-of-the-art accuracy. Additionally, we examine the potential effects of pretraining across multiple cognitive tasks, evaluating its impact on the accuracy of gaze predictions.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.