Abstract
Understanding where humans interact with objects in egocentric video is essential for action anticipation and human machine collaboration. While gaze has proven beneficial as a supervisory signal in indoor environments, its effectiveness in outdoor agricultural settings remains unexplored. This thesis investigates whether gaze can enhance prediction of plant interaction hotspots in greenhouse environments by adapting CNN-LSTM architectures with attention consistency modules to integrate gaze supervision. Using a newly annotated dataset of 46 egocentric videos with synchronized gaze recordings captured during plant manipulation tasks, I evaluate model performance through classification metrics and spatial attention metrics (KLD, SIM, AUC-Judd). The results reveal that gaze heatmaps exhibit strong alignment with ground-truth interaction regions and differ significantly from center-biased baselines, suggesting gaze contains task-relevant spatial information. However, fundamental dataset limitations - including structural ambiguity of plant components, class imbalance, and insufficient visual separability - prevent model form leveraging this signal effectively during training. Edge detection and semantic segmentation analyses demonstrate that CNNs struggle to distinguish visually similar plant parts, leading to poor classification performance across all model variants. Consequently, this work does not provide conclusive evidence for the benefit of gaze supervision or knowledge transfer from indoor to agricultural environments. Nevertheless, the quantitative and qualitative analyses establish that gaze is informative for interaction prediction, and the identified dataset challenges provide crucial insight for advancing egocentric vision research in agricultural domains.