To main content

Task-Conditioned Next-Fixation Prediction in Assembly Tasks

Abstract

Unlike bottom-up attention models, computational models of task-driven human attention have been less explored. In this work, we study task-driven eye-fixation prediction during the assembly of simple structures with educational building blocks. We introduce a new first person vision dataset of an assembly task, consisting of synchronized video frames, eye fixations, and template images of the manipulated objects. Building upon probabilistic scanpath modeling frameworks, we propose a neural network architecture that integrates scene frames, template object features, and past fixation heatmaps to predict the next fixation as a probability distribution over the scene. We evaluate the performance of our model against commonly used saliency metrics and further perform an ablation study that investigates how fixation history and the inclusion of center bias affect performance. In addition, we introduce a task-specific hit-rate evaluation metric. Our results demonstrate that the predicted fixations follow the gaze pattern of the specific task setting and, further, highlight the importance of task-specific data for studying top-down attention, and potentially advancing robotic autonomy by enabling systems with task-constrained visual perception.

Category

Academic chapter

Language

English

Author(s)

Affiliation

  • SINTEF Digital / Mathematics and Cybernetics
  • Norwegian University of Science and Technology

Date

20.01.2026

Year

2026

Publisher

Springer

Book

Advances in Visual Computing: 20th International Symposium, ISVC 2025, Las Vegas, NV, USA, November 17–19, 2025, Proceedings, Part I

ISBN

9783032144928

Page(s)

322 - 333

View this publication at Norwegian Research Information Repository