Method Article

Eye-Tracking Based Measurement Framework For Quantifying Visual Attention In Screen-Based Interactive Digital Art Environments

DOI:

10.3791/70801

June 5th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This protocol describes a comprehensive framework for quantifying visual attention in interactive digital art. It details the system architecture for synchronizing eye-tracking data with interaction logs, the method for dynamic Area of Interest (AOI) mapping, and the analytical pipeline for linking attention dynamics to user experience outcomes.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This article presents a gaze-enabled measurement framework designed to quantify visual attention during interactive digital art experiences and to link attention dynamics to experiential outcomes. Unlike traditional static art viewing, interactive digital art involves dynamic media, multi-source information, and continuous feedback loops, which challenge standard eye-tracking methodologies. The proposed system integrates eye-tracking acquisition, interaction logging, timestamp synchronization, dynamic Area of Interest (AOI) mapping, and metric computation into a unified pipeline. The protocol outlines the development of a reproducible data workflow that aligns gaze data with specific interaction events, enabling precise calculations of attention allocation, switching costs, and exploratory entropy. The experimental design involves both free exploration and goal-directed tasks, as demonstrated in a study with 37 participants. Results from applying this protocol indicate high perceived user-friendliness and satisfaction, with measurement modeling supporting a two-construct model of usability and satisfaction. Furthermore, outcome modeling confirms that perceived usability acts as a foundational predictor for overall user satisfaction, while the single-case demonstration validates the pipeline's capacity to quantify context-aware attention shifts. This protocol provides researchers and designers with a rigorous, reproducible toolset for analyzing how interaction design elements—such as feedback latency and guidance intensity—shape user attention and subsequent subjective experiences in immersive digital environments.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Interactive digital art encompasses a broad spectrum of modalities, from immersive spatial installations to screen-based interactive platforms1,2. While fully immersive setups represent the frontier of this evolution, screen-based interactive art remains a foundational medium that urgently requires rigorous, standardized evaluation methodologies prior to unconstrained spatial expansion. This modality effectively redefines the audience's role from passive observation to active co-creation, where participation directly shapes the artwork's meaning3. Classic frameworks in art psychology, notably the Vienna Integrated Model of Art Perception (VIMAP), offer a sophisticated lens for understanding the interplay between top-down cognitive control and bottom-up perceptual processing during aesthetic encounters. However, while these models provide a robust foundation for static observation, extending their latest applications to interactive digital environments requires advanced measurement techniques. Translating these theoretical constructs into interactive applications necessitates real-time physiological metrics to objectively capture the continuous feedback loop between user input and system response. Consequently, the transition from static viewing to active interaction introduces significant complexity in evaluating user experience. Current mainstream evaluation methods in the field primarily rely on post-event questionnaires, retrospective interviews, and on-site observations regarding viewers' viewing habits, sequence, and dwell time4. These are often supplemented by coarse-grained behavioral metrics like the total duration of stay and click frequency5. While such methods are effective at capturing general preferences and overall satisfaction, they fundamentally struggle to track the real-time, micro-level attention shifts that occur during the experience. Unlike retrospective questionnaires that rely on delayed recall and are therefore susceptible to cognitive bias, the proposed eye-tracking framework provides continuous, real-time, and objective physiological data. Furthermore, compared to coarse behavioral metrics such as total dwell time or superficial click frequencies, this gaze-enabled methodology captures micro-level cognitive processes—such as moments of hesitation, visual search strategies, and immediate perceptual reactions to system feedback. By doing so, it provides actionable, granular evidence for iterative interaction design that traditional evaluation methods fundamentally lack6. This technology offers process-level metrics for understanding attention allocation, maintenance, and shifting in interactive art, providing objective physiological metrics of the user's cognitive state7,8,9. By analyzing when, where, and for how long a user looks at specific elements, researchers can provide objective support for exhibition optimization and interactive design. However, implementing eye-tracking in the context of interactive digital art experiences presents multiple, non-trivial technical and methodological challenges that distinguish it from standard usability testing or static image viewing.

Interactive environments often involve dynamic media, multiple sources of information, and complex multi-step tasks. The visual stimuli are not fixed; they evolve based on the user's input. The constant changes in visual stimuli and system feedback cause rapid attention shifts between moving objects and interface areas. This dynamism makes traditional, static, interface-based segmentation and metric interpretation prone to distortion10,11. For instance, an Area of Interest (AOI) that is relevant at one moment may disappear or change its semantic meaning the next, rendering static heatmaps insufficient.

Additionally, real-world interactions frequently involve natural physical behaviors such as head movements, occlusions, and varying distances from the screen, leading to calibration drift12,13. These fluctuations in tracking quality introduce missing values and noise, reducing cross-subject comparability and the stability of statistical testing. Standard laboratory protocols often require rigid stabilization (e.g., chin rests) that detracts from the naturalistic experience required for art appreciation14. Therefore, an effective protocol must balance rigorous data quality control with the ecological validity of the user's posture. By utilizing remote screen-based eye-tracking, the proposed framework permits natural, unconstrained torso and head micromovements, serving as a pragmatic baseline for capturing aesthetic reactions without obtrusive physical restraints.

Critically, the temporal synchronization between eye movement data and interaction logs directly impacts causal interpretation. In a static viewing task, the stimulus onset is the only relevant time marker. In interactive art, every touch, gesture, or gaze trigger initiates a new temporal event reference. Without a unified time benchmark and event alignment mechanism, it becomes difficult to determine whether attention changes are triggered by system feedback or driven by individual exploration strategies15. To comprehensively evaluate digital art, researchers typically rely on two distinct data sources: interaction logs (capturing user inputs and system states) and physiological signals (such as gaze tracking for visual attention). However, these modalities are often analyzed in isolation. Integrating them early in the analytical pipeline is conceptually and practically vital because they are inherently complementary: interaction logs provide the precise contextual 'trigger' (the what and when of system feedback), while continuous gaze data reveals the corresponding cognitive 'reaction' (the where and how of attention allocation). Without a unified time, benchmark, and multimodal event alignment, it becomes difficult to determine whether an attention shift is triggered by system feedback (e.g., a visual ripple effect) or driven by individual exploration strategies. By fusing these data streams, this framework provides the theoretical benefit of closing the causal loop between interactive design features and user cognition.

Therefore, advancing eye-tracking research for interactive art requires more than just capturing fixation points; it necessitates rigorous quality control, synchronized alignment, and reproducible data-processing workflows. The limitations of existing solutions further underscore this research motivation. Existing eye-tracking studies in artistic contexts predominantly focus on static viewing tasks or treat textual descriptions and tour guides as external variables, lacking integration with interactive mechanisms16,17. While evidence suggests that information presentation methods and guidance strategies can significantly alter viewers' attention pathways and focal points, most findings remain at the level of single-intervention correlations, failing to explain how interactive feedback loops shape attention strategies over the duration of an experience.

Conversely, while interactive art practices have developed prototype systems that use gaze as a natural interaction input—emphasizing the transformation of viewers from passive observers into active participants—these efforts often focus on installation demonstrations or technical validation18,19. They frequently lack unified attention metrics, replicable experimental paradigms, and closed-loop research models that establish testable connections between attention dynamics and experiential outcomes. This gap hinders the development of transferable design principles that can reliably predict how a specific design choice (e.g., the speed of a visual transition) will impact the user's sense of immersion or usability.

Critically, existing literature lacks a cohesive framework that fundamentally aligns visual attention with real-time interactive systems. While advancements in dynamic AOI mapping and gaze-event synchronization have been explored in general HCI contexts, they are rarely integrated into the fluid and highly variable environment of digital art. The fundamental research gap lies in the inability of current single-modality evaluations to explain the causal loop between system feedback and user cognition. By early integration of multimodal data—specifically, fusing continuous gaze coordinates with millisecond-accurate interaction logs—researchers can transcend basic usability metrics. This integration provides the theoretical benefit of attributing specific micro-level attention shifts (e.g., search strategies vs. stabilized focus) directly to specific interaction design choices, transforming the experiential flow from a subjective feeling into a quantifiable cognitive state.

To address these gaps, this protocol proposes an eye-tracking-driven attention measurement framework specifically constrained to screen-based interactive digital art. While the underlying technical pipeline shares structural similarities with generic human-computer interaction (HCI) testing, this framework uniquely integrates non-utilitarian aesthetic tasks (e.g., free visual exploration of generative media) alongside goal-directed actions. This integration explicitly distinguishes the fluid, exploratory nature of art appreciation from the standard utilitarian approach to software evaluation. Theoretically, this study operates at the intersection of interactive digital art, HCI, and eye-tracking analysis. It addresses both the persistent challenge of objectively characterizing the continuous experiential flow in digital art evaluation and determining the correlation between interactive input and attentional output.

The research scope positions this study not only on the experiential context of interactive digital art but also on integrating eye-tracking data with interaction logs to establish a systematic evidence chain—from experiential input to attention metrics and ultimately to outcome variables. At the system level, we establish a reproducible data pipeline that integrates eye movement capture, interaction log recording, and temporal synchronization into a unified architecture. This framework standardizes preprocessing, region mapping, and attention metric generation while providing traceable data structures and analytical scripts.

The conceptual framework (Figure 1) elucidates how interaction design influences subjective outcomes by altering viewers' visual attention processes. The core logic treats interaction elements (such as feedback latency and guidance intensity) as controllable inputs, visual attention dynamics as key process variables, and experience outcomes (immersion, usability, satisfaction) as outputs. Individual differences are incorporated as moderating factors. For experimental design, interactive tasks are developed that balance free exploration with goal-oriented objectives. At the modeling level, we correlate attention dynamics with outcome variables to examine the mechanisms by which interaction design elements influence user experience through visual attention. This protocol provides detailed methodological steps for replicating this framework, ensuring that the complex interplay among user, system, and art can be measured with scientific rigor.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This research was performed in compliance with the institutional guidelines of the Faculty of Innovation and Design at the City University of Macau. All participants provided informed consent prior to data collection.

1. System development and architecture configuration

  1. Define the research workflow
    1. Adopt the four-stage closed-loop process as shown in Figure 2: System Development, Eye-Tracking Experiment, Analytics, and Outcome Modeling.
    2. Ensure that the workflow facilitates iterative refinement, where data from the analytics phase can inform adjustments in the system development for future iterations.
  2. Configure the system architecture
    1. Construct the system using the pipeline architecture illustrated in Figure 3, comprising five distinct modules: Experience (Tasks), Eye-Tracking, Logs, Synchronization (Sync), and Analytics.
    2. Experience module: Develop the interactive art interface to serve as the unified entry point. Configure it to generate two parallel data streams:
      1. Eye-tracking data stream: Output raw gaze coordinates and timestamps from the device.
      2. Interaction log stream: Document user actions (clicks, gestures) and system state changes (feedback triggers, scene transitions).
    3. Implement a User Datagram Protocol (UDP) script within the synchronization module to send event markers from the interaction rendering engine to the eye-tracker API at the exact millisecond a visual stimulus alters, or a user input is registered.
    4. Evaluate offset stability by running a 60-second synchronization test loop prior to the experiment. Verify that the standard deviation (SD) of the latency between the interaction clock and the eye-tracker clock remains strictly under 6.5 ms (as per Table 1). If the SD exceeds 6.5 ms, reset the network connection and recalibrate.
    5. Aligned data storage: Configure the database or file system to store aligned data where each data point contains synchronized gaze coordinates, event flags, and AOI indices.
  3. Establish the system runtime loop
    1. Implement the runtime loop as depicted in Figure 4.
    2. Program the sequence to strictly follow: Start Session → Calibration → Load Scene + AOIs → Real-time Loop (User Interaction / Gaze Capture / Render Feedback).
    3. Ensure the data pipeline (Write Logs → Timestamp Sync → Fixation Parsing → AOI Mapping) operates immediately following data acquisition to facilitate rapid quality checks.

2. Stimulus and interaction design

  1. Select interaction modalities
    1. Choose the interaction modalities based on the taxonomy. Select between Gaze-based Interaction, Touch/Gesture-based Interaction, or Multi-user Collocated Interaction depending on the research question.
    2. For this protocol, implement a combination of gaze and touch inputs to adhere to the principle of minimizing input difficulty while maximizing output interpretability.
  2. Define the interaction modules and trigger events.
    ​NOTE: Here, an "interaction module" is operationally defined as a discrete functional unit of the digital artwork (e.g., a specific interactive visual effect), while a "trigger event" refers to the specific user input or system state change that activates said module at a millisecond-accurate timestamp.
    1. Program the specific interaction modules defined in Table 2.
    2. Ambient attractor: Set the trigger to scene entry or idle state for ≥ 3.2 s. Define the expected attention response as wider exploration and higher gaze dispersion.
    3. Gaze highlight: Configure the trigger as gaze dwell on an object AOI ≥ 420 ms. Ensure this triggers a visual highlight to reinforce focus.
    4. Touch ripple: Implement a visual ripple effect at coordinates (x,y) upon a tap on the canvas event. Record saccade amplitude changes following the wavefront.
    5. Guided cue: Program a visual cue to appear for 1.4–2.1 s at a pre-set AOI. Measure the Time to First Fixation (TTFF) using this cue.
    6. Choice gate: Create a decision point where the user selects 1 of 3 options within 7.5 s. Use this to measure dwell time on the chosen option versus rejected options.
  3. Design Areas of Interest (AOIs)
    1. Create static and dynamic AOIs as shown in Figure 5.
    2. Define static AOIs for fixed user interface (UI) elements (e.g., menus, borders).
    3. Define dynamic AOIs for moving visual elements (e.g., floating orbs, particle systems). Ensure the AOI coordinates update in the log file at the system refresh rate.
    4. Implement semantic tagging for AOIs (e.g., Target Object, Distractor, Instruction Text) to facilitate the calculation of semantic credibility during analysis.

3. Participant recruitment and screening

  1. Recruitment strategy
    1. Recruit participants from a relevant demographic (e.g., university students for digital art studies) to ensure a baseline of digital literacy.
    2. Obtain informed consent and brief participants on the nature of the eye-tracking procedure.
  2. Screening criteria
    1. Screen for visual acuity. Verify that participants have normal or corrected-to-normal vision.
    2. Exclude individuals with ocular diseases (e.g., strabismus, severe astigmatism) that interfere with standard eye-tracking calibration.
    3. Calibration check: During the initial setup, perform a 5 or 9-point calibration. Exclude participants who cannot achieve a mean calibration error of ≤ 0.85° or who exhibit persistent tracking loss (valid sample ratio < 85%).
    4. Record demographic data, including age, gender, handedness, and self-reported digital art exposure (see Table 3 parameters).

4. Experimental procedure

  1. Physical setup
    1. Conduct the experiment in a quiet room with stable ambient lighting. Eliminate direct sunlight or strong reflections on the screen.
    2. Seat the participant at a fixed viewing distance (e.g., 60–70 cm) appropriate for the specific commercial screen-based eye-tracking unit.
    3. Adjust the chair height so the participant's eyes are level with the center of the screen.
  2. Session execution
    1. Calibration: Execute the calibration routine. Verify the accuracy. If the drift is > 1.10°, perform a recalibration (Drift check).
    2. Practice trial: Administer a standardized practice session (approx. 2 minutes) to familiarize the participant with the interaction protocols (e.g., how to trigger the Gaze Highlight).
    3. Task execution: Present the two task types in a counterbalanced order to minimize sequence effects:
      1. Free exploration task: Instruct the participant to "Freely explore the artwork and interact with elements that interest you."
      2. Goal-directed task: Instruct the participant to "Find specific hidden elements" or "Complete the interaction sequence to unlock the final visual."
    4. Data recording: Ensure the system automatically records both the eye-tracking stream and interaction logs with unified timestamps throughout the tasks.
    5. Post-task questionnaires: Immediately following each task block, administer the experience scales (Immersion, aesthetic pleasure, usability, satisfaction).

5. Data processing and analysis pipeline

  1. Data preprocessing and quality control
    1. Refer to Figure 6 for a comprehensive overview of the gaze processing pipeline. This visual schematic delineates the sequential stages required to transform raw gaze coordinates into interpretable analytical data: initial signal quality control (including blink interpolation and spatial filtering), velocity-based parsing of eye movements into fixations and saccades, dynamic mapping of gaze points to active Areas of Interest (AOIs), and the final computation of quantitative attention metrics.
    2. Handle blinks and signal dropouts using a standardized interpolation and filtering pipeline. Identify blinks and signal loss segments where pupil diameter equals zero or tracking validity flags are invalid.
    3. For short gaps (≤75 ms), apply a linear interpolation algorithm to estimate the missing gaze coordinates. For gaps exceeding 75 ms, leave the data unmapped to prevent artificial artifact generation. After interpolation, apply a moving-average filter (window size = 3 samples) to smooth minor high-frequency noise.
    4. Detect and eliminate spatial outliers defined as gaze points falling outside the physical screen boundaries or demonstrating biologically impossible saccadic velocities (> 1000o/s).
    5. Calculate the final dropout fraction; exclude sessions where the unrecoverable loss exceeds 12% of the total task duration.
    6. Apply trial-level drift correction. If a systematic offset was detected during the validation check immediately prior to a specific task block, apply a linear coordinate correction vector to the raw gaze data of that specific trial to compensate for minor head postural shifts.
  2. Fixation and saccade parsing
    1. Apply the Identification by Velocity Threshold (I-VT) parsing algorithm. Classify raw data points as saccades if the eye movement velocity exceeds 30°/s; otherwise, classify them as fixations.
    2. Apply a minimum fixation duration threshold of 80 ms to filter out micro-saccades and artifact noise.
    3. Classify eye movements into fixations (stable gaze) and saccades (rapid movement). Generate a fixation table and saccade table.
  3. Temporal alignment and AOI mapping
    1. Align the gaze event timestamps with the interaction log timestamps using the UDP event markers as the absolute reference anchor.
    2. Correct for any inherent hardware transmission delays by subtracting the average offset recorded during the initialization test loop from all gaze timestamps.
    3. Refer to Table 4 for a schematic example of the synchronized data rows illustrating this integration at the analytical level. Note how this matrix fuses continuous gaze coordinates, discrete interaction event flags, and dynamic AOI indices at millisecond resolution to form the foundational dataset for subsequent metric computation.
    4. Implement Dynamic AOI Mapping to overcome the limitations of static spatial partitioning in fluid digital art environments.
    5. Map each fixation coordinate to the active AOIs specifically updated for that exact millisecond timestamp using a point-in-polygon algorithm.
    6. NOTE: This ensures that even as visual elements continuously move or change state, the gaze data is correctly attributed to the intended interactive object rather than a fixed, obsolete screen location.
    7. Assign priority to overlapping AOIs using a strict z-buffer hierarchical rule. Map the raw (x,y) gaze coordinate to the bounding box array updated at that specific millisecond timestamp.
    8. If the coordinate falls within multiple active bounding boxes, assign the fixation exclusively to the AOI with the highest rendering layer index (the foremost visual element). Discard overlapping data from lower layers to prevent double-counting.
    9. Calculate the percentage of unmapped fixations. Ensure this remains ≤ 0.18 (18%) to verify the completeness of the AOI definitions.
  4. Metric computation
    1. Compute metrics across three dimensions as shown in the Figure 1 framework:
    2. Calculate TTFF, defined strictly as the elapsed time from the onset of a specific visual stimulus to the start of the first fixation within its corresponding target AOI. Compute Dwell Time (total fixation duration within an AOI) and Revisit Rate.
    3. Compute Transition Matrices to quantify attention switching. Calculate the transition probability PAB of moving from AOI A to AOI B by dividing the number of gaze transitions from to by the total number of transitions originating from A. Calculate the Switch Rate by dividing the total number of inter-AOI transitions by the valid task duration.
    4. Calculate Global Gaze Entropy (H) using Shannon's Information Theory formula to quantify the randomness of the visual exploration. Formulate the equation as follows:
      Entropy formula, H = -Σ(pi*log2(pi)), diagram; information theory, probability analysis. (1)
      where the variables are defined as follows: : Global Gaze Entropy, representing the dispersion of visual attention. N: The total number of defined dynamic AOIs within the scene. pi: The probability of the gaze landing within a specific AOI . This is derived from the proportion of dwell time in AOI relative to the total valid task duration.
    5. Aggregate these metrics by task phase (Free Exploration vs. Goal-Directed).
  5. Statistical modeling
    1. Assemble the final dataset merging Computed Metrics with the Survey Scores.
    2. Perform Confirmatory Factor Analysis (CFA) to validate the constructs of Usability and Satisfaction.
    3. Use Mixed-Effects Models or Structural Equation Modeling (SEM) to test the hypotheses (e.g., Interaction Latency predicts Attention Switching, Attention Switching predicts Usability).

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Sample characteristics and data quality

The application of this protocol recruited 43 students, of whom 37 were ultimately included in the final analysis (Inclusion rate: 86.0%). As detailed in Table 3, the exclusion of 6 participants resulted exclusively from adherence to data quality protocols: tracking loss/poor calibration (n=2), excessive missing samples (n=2), incomplete tasks due to self-reported motion sickness during the interactive phase (n=1), and v...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study establishes a validated eye-tracking measurement framework specifically tailored for interactive digital art. At the core, findings indicate that dynamic Area of Interest (AOI) mapping coupled with millisecond-level multimodal synchronization successfully quantifies micro-level visual attention shifts previously obscured in fluid media. Theoretically, this research bridges the conceptual gap between passive observation and active participation models by establishing functional usability as a quantifiable upstr...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no conflicts of interest to declare.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors thank the students of the Faculty of Innovation and Design at the City University of Macau for their participation in this study. We also extend our gratitude to the technical support staff for their assistance with the eye-tracking equipment setup. This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Eye TrackerTobii ABPro Nano60 Hz screen-based gaze capture; 5/9-point calibration
MonitorASUSVP249QGR23.8'', 1920×1080; viewing distance ~65 cm
Interaction PCDellOptiPlex 7010Windows 10; runs interface + logs
Sync ClockCustomv1.0Unifies gaze & interaction timestamps
Fixation ParserOpen-sourceI-VTFixation threshold 80–120 ms.
AOI MapperCustom2D Point-in-PolygonMaps gaze to static/dynamic AOIs

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Eye TrackingVisual AttentionInteractive Digital ArtAttention DynamicsGaze DataArea Of InterestInteraction LoggingUsability ModelingFeedback LatencyUser Satisfaction

Related Articles