IntroductionEye movements provide a window into overt spatial attention during visual examination. Across domains ranging from reading to medical diagnosis, patterns of fixations and saccades reveal how attention is distributed across space and how information is sampled over time. In radiology, eye tracking research has shown that expert observers exhibit distinctive search behaviors compared with novices, including more systematic scanpaths and adaptive modulation of fixation duration and saccade length in response to task demands and lesion presence1,2,3. In volumetric imaging, experts have been described as adopting drilling or scanning strategies that differ in how gaze is distributed across spatial and depth dimensions and are associated with performance differences4. Despite extensive documentation of these expertise related differences, the short timescale directional dynamics that contribute to these patterns remain largely unmeasured. Most standard gaze metrics quantify temporal allocation or movement magnitude. Fixation duration, saccade amplitude, velocity, and acceleration capture how long and how far the eyes move, but not how viewing direction changes within brief temporal intervals. Computational approaches to scanpath analysis have emphasized sequential structure and similarity. Methods such as ScanMatch and MultiMatch align fixation sequences to compare spatial and temporal order5,6, and network based analyses transform scanpaths into graph representations whose topology relates to reading skill and comprehension7. These approaches demonstrate that ordering and transition structure carry meaningful information. However, they operate at the level of fixation sequences and aggregate across the fine scale directional corrections that occur within and between fixations. No existing metric directly quantifies how much viewing direction changes within short temporal windows, a property that may distinguish sustained local inspection from broader spatial transitions. Short timescale components of eye movement have been shown to be functionally significant. Even during intended fixation, the eyes exhibit microsaccades, drifts, and corrective movements whose directional properties relate to attentional allocation and perceptual processing8,9. Microsaccades are systematically biased toward attended locations10,11 and modulated by task demands8, indicating that directional adjustments at fine temporal scales are behaviorally meaningful. Yet these dynamics are typically collapsed within fixation labels or averaged across larger segments in conventional analyses12. A metric that quantifies directional variability at sub fixation timescales may therefore capture an aspect of visual examination that is not reflected in duration or amplitude based measures. Across domains, theoretical distinctions between focal and ambient processing13,14, global versus local search in radiology1,4,15, and exploration versus exploitation in visual foraging16 share a common structural contrast: concentrated inspection within local regions versus broader sampling across distant locations. Operationalization of these constructs often relies on heuristic thresholds or dispersion based segmentation procedures17. While useful, such approaches do not derive regime structure directly from the directional dynamics of gaze trajectories. Directional variability, defined as the variability in angular change between consecutive gaze displacement vectors, provides a geometric description of how viewing direction evolves over time. Behaviorally, high directional variability corresponds to frequent directional corrections within constrained spatial regions, as occurs during sustained inspection, whereas low directional variability reflects directional persistence during transitions between spatially separated locations. Because it is computed independently of movement magnitude, directional variability isolates changes in orientation rather than distance traveled. In eye movement research, measures of scanpath curvature and structural complexity have been used to characterize viewing strategies and individual variability18,19. However, these measures are typically computed over extended sequences or entire scanpaths. In contrast, windowed directional variability operates at sub fixation to fixation timescales, enabling temporal segmentation of examination dynamics during ongoing viewing. In this study, we test whether directional variability reveals stable regime structure that is not present in conventional magnitude based metrics. We demonstrate three findings. First, directional variability exhibits a stable bimodal distribution across temporal windows from 10 to 50 ms and across both medical image interpretation and general viewing tasks, whereas velocity, acceleration, and step length variability show asymmetric spike plus tail distributions without comparable mode separation. Second, regimes defined by directional variability correspond to systematic differences in directional entropy, angular reversal density, and return rate that persist after matching windows by spatial dispersion. Third, at the study level, directional regime composition provides incremental predictive value beyond comprehensive magnitude based baselines, improving prediction of scanpath topology metrics and time to report. Together, these findings indicate that directional variability represents a distinct dimension of gaze dynamics that captures how frequently viewing direction changes within short temporal windows. By identifying regime structure directly from directional statistics rather than from heuristic dispersion thresholds or predefined task phases, this approach provides a data driven geometric framework for characterizing examination dynamics. Such directional markers may support expertise research, training assessment, and human AI interaction, provided that future work establishes their relationship to performance outcomes (Fig. 1).Fig. 1Full size imageConceptual roadmap of directional variability and emergent examination regimes. (1) Directional variability defined: Directional variability is computed as the standard deviation of the angular change between consecutive gaze displacement vectors within a short temporal window. Low directional variability reflects directionally persistent movement with small angular changes, whereas high directional variability reflects frequent directional corrections. (2) Emergent regimes: When computed across time, windowed directional variability exhibits a two-mode distribution, yielding low- and high-variability regimes defined directly from the empirical density rather than from heuristic dispersion thresholds. (3) Examination dynamics within a lesion: Illustrative examples from chest radiograph interpretation show how these regimes correspond to distinct patterns of gaze organization during local inspection. Low velocity combined with low directional variability reflects directionally persistent sampling within a region, whereas low velocity combined with high directional variability reflects sustained local inspection characterized by frequent directional adjustments. Images are schematic examples for conceptual illustration.ResultsWe tested whether directional variability exhibits distributional and behavioral structure distinct from magnitude-based gaze metrics across two datasets: REFLACX radiograph interpretation (\(n = 3052\) studies) and GazeBase general viewing tasks (\(n = 12,334\) recordings). First, we examined whether directional variability displays stable bimodal structure across temporal windows from 10–50 ms. Bimodality was quantified using two-component Gaussian mixture modeling with separation strength assessed by Cohen’s d, and compared against velocity, acceleration, and step length variability under identical windowing procedures. Second, we evaluated whether the two modes identified by directional variability, corresponding to high and low directional variability regimes, reflect differences in how gaze direction evolves. Specifically, we tested for regime differences in directional entropy, angular reversal density, and return rate, and assessed whether these differences persist after matching windows by spatial dispersion. Third, we tested whether study-level proportions of directional regimes predict scanpath complexity metrics, including transition entropy and unique spatial cells visited, and time to report in REFLACX, beyond mean velocity alone. Nested regression models were used to quantify incremental variance explained (\(\Delta R^2\)) .Directional variability exhibits robust bimodal structure across temporal scales and tasksWe tested whether directional variability exhibits stronger bimodal separation than magnitude-based metrics. For each metric (directional variability, velocity, acceleration, and step length), we computed windowed variability across temporal windows from 10 to 50 ms, yielding large distributions of windows in each dataset. In REFLACX, 10 ms windows yielded 21.04 million samples, decreasing to 4.20 million for 50 ms windows due to fewer long windows fitting within fixed recording durations. In GazeBase, analysis of each task yielded 1.12 to 7.40 million windows across window sizes (Fig. 2). Directional variability exhibited substantially stronger separation than magnitude-based metrics across both datasets (Fig. 2a). By conventional thresholds (\(d > 0.8\) large, 0.5–0.8 moderate, 0.2–0.5 small), directional variability consistently achieved very large separation, whereas magnitude-based metrics typically showed small to moderate effects. In REFLACX, directional variability showed very large separation (mean Cohen’s \(d = 3.40\), range: 2.41–4.54 across window sizes), whereas velocity, acceleration, and step length variability showed smaller separation (mean \(d = 0.85\), range: 0.62–1.01; mean \(d = 0.75\), range: 0.64–0.85; mean \(d = 0.49\), range: 0.34–0.60, respectively). In GazeBase, directional variability similarly showed very large separation across reading, gaming, and video tasks (mean \(d = 3.65\)–3.82, ranges: 2.81–4.94), while magnitude-based metrics remained small to moderate (velocity mean \(d = 0.29\)–0.31; acceleration mean \(d = 0.36\)–0.42; step length mean \(d = 0.30\)–0.35). On average, separation strength for directional variability was approximately four to ten times larger than for magnitude-based metrics. In contrast to directional variability, magnitude-based metrics exhibited asymmetric spike-plus-tail distributions (Fig. 2b), characterized by a dominant peak near zero and heavy upper tails. This asymmetry was reflected in large tail-to-median ratios (\(p_{99}/p_{50}\)). In REFLACX, mean \(p_{99}/p_{50}\) was 22.63 for velocity, 11.19 for acceleration, and 29.08 for step length, compared with 1.58 for directional variability. In GazeBase, mean \(p_{99}/p_{50}\) ranged from 35.72–52.12 for velocity, 24.79–33.84 for acceleration, and 60.11–111.22 for step length, compared with 1.43–1.44 for directional variability across tasks. These extreme ratios indicate that a small fraction of windows contain very large values, producing heavy-tailed distributions unsuitable for bimodal characterization. This pattern reflects amplitude heterogeneity across windows rather than distinct behavioral regimes. If bimodality reflects scale-dependent but stable regime structure rather than a window-size artifact, the threshold separating the two modes should remain interpretable across temporal scales. We therefore examined the regime boundary separating the two directional variability modes, which we defined as the minimum density point between mixture components (Fig. 2c). In REFLACX, the regime boundary increased from \(6.40^\circ\) (10 ms windows) to \(12.30^\circ\) (50 ms windows), consistent with longer windows aggregating more directional change. In GazeBase, boundary ranges showed similar scaling patterns: \(5.33^\circ\)–\(16.50^\circ\) for gaming, \(5.03^\circ\)–\(16.50^\circ\) for reading, and \(5.01^\circ\)–\(15.90^\circ\) for video. Despite this predictable window-size scaling, the two-mode structure persisted across all temporal scales and tasks, with separation remaining large at every window size (minimum Cohen’s \(d = 2.41\) across all conditions). To determine whether the observed bimodality reflects pooling across heterogeneous observers rather than within-subject dynamics, we quantified between-subject variance relative to total variance in GazeBase using the subject variance ratio (SVR), defined as the variance of per-subject means divided by the variance of the pooled distribution. Across tasks and window sizes, SVR values were low (mean SVR = 0.013, range: 0.004–0.046), indicating that only approximately 1–4% of total variability in directional variability is attributable to stable between-subject differences. The vast majority of variance therefore arises within individuals across time. This finding demonstrates that the bimodal structure is not an artifact of mixing distinct subject-level distributions, but instead reflects dynamic alternation between regimes within individual viewing sequences. Together, these results show that directional variability yields robust two-mode separation across tasks and datasets, whereas magnitude-based metrics primarily reflect asymmetric amplitude distributions without comparable mode separation(Fig. 3). We next test whether the regimes defined by directional variability correspond to differences in scanpath directional organization beyond spatial extent 3)Fig. 2Full size imageDirectional variability exhibits robust two-mode separation across datasets, unlike magnitude-based metrics. Top row: GazeBase. Bottom row: REFLACX. (a), Separation strength. Cohen’s d for two-component Gaussian mixture fits to windowed variability measures, averaged across temporal windows (10–50 ms). Error bars indicate standard deviation across window sizes. Directional variability shows substantially stronger component separation than velocity, acceleration, or step length variability in both datasets. (b), Distribution asymmetry. Tail-to-median ratio (\(p_{99}/p_{50}\); log scale), summarizing spike-plus-tail structure. Magnitude-based metrics exhibit pronounced asymmetry, whereas directional variability shows comparatively low asymmetry under identical windowing. (c), Boundary scaling. Estimated regime boundary (valley between mixture components) as a function of temporal window size (10–50 ms). In GazeBase, boundaries are shown separately for gaming, reading, and video tasks; in REFLACX, the boundary is shown for chest radiograph interpretation. Although boundary location scales with window size, two-mode separation persists across tasks and datasets .Fig. 3Full size imageWindowed variability distributions across metrics and temporal scales. Histograms show the distribution of windowed standard deviation values computed at 10, 20, and 30 ms (columns) for directional variability, velocity variability, and step length variability (rows). Directional variability exhibits a clear two-mode structure across window sizes, with low- and high-variability modes separated by the estimated regime boundary (vertical line). In contrast, velocity and step length variability show spike-plus-tail distributions, characterized by a dominant peak near zero and extended right tails without a comparably separable second mode. The persistence of directional variability bimodality across window sizes indicates scale-consistent two-mode structure, whereas magnitude-based metrics primarily reflect asymmetric amplitude distributions.Directional variability regimes reflect directional organization independent of spatial extentWe next tested whether directional variability regimes capture differences in scanpath directional organization, operationalized as directional entropy, angular reversal density, and return rate, beyond differences in spatial extent. If directional variability merely reflects spatial spread, controlling for dispersion should eliminate regime differences. Conversely, if it captures directional dynamics independent of spatial scale, entropy differences should persist and may even increase if spatial spread acts as a confounder. To ensure direct comparability with magnitude based baselines, all regimes were defined using median splits. Windows were classified into low and high variability groups separately for directional variability, velocity, acceleration, and step length. Using a common thresholding rule avoids methodological asymmetries and allows effect sizes to be compared directly across metrics. Across both datasets, median splits on directional variability produced the largest effects for directional entropy and return metrics (Fig. 4). These effect sizes quantify differences in directional entropy between regimes, not the bimodal separation in directional variability itself, which was very large as shown in the preceding section. By conventional thresholds (\(d > 0.8\) large, 0.5 to 0.8 moderate, 0.2 to 0.5 small), directional variability splits yielded moderate entropy effects in REFLACX (\(d = 0.78\)) and moderate effects in GazeBase at 20,ms (Reading: \(d = 0.63\); Gaming: \(d = 0.68\); Video: \(d = 0.65\)). In contrast, velocity, acceleration, and step splits produced effects below the small effect threshold in REFLACX (all \(|d| < 0.20\)) and moderate effects in GazeBase (\(d \approx 0.50\)). For return rate in REFLACX, directional variability yielded \(d = 0.45\), exceeding velocity (\(d = 0.19\)) and acceleration (\(d = 0.05\)). Angular reversal density showed similar patterns (REFLACX: \(d = 0.52\) for directional vs. \(d = 0.18\) for velocity; GazeBase Reading: \(d = 0.58\) vs. \(d = 0.31\)). Importantly, magnitude based splits often produced dispersion effects comparable to or larger than their entropy effects (Fig. 4), indicating that these splits primarily partition windows according to spatial spread rather than directional organization. Because directional variability and spatial dispersion are partially correlated, we next tested whether entropy differences persist when dispersion is approximately matched. Windows were stratified into spatial dispersion quartiles, and regime comparisons were repeated within each quartile. Within quartiles, dispersion differences between low and high variability regimes were small (Cohen’s d for dispersion within quartiles less than 0.15), confirming effective matching. Under matched dispersion conditions, directional entropy effects remained moderate to large across quartiles, whereas magnitude based effects were reduced to small levels. In REFLACX, directional entropy effect sizes ranged from \(d = 0.68\) to 1.19 across Q1 to Q4 (mean \(d = 0.96\)), while velocity based effects ranged from \(d = 0.18\) to 0.29 (mean \(d = 0.24\)). In GazeBase Reading, directional effects ranged from \(d = 0.65\) to 0.82 across quartiles (mean \(d = 0.74\)), whereas velocity effects ranged from \(d = 0.20\) to 0.25 (mean \(d = 0.22\)). Relative to unmatched analyses, magnitude based entropy effects were reduced by more than 50% on average (for example, GazeBase Reading velocity decreased from \(d = 0.50\) unmatched to \(d = 0.21\) in Q1, a 58% reduction), whereas directional effects remained stable or increased. In REFLACX, the directional entropy effect increased from \(d = 0.78\) unmatched to \(d = 1.19\) in Q1, representing a 52% increase, indicating that spatial variability reduces sensitivity to directional organization differences in the full sample. Together, these results support the hypothesis that directional variability regimes correspond to systematic differences in scanpath directional organization that are not reducible to spatial extent or movement magnitude. Entropy and reversal differences persist and in REFLACX strengthen by 52% when dispersion is controlled, demonstrating that directional variability captures directional organization independent of spatial extent and distinct from movement magnitude.Fig. 4Full size imageDirectional variability regimes capture directional organization beyond spatial extent. Absolute effect sizes (Cohen’s d) for differences in directional entropy, angular reversal density, return rate, and spatial dispersion between low- and high-variability regimes in REFLACX (left) and GazeBase at 20 ms (right; Reading, Gaming, Video shown separately). All regimes were defined using median splits for directional variability (Dir), velocity (Vel), acceleration (Acc), and step length (Step), ensuring direct comparability across metrics. Directional variability produces the largest effects for entropy and return metrics in both datasets, whereas magnitude-based splits primarily differentiate spatial dispersion.Directional regime composition predicts scanpath topology and reporting time beyond magnitude-based and mean turn angle baselinesWe next tested whether directional regime composition predicts scanpath topology metrics and reporting time beyond magnitude-based gaze features and mean turn angle. This analysis was conducted in REFLACX, which uniquely includes task completion time defined as time to report. Analyses were performed at the study level (\(n = 2572\) radiograph examinations). For each study, we computed the proportion of temporal windows assigned to each regime during the pre-reporting interval, defined as the time from the first recorded gaze sample to the onset of verbal report. Scanpath topology metrics included transition entropy (variability of grid cell transitions), unique spatial cells visited, stationary gaze entropy (spatial concentration), self-loop fraction (proportion of consecutive transitions within the same cell), and determinism (recurrence structure of transitions). Windows were cross-classified into four regimes using median splits on directional variability and velocity, yielding: (1) high directional variability and high velocity, (2) high directional variability and low velocity, (3) low directional variability and high velocity, and (4) low directional variability and low velocity. Because regime proportions sum to one, including all four would induce perfect multicollinearity. We therefore entered three regime proportions in regression models and omitted the low-directional-variability, high-velocity regime as the reference category. To assess robustness across temporal scales, we tested three window configurations: 10 ms with 5 ms stride, 20 ms with 10 ms stride, and 30 ms with 15 ms stride. Among the four regimes, the high-directional-variability, low-velocity (HiF–LoV) regime corresponds most closely to sustained local inspection characterized by frequent directional corrections within constrained spatial regions. We therefore highlight this regime in correlation analyses to contrast its associations with those of mean velocity. For each outcome, we compared nested linear models with progressively comprehensive baselines: M1 included mean velocity; M2 added velocity and acceleration variability; M3 added mean turn angle; and M4 added regime proportions. The critical test was whether M4 improved prediction over M3, isolating regime composition effects beyond magnitude-based metrics and average directional change (Fig. 5A–B). For 20 ms windows, M3 explained modest variance for transition entropy (\(R^2 = 0.024\)), which increased with regime composition (\(R^2 = 0.116\), \(\Delta R^2 = 0.093\), \(p < 0.001\), \(f^2 = 0.105\)). Similar improvements were observed for unique cells visited (\(R^2_{\textrm{M3}} = 0.059\), \(R^2_{\textrm{M4}} = 0.128\), \(\Delta R^2 = 0.069\), \(p < 0.001\), \(f^2 = 0.079\)), self loops (\(\Delta R^2 = 0.064\), \(p < 0.001\), \(f^2 = 0.103\)), stationary gaze entropy (\(\Delta R^2 = 0.076\), \(p < 0.001\), \(f^2 = 0.085\)), and determinism (\(\Delta R^2 = 0.023\), \(p < 0.001\), \(f^2 = 0.028\)). Regime composition also improved prediction of reporting time (\(R^2_{\textrm{M3}} = 0.113\), \(R^2_{\textrm{M4}} = 0.138\), \(\Delta R^2 = 0.025\), \(p < 0.001\), \(f^2 = 0.028\)). All increments remained significant after false discovery rate correction. Effect sizes fell within the small range (\(f^2 = 0.03\) to 0.10), indicating modest but reliable gains in explanatory power. Predictive gains remained stable across temporal scales (Fig. 5C). For transition entropy, \(\Delta R^2\) ranged from 0.079 at 10 ms to 0.099 at 30 ms. Self loops showed consistent increments ranging from 0.051 to 0.064, while determinism exhibited smaller but stable gains ranging from 0.021 to 0.025. Spearman correlations (Fig. 5D) further illustrate that regime composition captures structure largely independent of overall speed. For 20 ms windows, mean velocity showed negative associations with self loops (\(\rho = -0.49\), \(p < 0.001\)) and determinism (\(\rho = -0.33\), \(p < 0.001\)), and a modest positive correlation with transition entropy (\(\rho = 0.09\), \(p < 0.001\)). In contrast, the HiF–LoV regime proportion showed positive associations with self loops (\(\rho = 0.18\), \(p < 0.001\)) and determinism (\(\rho = 0.24\), \(p < 0.001\)), while its association with transition entropy was weak and not statistically significant (\(\rho = 0.02\), \(p = 0.43\)). These divergent patterns indicate that directional regime composition reflects structured examination dynamics not captured by aggregate movement magnitude alone. Together, these findings demonstrate that the temporal distribution of directional regimes provides incremental explanatory value for scanpath topology and reporting time beyond magnitude-based gaze metrics and mean turn angle.Fig. 5Full size imageDirectional regime composition predicts scanpath topology and reporting time beyond magnitude-based features and mean turn angle. (a) Incremental predictive gain (\(\Delta R^2\)) obtained by adding directional regime proportions (M4) to a model including mean velocity, velocity variability, mean acceleration, acceleration variability, and mean turn angle (M3). Points show \(\Delta R^2\) with 95% bootstrap confidence intervals for 30 ms windows. All increments remain significant after false discovery rate correction (\(p < 0.05\)). (b) Cumulative performance across nested models: M1 (mean velocity), M2 (+ magnitude variability), M3 (+ mean turn angle), and M4 (+ regime proportions). Stacked bars show incremental \(R^2\) contributions. Self loops show the highest absolute prediction (\(R^2 = 0.38\)), whereas transition entropy shows the largest relative gain from regime features. (c) Robustness across temporal scales. \(\Delta R^2\) values remain stable across 10 ms, 20 ms, and 30 ms window configurations. (d) External validity correlations. Heatmap shows Spearman correlations between outcomes and mean velocity (left column) versus the high-directional-variability, low-velocity (HiF–LoV) regime proportion (right column). The HiF–LoV regime reflects sustained low-speed inspection characterized by frequent directional adjustments. Divergent correlation patterns indicate that regime composition captures structured examination dynamics not reducible to overall movement speed.DiscussionWe demonstrate that directional variability in gaze trajectories exhibits a stable bimodal structure across medical image interpretation and general viewing tasks. This two regime organization persists across temporal windows from 10 ms to 50 ms and remains evident after matching spatial dispersion, indicating that it does not arise from window selection or spatial extent alone. Regime composition predicts scanpath topology metrics and reporting time beyond comprehensive magnitude based measures including velocity, acceleration, and turn statistics. Together, these findings indicate that directional variability represents a behaviorally informative dimension of gaze dynamics that is not captured by how fast or how far the eyes move.Oculomotor interpretationHigh directional variability windows, characterized by frequent angular changes within constrained spatial regions, are compatible with sustained periods of local inspection during which small corrective displacements maintain foveal positioning and perceptual stability13. Individual fixational movements occur at various timescales, and the 10 to 50 ms windows examined here are shorter than typical fixation durations of 200 to 300 ms. Rather than isolating individual microsaccades, these windows capture the net directionality of multiple small displacements within or between fixations, reflecting whether gaze remains locally concentrated or transitions spatially. Low directional variability windows reflect directional persistence across successive displacements, consistent with saccadic transitions between spatially separated regions. Unlike global scanpath similarity or curvature metrics computed over entire viewing sequences18,19, windowed directional variability operates at sub fixation to fixation timescales. This enables temporal segmentation of examination modes, identifying when viewers transition between inspection and spatial relocation, rather than characterizing overall path complexity. The bimodal distribution suggests two predominant modes of directional control but does not require discrete physiological states. Viewing behavior likely varies continuously along a directional variability spectrum, with bimodality emerging from behavioral context. Sustained inspection naturally produces frequent directional adjustments, whereas spatial transitions produce directional persistence.Relation to examination strategiesDirectional regimes provide a quantitative framework for describing examination strategies previously characterized qualitatively. In radiology, driller and scanner behaviors have been reported in volumetric imaging4. The high directional low velocity regime observed here is consistent with intensive local inspection involving frequent directional adjustments, whereas low directional regimes correspond to broader exploratory transitions. However, the present study did not directly classify observers into driller or scanner categories or test whether regime proportions differ between such groups. Establishing this mapping requires datasets that include independent behavioral classification alongside eye tracking. Similar distinctions appear in reading and visual foraging research. Increased transition entropy has been associated with flexible information sampling7, and exploration versus exploitation dynamics characterize foraging tasks16. Directional variability extends these frameworks by providing a continuous, data driven statistic for identifying examination modes without predefined task phases, region of interest boundaries, or expert labeling. This enables within trial temporal segmentation of viewing behavior, allowing identification of mode transitions during ongoing examination.Robustness of the regime structureMultiple analyses address alternative explanations. First, bimodality was stable across window sizes, indicating that the regime structure is not an artifact of temporal aggregation. Second, entropy and return differences persisted under dispersion matched stratification, demonstrating that directional effects are not reducible to spatial spread. Third, magnitude based metrics exhibited asymmetric spike plus tail distributions, as shown in the Results, rather than comparable mode separation, indicating that bimodality is not a generic property of windowed variability. Together, these results indicate that directional variability captures a distinct dimension of gaze dynamics, namely how frequently viewing direction changes over short timescales, that is separable from movement magnitude such as velocity, spatial displacement such as step length, or kinematic acceleration.Implications for radiology and human–AI interactionDirectional regime composition may serve as an additional descriptor of examination behavior in radiology education and quality assessment. Windows characterized by high directional variability reflect sustained local inspection, whereas low variability windows reflect broader sampling. These metrics could complement dwell time and spatial coverage measures. However, the present study does not establish whether directional regimes predict diagnostic accuracy. While regime proportions predict scanpath topology and task duration, the relationship to clinically relevant outcomes including detection rates, false positives, and diagnostic confidence remains unknown. Prospective studies linking regime composition to performance metrics are required before clinical utility can be claimed. Until such validation is conducted, directional metrics should augment, not replace, accuracy based evaluation. Directional information could also inform human AI collaboration by distinguishing sustained inspection from transient scanning. Current attention models typically weight fixations by duration or frequency, and incorporating directional statistics may provide interpretable behavioral features. This application remains hypothetical and requires empirical testing to determine whether directional features improve model performance or introduce unintended biases.Limitations and future directionsSeveral limitations constrain interpretation. Regime boundaries were estimated at the dataset level. Without subject level validation of boundary stability or classification reliability, the present findings do not support individual level inference, such as labeling individual radiologists by regime dominance. Establishing individual reliability requires within subject replication and test retest analysis. Most critically, the radiology dataset comprised expert observers only. We cannot determine whether directional regimes differentiate novice and expert behavior, change with training, or predict skill acquisition. Mixed expertise cohorts are necessary to evaluate training related applications. Furthermore, directional regime composition was not validated against diagnostic accuracy. Although regime proportions predict scanpath complexity metrics and reporting time, their relationship to detection performance remains unknown. Evaluating whether regime composition predicts diagnostic accuracy is the highest priority for future work. Additional priorities include assessing individual level reliability, extending analyses to mixed expertise cohorts and additional imaging modalities, and exploring alternative continuous or graph based topology representations. Topology metrics relied on spatial discretization using a 10 by 10 grid, and alternative continuous or graph based representations may provide complementary insights. The present study focuses on overt gaze dynamics, and multimodal influences on attention14 were not examined.MethodsDatasetsGazeBase. GazeBase20 is a large-scale, longitudinal eye-tracking dataset comprising 12,334 monocular (left-eye) recordings from 322 college-aged participants, collected over three years across nine recording rounds using an EyeLink 1000 eye tracker at 1,000 Hz sampling rate. Each participant completed seven tasks in two sessions per round (with a 20-minute interval): random saccades (RAN), reading (TEX), fixation (FXS), horizontal saccades (HSS), two video-viewing tasks (VD1 and VD2), and a gaze-driven gaming task (BALURA/BLG). We analyzed the BALURA, READING (TEX), and VIDEO 1 (VD1) tasks to evaluate windowed directional variability under diverse non-clinical conditions, such as gaming, text processing, and dynamic scene viewing. Pixel coordinates were derived from the dataset’s published EyeLink 1000 calibration geometry.REFLACX21 is a specialized dataset of radiologist gaze and speech during chest X-ray interpretation, containing 3,032 synchronized recordings from five board-certified thoracic radiologists who interpreted 2,616 frontal chest X-rays sampled from the MIMIC-CXR database. Eye-tracking data were captured using an EyeLink 1000 Plus system at 1,000 Hz in remote mode, paired with timestamped dictation transcriptions, zooming/panning/windowing states, and auxiliary manual annotations for validation. These include image-level abnormality labels (e.g., atelectasis, consolidation, pleural effusion), ellipses localizing abnormalities with certainty scores (from “Unlikely” 90%), and bounding boxes around the lungs and heart.Gaze preprocessingRaw gaze coordinates were converted to pixel space using each dataset’s calibration geometry. Samples flagged as invalid by the eye tracker were removed. After removal of invalid samples, the effective sampling density was reduced relative to nominal 1000 Hz due to blinks and tracking loss; window sizes from 10 to 50 ms were analyzed, corresponding approximately to 5 to 25 valid gaze samples per window. Gaze trajectories were segmented into overlapping temporal windows with 50 percent overlap. Unless otherwise stated, 20 ms windows with 10 ms stride were used for regime composition and predictive analyses.Directional variability computationGaze position at time index i is denoted by \(\textbf{g}_i \in \mathbb {R}^2\), where samples were recorded at 1000 Hz (i.e., consecutive indices correspond to 1 ms intervals). Consecutive displacement vectors were defined at the native sampling resolution as$$\vec {v}_i = \textbf{g}_i - \textbf{g}_{i-1}.$$The unsigned turn angle between successive displacement vectors \(\vec {v}_i\) and \(\vec {v}_{i+1}\) was computed using the dot product identity,$$\begin{aligned} \theta _i = \arccos \left( \frac{\vec {v}_i \cdot \vec {v}_{i+1}}{\Vert \vec {v}_i \Vert \, \Vert \vec {v}_{i+1} \Vert } \right) , \end{aligned}$$(1)which yields the angle between consecutive movement directions. If either displacement vector had zero magnitude, \(\theta _i\) was undefined and excluded from computation. In practice, this occurred rarely and primarily during periods of minimal movement. Turn angles \(\theta _i\) were computed at the full sampling resolution prior to temporal aggregation. Directional variability within a temporal window w was then defined as the sample standard deviation of all turn angles whose indices fell within that window,$$\begin{aligned} \tau _w = \sqrt{ \frac{1}{n_w - 1} \sum _{i \in w} \left( \theta _i - \bar{\theta }_w \right) ^2 }, \end{aligned}$$(2)where \(n_w\) is the number of valid turn angles in window w, and \(\bar{\theta }_w\) is their mean. Because \(\theta _i\) depends only on the orientation of displacement vectors and not their magnitude, directional variability isolates changes in movement direction independently of displacement amplitude. Velocity, acceleration, and step length variability were computed analogously as the sample standard deviation within window w of instantaneous speed \(\Vert \vec {v}_i \Vert\), acceleration (first difference of speed), and step length \(\Vert \vec {v}_i \Vert\), respectively.Distributional analysisTo assess whether windowed variability metrics exhibit bimodal structure, we analyzed the empirical distributions of directional variability, velocity variability, acceleration variability, and step length variability across temporal windows of 10 to 50 ms in each dataset and task. For each condition, window-level variability values were pooled and histogram-based density estimates were computed using 150 bins spanning the observed range. Distributional asymmetry was quantified using the tail-to-median ratio \(p_{99}/p_{50}\), where \(p_{50}\) and \(p_{99}\) denote the 50th and 99th percentiles of the distribution. To evaluate bimodality formally, we fitted one- and two-component Gaussian mixture models using expectation–maximization. For directional variability, the two-component model was initialized using a data-driven valley estimate defined as the minimum of a Savitzky–Golay smoothed density curve (window length 51 bins, polynomial order 3) within the interval \(5^\circ\) to \(40^\circ\), a range that consistently contained the separation between modes across datasets and window sizes; sensitivity analysis varying this interval by \(\pm 5^\circ\) altered the valley location by less than \(1.5^\circ\). For magnitude-based metrics, two-component models were initialized randomly with 10 restarts to avoid local optima. For computational efficiency with very large window counts, mixture parameters were estimated on random subsamples of up to \(5 \times 10^5\) windows per condition, while descriptive statistics were computed on the full distributions. Model comparison between one- and two-component fits used the Bayesian Information Criterion, reporting \(\Delta \textrm{BIC} = \textrm{BIC}_1 - \textrm{BIC}_2\), where positive values favor the two-component model. Separation strength was quantified using Cohen’s \(d = |\mu _1 - \mu _2| / \sqrt{(\sigma _1^2 + \sigma _2^2)/2}\), where \(\mu _k\) and \(\sigma _k^2\) are the mean and variance of component k, and we additionally report the minimum mixture proportion to ensure that separation is not driven by a negligible secondary component. To determine whether apparent bimodality reflects pooling across heterogeneous observers rather than within-subject dynamics, we computed a subject variance ratio in GazeBase defined as \(\textrm{SVR} = \textrm{Var}(\mu _{\text {subject}})/\textrm{Var}(x_{\text {pooled}})\), where \(\mu _{\text {subject}}\) denotes each subject’s mean directional variability and \(x_{\text {pooled}}\) denotes all window-level values; low SVR values indicate that between-subject differences explain only a small fraction of total variance. Temporal windows were constructed with 50% overlap, and mixture modeling was used descriptively to characterize distributional structure rather than to conduct inferential tests assuming independent observations.Regime interpretation analysisTo test whether directional variability regimes reflect directional dynamics independent of spatial extent, we computed window level directional entropy, angular reversal density, return rate, and spatial dispersion. Directional entropy was computed by discretizing turn angles within each window into 36 equal bins spanning \(0^\circ\) to \(180^\circ\) and calculating Shannon entropy:$$\begin{aligned} H_{\text {dir}} = -\sum _{k=1}^{36} p_k \log _2(p_k), \end{aligned}$$(3)where \(p_k\) denotes the proportion of turn angles falling in bin k within window w. Higher values indicate greater variability in directional change. Angular reversal density was defined as the proportion of consecutive turn angles within a window exceeding \(90^\circ\), quantifying the frequency of directional reversals. Return rate was computed from window center transitions after discretizing spatial position to a fixed grid. For each window, we identified transitions that returned to a previously visited grid cell and defined return rate as the proportion of such revisits relative to total transitions within the window. Spatial dispersion was quantified as the root mean squared distance of gaze samples from the window centroid, providing a measure of spatial spread independent of directional structure. For regime comparisons, all metrics were split using median thresholds to ensure direct comparability across directional variability, velocity, acceleration, and step length variability. Windows were classified into low and high variability groups separately for each metric. Group differences were assessed using Mann Whitney U tests, and effect sizes were reported as rank biserial correlations. To test independence from spatial extent, spatial dispersion quartiles were computed from the pooled distribution across all windows. Regime comparisons for directional entropy, angular reversal density, and return rate were then repeated within each quartile to approximately match dispersion between groups and isolate directional effects from spatial spread.Regime taxonomy and topology prediction analysisPredictive analyses were conducted at the study level in REFLACX. After exclusion of examinations with insufficient pre reporting gaze data or implausible reporting times, the final sample comprised \(n = 2572\) radiograph examinations. For each study, temporal windows were extracted from the pre reporting interval defined as the time between the first recorded gaze sample and the onset of verbal dictation. Time to report was defined as the duration of this interval. Directional variability and velocity were computed within each window. For predictive analyses, windows were cross classified using median splits applied separately to directional variability and velocity within each window configuration. This yielded four regimes: high directional high velocity, high directional low velocity, low directional high velocity, and low directional low velocity. For each examination, regime proportions were defined as the fraction of windows assigned to each regime. Because regime proportions sum to one, regression models included three regime proportions with the low directional high velocity regime treated as the reference category to avoid perfect multicollinearity. Scanpath topology metrics were computed from sequences of window centers discretized to a \(10 \times 10\) spatial grid normalized to each examination’s spatial extent. Transition entropy was defined as$$H_{\text {trans}} = -\sum _{(i,j)} p_{ij} \log _2(p_{ij}),$$where \(p_{ij}\) denotes the empirical probability of transitioning from grid cell i to cell j. Unique cells counted the number of distinct grid cells visited. Self loop fraction was defined as the proportion of consecutive windows remaining within the same grid cell. Stationary gaze entropy was computed as$$H_{\text {SGE}} = -\sum _{k} p_k \log _2(p_k),$$where \(p_k\) denotes the proportion of time spent in grid cell k. Determinism quantified recurrence structure as the fraction of recurrent points forming diagonal lines of length at least two in the recurrence matrix. Predictive models were evaluated across three window configurations: 10 ms with 5 ms stride, 20 ms with 10 ms stride, and 30 ms with 15 ms stride. For each outcome, nested ordinary least squares models were fitted with progressively comprehensive magnitude baselines. Model M1 included mean velocity. Model M2 added velocity variability and acceleration variability. Model M3 added mean turn angle. Model M4 added three directional regime proportions. The critical comparison tested whether M4 improved prediction over M3, isolating the incremental contribution of regime composition beyond aggregate magnitude based metrics. Incremental variance explained was quantified as \(\Delta R^2 = R^2_{\text {M4}} - R^2_{\text {M3}}\). Statistical significance was assessed using nested F tests from analysis of variance. Ninety five percent confidence intervals for \(\Delta R^2\) were estimated using nonparametric bootstrap resampling with 1, 000 iterations. False discovery rate correction was applied within each outcome across window sizes using the Benjamini Hochberg procedure. Incremental effect sizes were summarized using Cohen’s \(f^2 = \Delta R^2 / (1 - R^2_{\text {M4}})\). Spearman rank correlations were computed to characterize bivariate associations between regime proportions, topology metrics, and reporting time.