Seeing Through Experts Eyes A Foundational Vision Language Model Trained on Radiologists Gaze and Reasoning

arXiv cs.AI / 4/17/2026

📰 NewsIdeas & Deep AnalysisModels & Research

共有:

Key Points

The paper argues that vision-language models for chest X-rays often fall short because they optimize for semantic correctness without mirroring how radiologists visually inspect and reason over images.
It introduces GazeX, which uses radiologists’ eye-tracking data as a behavioral prior by incorporating gaze trajectories and fixation patterns into pretraining.
GazeX is trained on a curated dataset with gaze key frames from five radiologists and evaluated using large-scale radiology study, QA, and caption/bounding-box datasets.
Results claim that GazeX improves accuracy, interpretability, and consistency with expert diagnostic workflows across report generation, disease grounding, and visual question answering.
Unlike fully autonomous systems, GazeX is designed to output verifiable evidence artifacts such as inspection trajectories and localized findings to support safer human–AI collaboration.

Abstract

Large scale vision language models have shown promise in automating chest Xray interpretation, yet their clinical utility remains limited by a gap between model outputs and radiologist reasoning. Most systems optimize for semantic information without emulating how experts visually examine medical images, often overlooking critical findings or diverging from established diagnostic workflows. Radiologists follow structured protocols (e.g., the ABCDEF approach) that ensure all clinically relevant regions are systematically examined, reducing missed findings and supporting reliable diagnostic reasoning. We introduce GazeX, a vision language model that leverages radiologists' eye tracking data as a behavioral prior to model expert diagnostic reasoning. By incorporating gaze trajectories and fixation patterns into pretraining, GazeX learns to follow the spatial and temporal structure of radiologist attention and integrates observations in a clinically meaningful sequence. Using a curated dataset of over 30,000 gaze key frames from five radiologists, we demonstrate that GazeX produces more accurate, interpretable, and expert consistent outputs across radiology report generation, disease grounding, and visual question answering, utilizing 231,835 radiographic studies, 780,014 question answer pairs, and 1,162 image sentence pairs with bounding boxes. Unlike autonomous reporting systems, GazeX produces verifiable evidence artifacts, including inspection trajectories and finding linked localized regions, enabling efficient human verification and safe human AI collaboration. Learning through expert eyes provides a practical route toward more trustworthy, explainable, and diagnostically robust AI systems for radiology and beyond.