Alt

SEDarc PhD Cohort 2025

Haibei Wang

Haibeiwang

What can large language models tell us about skilled reading?

The ability to read is central to an individual’s economic chances, their health and well-being, and their ability to self-advocate. Many theories of skilled reading argue that we use prediction to increase the speed and efficiency of information processing. For example, models of eye movement control during reading (e.g., the E-Z Reader model) assume that the time taken to identify words in context is governed by a word’s predictability. However, these theories have a serious limitation in that they have not articulated precisely what is meant by prediction and have instead used a metric developed over 70 years ago (cloze probability) as a crude proxy for predictability. This project takes advantage of immense strides in the development of large language models for AI and engineering purposes to advance understanding of how readers use predictive information. I will derive predictability metrics from recently-developed large language models for three large eye-movement corpora. I will then test (1) whether these account for readers’ eye-movements better than existing cloze probability approaches; and (2) whether implementation of these metrics in the E-Z Reader model increases its explanatory power. I will then (3) conduct new empirical studies investigating how top-down predictions interact with bottom-up visual input in reading. This work lies at the intersection of behavioural social sciences, computer science, and reading research, and has significant theoretical and practical implications for both reading research and the emerging field of large language models. It intersects the Transformative Technologies for Society, Big Data and Interdisciplinary strategic steers.