Abstract

Listening effort, which MeSH classifies under auditory perception, is the deliberate allocation of mental resources to the task of understanding speech, especially when the signal is degraded by noise, hearing loss, or an unfamiliar accent. It is distinct from intelligibility: a listener may understand every word and still be working hard to do so, and that hidden cost shows up as fatigue, slowed responses on a concurrent task, and a dilating pupil. The dominant frameworks — Kahneman's capacity model, the Framework for Understanding Effortful Listening (FUEL), and the Ease of Language Understanding (ELU) model — treat effort as a limited, motivation-mediated commitment of processing capacity drawn on when an impoverished signal fails to match stored linguistic representations. Measuring it has become central to audiology, because two listeners with identical word scores can bear very different loads.

Keywords: listening effort, cognitive load, pupillometry, speech perception in noise, effortful listening

What Listening Effort Is

Listening effort is the mental work a listener invests to recognise and understand speech. When the acoustic signal is clear and the listener's hearing is normal, comprehension feels effortless and automatic. When the signal is degraded — masked by background noise, distorted by hearing loss, thinned by a poor telephone line, or carried in an unfamiliar accent — understanding still may be achieved, but only by recruiting additional cognitive resources: attention, working memory, and executive control. That extra recruitment is listening effort.

The defining move in the field is to separate effort from intelligibility, the proportion of speech correctly identified. The two can diverge sharply. A listener in moderate noise may score 100% on a word-recognition test while drawing heavily on cognitive reserves to do so, and the cost of that investment — measurable as fatigue at the end of a working day, as poorer performance on anything else attempted at the same time, or as a physiological stress response — is invisible to the word score. This is why listening effort matters clinically: a hearing aid that restores audibility, and so restores the word score, does not necessarily restore ease, and two patients with the same audiogram may live with very different daily burdens.

The construct draws its theoretical backbone from the general psychology of attention and effort. Kahneman's capacity model held that mental effort is a graded, limited commitment of processing resources, allocated moment to moment according to the demands of the task and the motivation of the person, and read out physiologically by the dilation of the pupil (Kahneman, 1973). Listening-effort research imports that model wholesale: understanding degraded speech is a task that draws on the same limited pool, and the size of the draw is what the field sets out to measure.

Theoretical Frameworks

Three frameworks organise the modern study of listening effort. The first is Kahneman's capacity model itself, which supplies the core idea that comprehension competes for a finite pool of processing resources and that the resources committed can be indexed by physiological arousal. On this account a degraded signal raises the processing demand, and if the listener is motivated to understand, more capacity is allocated — up to the limit of what is available (Kahneman, 1973).

The second is the Framework for Understanding Effortful Listening (FUEL), which extends the capacity model by making motivation explicit. FUEL defines listening effort as the deliberate allocation of mental resources to overcome obstacles in a listening task, and insists that the amount allocated is not fixed by the difficulty of the signal alone but is mediated by the listener's motivation, the perceived importance of succeeding, and the demands competing for the same resources (Pichora-Fuller et al., 2016). It borrows the language of motivational-intensity theory: a listener disengages when the task is judged impossible or unimportant, so maximum effort is found not at maximum difficulty but at the point where success is hard yet still worth the cost.

The third is the Ease of Language Understanding (ELU) model, which specifies the cognitive mechanism. ELU holds that speech is normally understood by a rapid, implicit match between the input and stored phonological representations in long-term memory. When the signal is degraded, that match fails, and comprehension must fall back on slow, explicit, working-memory-dependent processing to reconstruct the message from context and partial cues (Ronnberg et al., 2013). This is why individual differences in working memory capacity predict who copes well with degraded speech: the explicit processing that a mismatch forces draws directly on working memory, and listeners with more of it carry the load more easily.

Demonstration 1

Capacity allocation and the dual-task cost

capacity pool (100%)speech 60%RT 650 ms
speech taskleft for secondary taskunmet demand (disengaged)
The speech task draws 60% of capacity, leaving 40% for the concurrent task, whose reaction time rises from 500 ms to 650 ms — a dual-task cost of 30%. All of the demand is met. At a demand of 20% the cost is 10% (RT 550 ms); at 60% it is 30% (RT 650 ms), the worked example in the text.
A fixed pool of mental capacity is split between understanding speech and a concurrent task. Raise the listening difficulty and the speech task claims more of the pool, slowing the secondary task — the dual-task cost. Lower the motivation ceiling and the listener disengages: demand above the ceiling goes unmet rather than drawing ever more effort.

Measuring Listening Effort

Because listening effort is a hidden cost rather than an overt error, its measurement has driven much of the field. Three families of measure are used, and they do not always agree (McGarrigle et al., 2014).

Behavioural measures infer effort from performance on a concurrent task. In the dual-task paradigm the listener recognises speech while simultaneously doing something else — pressing a key as fast as possible to a visual probe, say, or holding a list of digits in memory. As the speech task consumes more capacity, less is left for the secondary task, so a slower reaction time or poorer recall on the second task indexes the effort the first is costing, even when speech recognition itself stays near perfect. Subjective measures ask the listener directly, with rating scales and questionnaires of perceived effort and fatigue. Physiological measures read the body's response, and of these the task-evoked pupil response has become the signature of the field.

Table 1

Three Families of Listening-Effort Measure

FamilyExample measureWhat it indexes
BehaviouralDual-task reaction time; recall of a concurrent memory loadResidual capacity left over after the speech task is served
SubjectiveSelf-report rating scales; fatigue questionnairesThe listener's own experience of how hard they are working
PhysiologicalTask-evoked pupil dilation; skin conductance; cortical activityAutonomic and neural arousal accompanying resource allocation

Note. The three families often diverge — a listener may report low effort while the pupil shows high arousal — which is itself an active research problem. Full citations appear in the References (McGarrigle et al., 2014; Pichora-Fuller et al., 2016).

The task-evoked pupil response rests on the oldest strand of the theory: Kahneman showed that the pupil dilates in proportion to mental effort, and listening-effort research turned that instrument on degraded speech. Zekveld, Kramer, and Festen established that the peak dilation of the pupil during a sentence grows as the sentence becomes harder to understand, tracking the cognitive load of listening rather than the acoustic level of the speech (Zekveld et al., 2010). The pupil has since become the workhorse measure of the field, sensitive enough to reveal effort differences between conditions that produce identical word scores.

Figure 1

The Task-Evoked Pupil Response to a Single Sentence

Pupil diameter rising to a peak during a sentence and recovering afterward Pupil diameter plotted against time. Before the sentence the pupil holds a flat baseline. After sentence onset it dilates, rising to a peak a second or two after the sentence ends, then recovering slowly toward the baseline. A harder sentence produces a larger peak than an easier one. Pupil diameter Time baseline sentence onset peak (hard) peak (easy)
The pupil holds a baseline before the sentence, dilates to a peak a second or two after it ends, and then recovers slowly. The peak is larger for a harder sentence (solid) than an easier one (dashed), even when both are understood correctly, which is what makes peak dilation an index of effort rather than of intelligibility.

Demonstration 2

The task-evoked pupil response

sentence3.63.84.04.2Pupil diameter (mm)Time (s)0246pupil
At -3 dB SNR the peak task-evoked dilation is about 0.37 mm (a peak diameter of 3.97 mm on a 3.6 mm baseline). At intermediate difficulty the pupil shows a clear, graded response. The physical loudness of the speech is unchanged; the pupil is tracking effort, not sound level.
As the signal-to-noise ratio falls, the sentence becomes harder to understand and the pupil dilates more. The peak of the task-evoked pupil response tracks the cognitive load of listening, not the loudness of the speech, which is what makes it the signature physiological measure of listening effort.

Effort and Intelligibility

The most consequential finding in the field is that effort and intelligibility come apart. Intelligibility rises monotonically as the signal improves: the better the signal-to-noise ratio, the more words are understood, up to a ceiling of perfect recognition. Effort does not follow the same curve. At very poor signal-to-noise ratios, where almost nothing is intelligible, effort is paradoxically low — the listener disengages, because the task is hopeless. At very good ratios, where everything is easy, effort is also low, because little is required. Effort peaks in between, at the intermediate difficulty where comprehension is achievable but only by hard work.

Winn and colleagues demonstrated the dissociation directly, showing that pupil-indexed effort continues to rise as spectral resolution degrades even across a range where word-recognition accuracy stays at ceiling: the listeners were understanding everything, and working steadily harder to do it (Winn et al., 2015). The practical implication is sharp. An intervention evaluated only by its effect on the word score — the standard clinical outcome — can look like it has done nothing while in fact substantially lowering the effort a listener must pour in. Ohlenforst and colleagues, reviewing how hearing impairment and amplification affect listening effort, found that this is exactly where hearing aids may deliver a benefit the audiogram cannot capture (Ohlenforst et al., 2017).

Demonstration 3

Effort versus intelligibility

0255075100Percent of maximumSignal-to-noise ratio (dB)-16-80+8+16
intelligibilitylistening effort
At -2 dB, intelligibility is 50% and effort is 100% of its maximum. Comprehension is achievable but costly — this is where effort peaks, even as intelligibility is still climbing.
Intelligibility rises monotonically as the signal improves; effort does not. Effort peaks at intermediate difficulty — where comprehension is achievable but only by hard work — and falls at both ends, because a hopeless signal makes a motivated listener disengage and an easy one needs no work. Drag the marker to read both at once.

Neural and Affective Dimensions

Listening effort is not only a matter of allocated capacity; it has a neural signature and an affective cost. Neuroimaging shows that comprehending acoustically degraded speech recruits regions beyond the auditory cortex — frontal and cingulate areas associated with attention, working memory, and cognitive control — so that the effort of listening is visible as a redistribution of neural activity toward domain-general cognitive systems (Wild et al., 2012). Peelle's synthesis of the brain and behavioural evidence argues that this recruitment is the neural reflection of the capacity drain the frameworks describe: when the acoustic signal is poor, cognitive systems are conscripted to shore up perception (Peelle, 2018).

The affective dimension is equally real. Francis and Love argue that what is measured as listening effort may be partly cognitive load and partly the stress and physiological arousal of a demanding, aversive situation, so the pupil and skin-conductance signals index not only how much capacity is committed but how the listener feels about committing it (Francis & Love, 2020). Herrmann and Johnsrude reframe the whole construct around engagement, arguing that sustained listening is governed by the listener's fluctuating motivation and state over time, not by a fixed response to momentary difficulty (Herrmann & Johnsrude, 2020). On this view the question is not simply how hard a signal is, but whether and for how long the listener chooses to keep working at it.

Worked Example

Consider a dual-task measurement of listening effort. A listener performs a visual reaction-time task alone, giving a baseline response time of 500 ms. The same visual task is then performed while the listener also recognises sentences in two conditions. In the easy condition (a favourable signal-to-noise ratio) the visual response time rises to 550 ms; in the hard condition (an unfavourable ratio) it rises to 650 ms.

The dual-task cost is the proportional slowing of the secondary task relative to baseline. For the easy condition it is (550 − 500) / 500 = 50 / 500 = 0.10, a 10% cost. For the hard condition it is (650 − 500) / 500 = 150 / 500 = 0.30, a 30% cost. The hard condition therefore extracts three times the listening effort of the easy one (0.30 / 0.10 = 3), even if — and this is the crucial point — the listener recognised the same proportion of words in both. The word score can be flat across the two conditions while the effort measure reveals a threefold difference in the cognitive price paid, which is precisely the dissociation the dual-task paradigm exists to expose. These worked numbers match the Capacity Allocation demonstration above.

Current Directions

The field's liveliest current question is why its three measurement families — behavioural, subjective, and physiological — so often fail to converge. A listener may rate a condition as low-effort while the pupil signals high arousal, or show a large dual-task cost with no corresponding self-report. Francis and Love's argument that effort measures confound cognitive load with affective arousal is one attempt to explain the divergence, and it has pushed the field toward multidimensional measurement that reports cognitive and affective components separately rather than collapsing them into a single number (Francis & Love, 2020).

A second active direction is the move from momentary effort to sustained engagement. Herrmann and Johnsrude's model of listening engagement treats the listener as an agent whose motivation rises and falls over minutes and hours, so that the relevant clinical outcome is not the effort of a single sentence but whether a person stays engaged through a long conversation or withdraws from social listening altogether (Herrmann & Johnsrude, 2020). This reframing connects listening effort to the real-world consequences — fatigue, social withdrawal, the decision to stop wearing a hearing aid — that motivated the construct in the first place, and it remains an open problem to measure engagement as directly as the pupil now measures momentary load.

Discussion

Listening effort has reorganised how the understanding of speech is evaluated. For most of its history, audiology asked a single question — how much does the listener understand? — and answered it with a word score. The effort construct adds a second, orthogonal question: at what cost? The two are genuinely independent, and the independence is the point. By importing Kahneman's capacity model, FUEL and ELU gave the field a principled account of why a degraded signal should cost more even when it is fully understood, and pupillometry gave it an instrument sensitive enough to measure that cost where behaviour shows nothing.

The lasting contribution is a shift in what counts as a successful outcome. A hearing aid, a cochlear implant, or a noise-reduction algorithm may leave the word score untouched while substantially lowering the effort a listener must spend, and that reduction — felt as less fatigue, more capacity for everything else, a greater willingness to stay in the conversation — is a real benefit the older measures could not see. That listening effort is still contested at its foundations, divided between cognitive and affective readings and between momentary and sustained accounts, is a sign of how much work the construct is now being asked to do.

Common Misconceptions

If a listener understands every word, listening is effortless.
Intelligibility and effort are distinct. A listener can score 100% on a word-recognition test while drawing heavily on cognitive resources to do it, a cost invisible to the word score but measurable in fatigue, dual-task slowing, and pupil dilation (Winn et al., 2015).
Listening effort is greatest when the signal is worst.
Effort peaks at intermediate difficulty, not at maximum difficulty. When a signal is hopeless the motivated listener disengages and effort falls; maximum effort is found where success is hard but still achievable and worth the cost (Pichora-Fuller et al., 2016).
A hearing aid that restores the word score restores ease.
Restoring audibility need not restore ease. Two listeners with the same audiogram and the same word score can bear very different effort, which is why amplification is now evaluated for its effect on effort, not intelligibility alone (Ohlenforst et al., 2017).
The pupil dilates because the sound is loud.
The task-evoked pupil response tracks cognitive load, not acoustic level. Peak dilation grows as a sentence becomes harder to understand, even when its physical intensity is held constant, which is what makes the pupil a measure of effort rather than of sound (Zekveld et al., 2010).

Glossary

Capacity model.
Kahneman's account of attention as a single limited pool of processing resources, allocated according to task demand and motivation and indexed physiologically by pupil dilation.

Cognitive load.
The quantity of mental resources a task consumes; listening effort is the cognitive load specifically imposed by understanding speech.

Dual-task paradigm.
A method that infers effort from performance on a concurrent secondary task: as the speech task consumes capacity, the secondary task slows or suffers, indexing the effort the speech is costing.

Ease of Language Understanding (ELU).
Ronnberg's model holding that speech is normally understood by an implicit match to stored phonological representations, and that a mismatch forces slow, explicit, working-memory-dependent processing.

Effortful listening.
Listening that requires the deliberate allocation of cognitive resources because the signal, the listener's hearing, or the environment makes automatic comprehension impossible.

Engagement.
The listener's sustained motivation to keep working at a listening task over time, proposed as the governing variable in accounts that treat effort as fluctuating rather than fixed.

FUEL.
The Framework for Understanding Effortful Listening, which defines listening effort as the motivation-mediated deliberate allocation of mental resources to overcome obstacles in a listening task.

Intelligibility.
The proportion of speech correctly identified; the traditional outcome measure of speech perception, distinct from and often dissociable from effort.

Listening-related fatigue.
The tiredness and depletion that follow sustained effortful listening, one of the real-world costs the effort construct was developed to capture.

Motivation.
The listener's drive to succeed at a listening task, which in FUEL determines how much of the available capacity is actually committed to understanding a given signal.

Pupillometry.
The measurement of pupil diameter over time; the task-evoked pupil response is the signature physiological index of listening effort.

Signal-to-noise ratio (SNR).
The level of a target speech signal relative to background noise, in decibels; the principal variable manipulated to make listening harder or easier.

Speech perception in noise.
The recognition of speech against a background of competing sound, the canonical degraded-listening situation in which effort is studied.

Task-evoked pupil response.
The transient dilation of the pupil during a cognitive task, whose peak amplitude grows with the mental effort the task demands.

Working memory.
The system that holds and manipulates information over short intervals; in ELU it carries the explicit processing a degraded signal forces, so its capacity predicts who copes with effortful listening.

Key Researchers

Daniel Kahneman

(1934–2024). Israeli-American psychologist and winner of the 2002 Nobel Memorial Prize in Economic Sciences. His 1973 monograph Attention and Effort set out the capacity model in which mental effort is a graded, limited commitment of processing resources read out by pupil dilation — the theoretical foundation on which the modern listening-effort frameworks were built. See his Wikipedia biography.

Sophia E. Kramer

(living). Professor of Auditory Functioning and Participation in the Otolaryngology–Head and Neck Surgery department of Amsterdam UMC (Vrije Universiteit Amsterdam). Her work established pupillometry as a measure of the cognitive load of listening to degraded speech and she co-authored the FUEL framework. ORCID 0000-0002-0451-8179.

Jonathan E. Peelle

(living). Professor in the Department of Communication Sciences and Disorders at Northeastern University. His neuroimaging work showed that comprehending acoustically degraded speech recruits cognitive and frontal resources beyond the auditory cortex, and his review synthesised the brain and behavioural signatures of listening effort. ORCID 0000-0001-9194-854X.

M. Kathleen Pichora-Fuller

(living). Professor Emerita of Psychology at the University of Toronto and lead author of the Framework for Understanding Effortful Listening (FUEL), which imports Kahneman's capacity model and motivational-intensity theory to define listening effort as the deliberate, motivation-mediated allocation of resources. See her faculty page.

Jerker Ronnberg

(living). Senior Professor of Psychology at Linkoping University and architect of the Ease of Language Understanding (ELU) model, which explains effortful listening as the explicit, working-memory-dependent processing triggered when a degraded signal fails to match stored phonological representations. ORCID 0000-0001-7311-9959.

Adriana A. Zekveld

(living). Associate Professor in the Otolaryngology–Head and Neck Surgery department of Amsterdam UMC (Vrije Universiteit Amsterdam). Her pupillometry studies are central to the measurement of listening effort, establishing the task-evoked pupil response as an index of the cognitive load of understanding degraded speech. ORCID 0000-0003-1320-6908.

Frequently Asked Questions

What is listening effort?

Listening effort is the mental work a listener invests to understand speech, especially when the signal is degraded by noise, hearing loss, or an unfamiliar accent. It is the deliberate allocation of attention, working memory, and executive resources to a task that would otherwise be automatic, and it can be substantial even when the listener understands every word.

How is listening effort different from intelligibility?

Intelligibility is how much speech is correctly understood; effort is how hard the listener works to understand it. The two are distinct and often diverge: a listener can score perfectly on a word-recognition test while expending large amounts of cognitive resource, so a flat word score can hide a large difference in effort.

How is listening effort measured?

Three families of measure are used. Behavioural measures infer effort from slowing on a concurrent task; subjective measures ask the listener to rate how hard they are working; and physiological measures read the body's response, most often the dilation of the pupil during a sentence. The three do not always agree, which is itself an active research problem.

Why does the pupil dilate during listening?

The task-evoked pupil response reflects mental effort, not the loudness of the sound. Kahneman showed that the pupil dilates in proportion to the cognitive resources a task demands, and listening-effort research uses this: peak pupil dilation grows as a sentence becomes harder to understand even when its physical intensity is held constant.

When is listening effort greatest?

At intermediate difficulty, not at maximum difficulty. When a signal is hopeless, a motivated listener disengages and effort falls; when it is easy, little effort is needed. Maximum effort is found in between, where comprehension is achievable but only by hard work and the listener judges the effort worthwhile.

What is the FUEL framework?

The Framework for Understanding Effortful Listening defines listening effort as the deliberate allocation of mental resources to overcome obstacles in a listening task. Its key addition is motivation: the amount of effort allocated depends not only on how hard the signal is but on how important success is to the listener and what else is competing for the same resources.

Why does listening effort matter clinically?

Because restoring the word score does not necessarily restore ease. A hearing aid can leave recognition unchanged while substantially lowering the effort a listener must spend, a benefit the audiogram cannot capture. Measuring effort reveals the daily burden — fatigue, social withdrawal — that two patients with identical word scores may not share.

What does working memory have to do with listening effort?

In the Ease of Language Understanding model, a degraded signal fails to match stored representations automatically and forces slow, explicit processing that draws on working memory. Listeners with greater working-memory capacity therefore carry the load of effortful listening more easily, which is why working memory predicts who copes well with speech in noise.

References

Francis, A. L., & Love, J. (2020). Listening effort: Are we measuring cognition or affect, or both? WIREs Cognitive Science, 11(1), e1514. https://doi.org/10.1002/wcs.1514

Herrmann, B., & Johnsrude, I. S. (2020). A model of listening engagement (MoLE). Hearing Research, 397, 108016. https://doi.org/10.1016/j.heares.2020.108016

Kahneman, D. (1973). Attention and effort. Prentice-Hall.

McGarrigle, R., Munro, K. J., Dawes, P., Stewart, A. J., Moore, D. R., Barry, J. G., & Amitay, S. (2014). Listening effort and fatigue: What exactly are we measuring? A British Society of Audiology Cognition in Hearing Special Interest Group white paper. International Journal of Audiology, 53(7), 433–445. https://doi.org/10.3109/14992027.2014.890296

Ohlenforst, B., Zekveld, A. A., Jansma, E. P., Wang, Y., Naylor, G., Lorens, A., Lunner, T., & Kramer, S. E. (2017). Effects of hearing impairment and hearing aid amplification on listening effort: A systematic review. Ear and Hearing, 38(3), 267–281. https://doi.org/10.1097/AUD.0000000000000396

Peelle, J. E. (2018). Listening effort: How the cognitive consequences of acoustic challenge are reflected in brain and behavior. Ear and Hearing, 39(2), 204–214. https://doi.org/10.1097/AUD.0000000000000494

Pichora-Fuller, M. K., Kramer, S. E., Eckert, M. A., Edwards, B., Hornsby, B. W. Y., Humes, L. E., Lemke, U., Lunner, T., Matthen, M., Mackersie, C. L., Naylor, G., Phillips, N. A., Richter, M., Rudner, M., Sommers, M. S., Tremblay, K. L., & Wingfield, A. (2016). Hearing impairment and cognitive energy: The Framework for Understanding Effortful Listening (FUEL). Ear and Hearing, 37(Suppl 1), 5S–27S. https://doi.org/10.1097/AUD.0000000000000312

Ronnberg, J., Lunner, T., Zekveld, A., Sorqvist, P., Danielsson, H., Lyxell, B., Dahlstrom, O., Signoret, C., Stenfelt, S., Pichora-Fuller, M. K., & Rudner, M. (2013). The Ease of Language Understanding (ELU) model: Theoretical, empirical, and clinical advances. Frontiers in Systems Neuroscience, 7, 31. https://doi.org/10.3389/fnsys.2013.00031

Wild, C. J., Yusuf, A., Wilson, D. E., Peelle, J. E., Davis, M. H., & Johnsrude, I. S. (2012). Effortful listening: The processing of degraded speech depends critically on attention. The Journal of Neuroscience, 32(40), 14010–14021. https://doi.org/10.1523/JNEUROSCI.1528-12.2012

Winn, M. B., Edwards, J. R., & Litovsky, R. Y. (2015). The impact of auditory spectral resolution on listening effort revealed by pupil dilation. Ear and Hearing, 36(4), e153–e165. https://doi.org/10.1097/AUD.0000000000000145

Zekveld, A. A., Kramer, S. E., & Festen, J. M. (2010). Pupil response as an indication of effortful listening: The influence of sentence intelligibility. Ear and Hearing, 31(4), 480–490. https://doi.org/10.1097/AUD.0b013e3181d4f251