
The verbal fluency test asks a patient to generate as many words as possible in 60 seconds, either starting with a given letter (letter fluency, commonly F-A-S) or belonging to a given category (semantic/category fluency, commonly "animals"). Letter fluency taxes executive strategic search; category fluency taxes semantic memory access. Both are scored by raw word count converted to age- and education-normed T-scores.
What the Verbal Fluency Test Measures
Verbal fluency is a timed word-generation task with two common variants that probe different cognitive systems:
- Letter (phonemic) fluency — the patient names as many words as possible starting with a specified letter (typically three trials: F, A, S), excluding proper nouns, numbers, and repeated words with only a suffix change. Because there's no obvious semantic organizing principle, this variant leans on executive function — self-generated search strategy, cognitive flexibility, and inhibition of intrusion errors.
- Category (semantic) fluency — the patient names as many items as possible within a category (most commonly "animals," sometimes "fruits" or "supermarket items") in 60 seconds. This variant leans more heavily on semantic memory network integrity and access, which makes it comparatively more sensitive to early semantic memory decline in dementia workups than letter fluency is.
Both variants are brief, low-burden, and language-dependent — the F-A-S letter set is specific to English; other languages use different, psychometrically matched letter sets. A verbal fluency score in isolation says nothing about why output was reduced; it needs to be interpreted alongside the rest of a battery to distinguish an executive-search deficit from a semantic-memory deficit from a word-finding/language deficit.
Administration Time
Each trial runs exactly 60 seconds. A standard letter fluency administration (three trials: F, A, S) takes about 3 minutes of pure task time; adding a category fluency trial brings total task time to roughly 4-5 minutes. Instructions, practice, and recording of responses (including intrusions and perseverations) add a few minutes on top, so plan for 5-10 minutes total within a larger battery.
Scoring and Interpretation
Raw scores are the total number of correct, non-repeated, non-intrusion words produced per 60-second trial (letter fluency is usually summed across the three F-A-S trials; category fluency is scored per category). Raw totals are converted to age- and education-normed T-scores or z-scores using a published normative dataset.
There is no single universal pass/fail cutoff for verbal fluency. Interpretation depends entirely on which normative reference is used — Heaton norms and other demographically corrected datasets produce different expected ranges, and scores from different normative sources are not directly interchangeable. In most protocols, a T-score meaningfully below the normative mean (commonly framed as below average or impaired relative to same-age, same-education peers) is what's flagged, not a fixed raw number. Clinicians should also track qualitative features — perseverations (repeating a prior response), intrusions (off-category words), and clustering/switching patterns — which can carry diagnostic signal independent of the total count.
Because category fluency draws more on semantic memory networks, a pattern of disproportionately low category fluency relative to letter fluency is sometimes noted as more consistent with semantic memory network involvement (as seen in some dementia syndromes), while proportionally low letter fluency with preserved category fluency points more toward executive/frontal-systems dysfunction. This is a pattern to note, not a standalone diagnostic rule.
Clinical Use Case
Verbal fluency is a standard component of general neuropsychological batteries and is used in:
- Post-concussion / TBI evaluation — as part of a broader executive function and language assessment when cognitive-health concerns persist beyond the acute injury window.
- Dementia and MCI workup — category fluency in particular is sensitive to early semantic memory decline and is commonly paired with global screens; see MCI of unclear cause.
- Differentiating executive vs. semantic-memory contributions to reduced verbal output, which helps route a patient toward the right follow-up (cognitive rehab, further language workup, or referral).
It is rarely administered alone — it's almost always one component test within a battery that also covers memory, attention, and processing speed, such as Trail Making Test B, RAVLT, or Digit Span.
CPT Code for Administering Verbal Fluency
Verbal fluency is not billed as a standalone code. As a component test within a neuropsychological battery, its administration and scoring falls under the standard test-administration codes based on who administers it:
- 96136/96137 — administration and scoring by a physician or other qualified health professional (QHP), first 30 minutes / each additional 30 minutes, for two or more tests by any method.
- 96138/96139 — the same administration and scoring, performed by a trained technician, first 30 minutes / each additional 30 minutes.
The interpretive work — integrating the fluency result with the rest of the battery, writing the report, and delivering feedback — falls under 96132/96133 (neuropsychological testing evaluation services, physician/QHP, first hour / each additional hour). Whether a technician can administer fluency testing under general supervision, and how that interacts with documentation requirements, is addressed in Technician-Administered Testing Supervision Rules. If a physician or nurse practitioner without a psychology background is billing this battery, see Billing Neurocognitive Testing as a Non-Psychologist.
Where Verbal Fluency Sits in the Kavera Protocol
Verbal fluency is one instrument within Kavera's broader assessment battery, which spans the Concussion, Mental Health, Cognitive Health, and Headache modules. Within the Cognitive Health module, it's typically deployed alongside other executive-function and processing-speed measures — see Executive Function — to build a between-visit picture of cognitive trajectory that the treating clinician reviews at follow-up.
How Kavera Handles This
Verbal fluency is administered in the office and scored in Kavera by letter and category, stored against baseline for serial comparison in the Cognitive Health module. The record supports 96132 and 96138 documentation. Self-Serve practices run this with their own staff. On Managed, Juliet Mott's team runs it and bills it under your credentials.
FAQ
What is the difference between letter fluency and category fluency?
Letter fluency (e.g., F-A-S) asks patients to generate words starting with a specific letter and relies more on executive/strategic search. Category fluency (e.g., "animals") asks patients to generate words within a semantic category and relies more on semantic memory access. Both are commonly administered together for a fuller picture.
Is there a normal score on the verbal fluency test?
There is no single universal normal score. Raw word counts are converted to age- and education-normed T-scores or z-scores using a specific normative dataset, and what counts as "below average" depends on which norms are used. Scores should always be interpreted relative to demographically matched peers, not a fixed number.
How long does the verbal fluency test take to administer?
Each trial is exactly 60 seconds. A standard three-trial letter fluency administration takes about 3 minutes of task time; adding a category fluency trial and instructions typically brings the full administration to 5-10 minutes.
What CPT code covers verbal fluency testing?
Administration and scoring fall under 96136/96137 (physician/QHP) or 96138/96139 (technician), and interpretation/report falls under 96132/96133, since verbal fluency is billed as part of a broader neuropsychological battery rather than as a standalone code.
Can a low verbal fluency score alone diagnose a cognitive impairment?
No. Verbal fluency is one component test and should always be interpreted alongside the rest of the battery and clinical history. A single low score can reflect fatigue, mood, language background, or test-taking factors as well as an underlying cognitive issue.
See this on your own patient population
One field. 30 minutes. Live demo with a clinician.