Test your website →

Blog

Reading Level Tools and How to Actually Use Them

Reading level scores are useful instruments, but only if you know what they measure and what they miss.

There’s a moment that happens to almost everyone who discovers readability formulas. You paste your homepage copy into a tool, it tells you your text reads at a twelfth-grade level, and you feel a small jolt of alarm. Then you rewrite a few sentences, the number drops to eighth grade, and you feel like you’ve fixed something. The question is whether you actually have.

I think about this a lot because I built a tool, GazeSite, that grades websites, and readability is one of the things it grades. So I’ve had to think hard about what these scores mean and what they don’t. The honest answer is that reading level formulas are crude instruments. But crude instruments can still be useful, the way a bathroom scale is useful even though it can’t tell muscle from fat. You just have to know what you’re measuring.

Most readability formulas, including the famous Flesch-Kincaid one, look at exactly two things: how long your sentences are and how long your words are. That’s it. They don’t know whether your argument makes sense. They don’t know whether your words are the right words. They count syllables and periods. A formula would score “Colorless green ideas sleep furiously” as wonderfully readable, because the sentence is short and the words are common. It’s also nonsense.

This tells you something important about how to use these tools. They’re not measuring comprehension. They’re measuring two of the most common causes of poor comprehension. Long sentences make readers hold more in their heads at once. Long words are, statistically, rarer words, and rare words make readers stop and think. When a formula says your text reads at a college level, what it’s really saying is: your sentences are long and your vocabulary is fancy. That’s often a fair proxy for “hard to read,” but it’s a proxy, not the thing itself.

The mistake people make is treating the score as a target instead of a diagnostic. If your goal becomes “get the number to eighth grade,” you can hit it by mechanically chopping sentences in half and swapping every three-syllable word for a shorter one. Sometimes this improves the writing. Often it produces choppy, dumbed-down prose that’s actually harder to follow, because you’ve amputated the connective tissue between ideas. “We reduced costs. This helped customers. They stayed longer” scores beautifully. But the reader now has to reconstruct the causal chain you deleted.

The better way to use a readability score is as a smoke detector. When it goes off, you don’t spray water at the detector. You go find the fire. A high score means somewhere in your text there are sentences doing too much work. So you go looking for them. Usually you’ll find a sentence that started life as one idea and accumulated three more through revisions, each attached with a comma or a “which.” The fix isn’t to shorten it mechanically. The fix is to figure out what you were actually trying to say and say it in the order the reader needs to hear it.

There’s another subtlety worth understanding. Reading level and reader intelligence are not the same axis. People assume writing at an eighth-grade level means writing for eighth graders. It doesn’t. It means writing for people who are doing something else. Your website visitor is on a phone, in line for coffee, half-listening to a conversation, giving you maybe a quarter of their attention. A neurosurgeon skimming your pricing page has the reading capacity of a distracted teenager, because that’s how much capacity she’s allocating to you. Simple prose isn’t a concession to dumb readers. It’s a concession to the actual conditions under which reading happens on the web, which is to say, bad ones.

This is also why I’m skeptical of running readability tools only on blog posts and long-form pages. The text that matters most on a website is usually the shortest: headlines, button labels, form instructions, error messages. Formulas do badly on short text because the statistics get noisy with small samples. But the underlying principle still applies. “Commence your complimentary evaluation” and “Start your free trial” mean the same thing, and one of them makes the reader work. No formula will flag a three-word button, so you have to develop the ear yourself. The formula is training wheels for that ear.

If you want a practical routine, here’s what I’d suggest, and it’s roughly what I try to do myself. Run the score once to find out where you stand. If it’s high, read your text aloud and notice where you run out of breath or stumble. Those spots and the formula’s complaints will usually coincide. Fix the ideas, not the syllables. Then run the score again, not to hit a target, but to confirm the trend went the right direction. And then stop. Past a certain point, optimizing the number is like optimizing your weight by cutting off a limb. The scale says you succeeded.

The tools are worth using. I built one into my product because most site owners have never once measured this, and the first measurement is where nearly all the value is. It’s the difference between not knowing you have a problem and knowing. But after that first reading, the tool has done most of its job. The rest is writing, and writing is the old problem it always was: knowing what you mean, and saying it plainly, one idea at a time.

More articles

← All posts