Two thousand three hundred and ninety ways to say twenty-nine things.

Hundreds of free-text business descriptions already contained the site's service taxonomy — written in 2,390 spellings. Counting it beat designing it, and the evidence had to be graded before it was published.

A directory site I run lists a few hundred trade businesses, each with a free-text description they or someone before them wrote. I wanted to turn the services those descriptions mention into something you can click — pick a service, see everyone who offers it.

So I counted first. Three thousand four hundred and thirty-three service mentions, written in two thousand three hundred and ninety distinct spellings.

Not two thousand three hundred and ninety services. Twenty-nine. Everything else was spelling, word order, compound nouns, regional preference, and the ordinary human habit of describing the same job four different ways in one paragraph.

The taxonomy already existed

The temptation with a project like this is to sit down and design the categories. You know the industry, you have opinions, you can produce a clean list of twenty in an hour and it will look excellent.

It will also be yours. And the gap between the words the people who build a thing use and the words the people who use it use is the oldest problem in information architecture. It is the whole reason card sorting exists as a method — you stop guessing at the structure and watch the patterns emerge from what people actually do, because the people who design a product think about its content differently from the people who use it.

Here I did not need to run a study. Several hundred businesses had already told me, in their own words, what they do. The research had been sitting in the database for years, unread, because nobody had thought to count it.

Your categories are usually already written down. They are just spelled two thousand different ways.

Clustering those spellings into twenty-nine canonical services took a regular expression per service and a couple of hours of checking. Every single business ended up carrying at least one tag. Roughly fourteen percent of the individual lines stayed unassigned — mostly filler phrases, but with a few genuine niches hiding in there that nobody would have put on a designed list.

Strong evidence and weak evidence are not the same tag

The decision I am most glad we made was to split the evidence into two grades and only publish one of them.

A service named in the explicit services block is strong evidence — the business is claiming it. The same word appearing in a paragraph of prose is weak evidence, and it means something completely different. One landscaping term showed up as a strong claim for about a hundred businesses and as a passing mention for more than twice that many. Publishing the combined number would have listed hundreds of businesses under a service most of them do not offer.

That is the kind of quiet corruption that makes a directory useless. Not a bug — a definition that was too generous, applied consistently, at scale. So the pages list strong matches only, and the weak count is shown separately where it is informative.

It is the same discipline as measuring your live inventory before planning around it, which is how ten minutes of querying once killed a two-week content sprint. The alternative is building on an assumption and finding out in month three.

What the count then changes

Counting reshaped the plan twice.

First, the five biggest services turn out to apply to almost every business in the directory. As a filter they are worthless — a filter that returns everyone has filtered nothing. As search entry points they are the most valuable pages on the site, because that is what people type. So those pages got built with a geographic breakdown as their main content instead of a list of businesses. Same page type, different job, and only the numbers told us which was which.

Second, we deliberately did not build a formal taxonomy in the CMS yet. Populating hundreds of listings with terms and designing an archive template is a real project; generating the pages from the classification is a script. The mapping is finished either way, and it can be written into proper terms later unchanged. Doing the cheap version first is not laziness when the expensive version depends on the same output.

The icons taught me something I should have known

Each service got a small line icon. The first set was drawn at what looked like a sensible level of detail, and at display size it fell apart. Tree care read as a Venus symbol. Excavation read as a shopping trolley. Irrigation read as an umbrella. One of them read as the broken-image placeholder, which is a special kind of insult.

Three correction rounds, and the thing that finally worked was checking them against rendered pages at the real size rather than against the source files. Nielsen Norman Group have been saying the underlying thing for years: most icons are ambiguous to users, only a handful enjoy near-universal recognition, and a text label is what removes the ambiguity. Ours have labels, which is why the first round survived long enough to be embarrassing rather than harmful.

Test the artefact in the state the user meets it. Not the file, not the design canvas — the rendered page at the actual size. That rule has now cost me enough times that I write it down.

The unglamorous part

Nine hundred-odd pages of directory content, and the thing that nearly sank the launch was a stale sitemap cache that had frozen weeks earlier and quietly stopped listing anything new — including, it turned out, every page created since. Not caused by this work, but discovered by it, and it would have made the whole exercise invisible to search.

Which is the pattern, really. The interesting part was counting the words people already used. The part that decides whether any of it matters was a cache.

Before you design a structure, go and count the one your customers have already written for you.

Sources & further reading

External
Nielsen Norman Group — Card Sorting: articles, videos and research
Nielsen Norman Group — Card Sorting: Pushing Users Beyond Terminology Matches
Nielsen Norman Group — Icon Usability

Related posts
I planned a two-week content sprint. Ten minutes of querying killed it.
Six search impressions told me exactly what to write next.
Why pain points, not keywords, will decide your SEO future.

Subscribe to Remco Livain

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe
Work with me →×