Guide
How to measure brand voice
Twenty-two things you can count. Four of them tell you anything.
Brand voice is usually assessed the way wine is: somebody reads it and reaches for adjectives. That is fine for admiring and useless for managing, because an adjective cannot be checked on next week's draft. Voice can be counted. The surprise, when it was, is how many of the obvious counts turned out to reward the wrong thing.
The experiment
A control written to beat real writing
The method behind smarticulate's score was built by measuring six real businesses across five sectors, then writing a control passage designed to sound competent and mean nothing, and asking which measurements told them apart.
Several of the expected ones did not. The control used more contractions than two of the three most distinctive real voices. Its sentences were longer than all of them. It used no long dashes, which is exactly like every real voice measured, and disposes of the internet's favourite AI tell in one line. Any score built on those three would have ranked the fake above the real writing.
Four measures separated them every time. They carry all the marks. Eighteen others are worth reporting, and worth nothing as a test.
A measure that flatters a fake is not a test.
The four
What actually separates a voice from the average
- Own words. How much of the text came from the phrasebook every business in the sector shares: the two or three hundred phrases that read like a proposal template and that nobody could attribute to anybody. The single most reliable signal, the one that costs most sites their marks, and the one ChatGPT, Claude, Copilot and Gemini all push hardest in the wrong direction, because the phrasebook is what they have read most of.
- Rhythm. Sentence-length variance. Real voices measured between 7.7 and 9.7; the control scored 4.3, and the ranges never overlapped. Writing that clusters around one length reads as wallpaper however good the individual sentences are.
- Landing. The share of sentences at five words or fewer. Real voices ran between 15% and 29%. The control managed 0%. Roughly one in six is the target, placed where the argument turns.
- Evidence. Figures, dates, names and results. Of three sites measured while building the method, two contained no figure of any kind, and one of those led with the word evidence in its headline. That is a claim arguing against itself.
The weights are 35, 30, 25 and 10, and the bands run from established voice at the top to no voice present at the bottom. The full definitions, the calibration ranges and the refusal rules are published, because a number without a published basis is an assertion.
Doing it by hand
A spreadsheet and an hour
You can get most of the way without any tool. Take a page of prose you have published, strip the headings and the bullet fragments, because they are furniture rather than writing, and count.
- Split it into sentences and record the word count of each. The median is your typical sentence; the standard deviation is your rhythm. Under about five, the page is flat.
- Count the sentences of five words or fewer and divide by the total. Under 10%, nothing lands.
- Highlight every phrase you could imagine on a competitor's site. Count them per hundred words. That is your borrowed vocabulary, and it is usually worse than expected.
- Count every sentence carrying a figure, a date, a name or a result. Fewer than one in ten and the page is asking to be trusted rather than showing why.
- Do all four again for each section separately. A page that scores well overall can change costume entirely when it starts selling: one measured business went from a variance of 10.8 in its origin story to 6.2 on its services.
The free test does the same counting in about a minute, in your browser, and quotes back the sentences that are genuinely yours. Under 25 sentences the rhythm and landing figures get noisy, and it says so rather than pretending.
What the number is not
Four honest limits
- Not a quality score. It cannot tell whether what you said is true or worth saying.
- Not an AI detector, and nothing claiming to be one works. It measures distinctiveness, never authorship. It does not know or care whether you typed the words or prompted ChatGPT, Claude, Gemini or Copilot for them, and neither does your reader.
- Not a grammar or readability check. Those measure ease. This measures whether the writing could have come from anybody.
- Not an opinion. No model is consulted at any point in the scoring, which is what makes comparing two businesses defensible: identical code, same day, same answer next Tuesday.
A higher score is also not a bigger business. A firm can score badly here and out-sell everybody on price, reach or the work itself. What the number tells you is how hard it is for a buyer to tell you apart, and that is a diagnosis rather than a verdict.
Count first. Describe second. Everything else is wine tasting.
Measure a page of your own
Paste something you have published and get all four scored measures, the eighteen reported ones, and one thing to fix. If the answer is that nothing needs fixing, you will be told that too.
More guides: how to write brand voice guidelines, ChatGPT custom instructions for your brand voice, tone of voice examples.