Skip to content

How the scoring works

Every number on this site is measured from the text of the message it describes. Nothing is model reported, and nothing is invented. This page explains each measurement, states what it cannot do, and then scores a real message from the corpus in front of you.

The cliche score

Every message is checked against a fixed list of 201phrases this genre runs on: “you complete me”, “my other half”, “to the moon and back”, “words cannot express”. The score is the share of the message covered by those phrases. Overlapping matches count once, so the number is a proportion of the message rather than a tally.

A high score is not a judgement about how someone feels. It measures how many other people have already sent the same words, which is a different and more useful thing to know when you are about to send one.

What it does not do: the list is fixed, so the score catches known cliches and nothing else. A sentence that is novel and hollow scores zero. This is not an originality metric and is never presented as one.

Text message segments

A plain text message fits 160 characters in one segment. That limit is not about character count in the way people assume: it is about encoding. Standard Latin text uses a 7-bit alphabet, but the moment a message contains one character outside it, including any emoji, the whole message switches to a 16-bit encoding and the limit drops to 70.

So a 68-character message with one emoji still sends as one text. A 71-character one does not. Every message on this site is measured against the real rules, which is why the counts occasionally look surprising.

The other numbers

Word and character counts, sentence counts, and read-aloud time are direct measurements. Reading level uses a standard readability formula with a syllable heuristic that is right about 90 to 95 percent of the time per word: fine for the collection averages shown here, not a substitute for a pronouncing dictionary.

Concreteness is the density of common physical nouns. It is the single best signal that a message was written about a real person rather than assembled from superlatives, which is why collections are sorted the way they are.

One message, scored

This is a real message from the corpus — the most cliched one currently published, chosen so there is something to highlight. The marked phrases are its stored cliche hits; everything below them is computed from the text you are reading.

I cannot imagine my life without you; your small, steady presence is my anchor.
2 phrases from the fixed list: “i cannot imagine my life without you”, “my anchor”
Characters
79
Words
14
Sentences
1
SMS segments
1 (GSM-7 encoding)
Emoji
0
Read aloud
6s
Flesch reading ease
65.73
Grade level
7.57
Warmth
14%
Cliche score
64%
Concreteness
0%
Direct address
4%
Fits one text (160)
yes
Fits a tweet (280)
yes
Fits a card (250)
yes
Fits an Instagram caption (2,200)
yes

The corpus right now

Across the 10,612 messages currently published, the average cliche score is 0%, the average message is 17 words, and the average reading level is grade 6.

Computed from the live corpus when this page loaded, not written in advance.

The list itself

A sample of the phrases every message is checked against. The full list has 201 entries, it is hand-maintained, and it is fixed: the same list scores every message, so two messages with the same score earned it the same way.

  • you complete me
  • you are the one
  • destiny brought us together
  • butterflies in my stomach
  • my heart melts
  • more than words can say
  • my guiding light
  • i will never let you go
  • you light up the room
  • you are perfect
  • a dream come true
  • lost in your eyes
  • thinking of you always
  • against all odds
  • my ride or die
  • each passing day
  • i love you more
  • you saved me
  • burning desire
  • i feel alive when i am with you

How messages get here

The messages are written by a language model, offline, and then measured by this code. Nothing is generated when you load a page, and no message is a quotation from a named author.

A collection only appears once enough messages genuinely match it, and every one has to pass three duplicate checks first: an exact match after normalisation, a similarity check against the entire corpus, and a per-collection check that stops the same sentiment appearing thirty times in different words.

More about who publishes this and what it does not claim is on the about page.

Method last reviewed .