Every time we publish a post on this blog, we run an SEO audit script that checks the title, meta description, canonical URL and hreflang tags. On the English posts, one number keeps showing up: 164 chars. The theme is written to cap meta descriptions at 160 characters, yet the audit keeps reporting more. Is the theme broken, or is the audit wrong? Neither. They are simply counting two different strings.
This post uses that “phantom overflow” as a way into two questions: why meta descriptions have a length guideline at all, and how many different things “length” can mean.
What a meta description is (and is not)
A meta description is a short summary placed in the page <head>:
<meta name="description" content="A one- or two-sentence summary of this page." />
It is not shown on the page itself. It is mainly used as a candidate for the snippet under the title in search results, and often as preview text when a URL is shared (many sites put the same text in og:description).
Two things are worth keeping in mind. Google has said the meta description is not used as a ranking signal, and it is not guaranteed to be displayed — search engines frequently pull a different passage from the page to match the query. Its value is on the click side: when it is shown, it helps a reader decide whether to open the page.
Who decides the length limit?
You will see advice like “keep it under 155–160 characters.” But neither the HTML specification nor Google’s documentation sets a maximum length. A 1,000-character description is valid HTML and will not be treated as an error.
The guideline comes from display width. Snippets are truncated with “…” to fit the space on the results page, and that truncation is based on pixels, not characters. Wide characters (Japanese full-width text takes roughly twice the width of a Latin letter) get cut sooner, mobile results tend to be shorter than desktop ones, and the layout itself changes over time. The “limit” is a rule of thumb, not a rule. The practical takeaway is simple: put the important information first, so the text still makes sense if the end is cut off.
How this blog’s theme generates descriptions
Our theme, wpmm-blog, builds the description in a function called wpmm_blog_meta_description(). For a single post, it does roughly this:
$excerpt = has_excerpt() ? get_the_excerpt() : '';
if (!$excerpt) {
$body = wp_strip_all_tags(strip_shortcodes($post->post_content));
$excerpt = trim(preg_replace('/\s+/u', ' ', $body));
}
$desc = wp_trim_words($excerpt, 60, '…');
// cap at 160 characters (characters, not bytes)
if (mb_strlen($desc, 'UTF-8') > 160) {
$desc = mb_substr($desc, 0, 158, 'UTF-8') . '…';
}
And it outputs it like this:
echo '<meta name="description" content="' . esc_attr($desc) . '" />';
One subtle point: wp_trim_words() means different things in different languages. In English it counts 60 words. On our Japanese site, WordPress’s translation settings make it count characters instead, because Japanese does not separate words with spaces. So Japanese descriptions come out at 60 characters plus “…” (61), while English descriptions usually exceed 160 characters at 60 words and hit the second cap: 158 characters plus “…” (159).
Where 164 comes from: esc_attr() and the apostrophe
Here is the actual output from our previous post on robots.txt:
<meta name="description" content="When a search engine crawler visits a site, it usually doesn't jump straight to your articles. It first checks robots.txt to learn where it may go, and it rea…" />
The apostrophe in doesn't has become ' — an HTML entity. Inside an attribute value, characters like ", ' and < could be confused with markup, so esc_attr() replaces them with safe notation just before output (we covered this “escape on output” approach in our esc_html()/esc_attr() post).
| Character | Entity | Added length |
|---|---|---|
& |
& |
+4 |
< |
< |
+3 |
> |
> |
+3 |
" |
" |
+5 |
' |
' |
+5 |
One apostrophe adds five characters: 159 + 5 = 164. The audit script extracts content="..." from the raw HTML with a regular expression and measures the still-escaped string. Browsers and search engines decode entities when they parse the page, so readers see 159 characters. English, with all its contractions, triggers this on nearly every post; our Japanese posts rarely do.
A deeper version of the trap
The English esc_html()/esc_attr() post also audits at 164 characters, but decodes to only 156:
content="... Instead you see echo esc_html( $title ); or <a h…"
That < was not added by esc_attr(). The post body contains a code example, so the stored HTML already had <a href.... wp_strip_all_tags() removes tags but leaves entities as text, so the 160-character check counted < as four characters. Step by step:
- At the cap check: 158 characters (including
<as 4 and'as 1) plus “…” = 159 - After
esc_attr():'becomes'(+5) = 164, the audited value.esc_attr()does not double-encode the existing< - After decoding: −5 and −3 = 156, what readers actually see
The cap was measuring a string that already contained entities, so posts with lots of code produce shorter visible descriptions than intended. It is harmless here, but it shows that you need to know which stage of the string you are measuring.
“Length” has at least four meanings
For the robots.txt post, measured on the live pages:
| Measure | English | Japanese |
|---|---|---|
| Characters in HTML source (escaped) | 164 | 61 |
| Characters after decoding | 159 | 61 |
| UTF-8 bytes (decoded) | 161 | 177 |
| Display width | depends on device and font | depends on device and font |
The Japanese description is 61 characters but 177 bytes, because most Japanese characters take three bytes in UTF-8 (the “…” does too). Languages disagree as well: PHP’s strlen() counts bytes while mb_strlen() counts characters, and JavaScript’s .length counts UTF-16 code units, so some emoji count as two. That is why the theme comment insists on “characters, not bytes”: truncating Japanese with a byte-based substr() can cut a character in half and emit broken output.
If the goal is to check what readers see, decode first, then count:
import html
raw = "it usually doesn't jump straight"
print(len(raw)) # 37 (escaped)
print(len(html.unescape(raw))) # 32 (what readers see)
Our audit still reports the escaped length, so we treat a small overshoot on English posts as a known, acceptable difference rather than something to “fix” by trimming text.
Common pitfalls
- Treating HTML-source length as visible length — entities inflate the count
- Capping an already-escaped string — the visible result ends up shorter than planned
- Truncating multibyte text by bytes — use
mb_substr()in PHP - Ignoring mid-word cuts — character-based truncation produces endings like
it rea…; a hand-written excerpt avoids this - Assuming the description will be shown as written — it is a candidate, not a guarantee
- Forgetting language differences —
wp_trim_words()counts words in English and characters in Japanese
Summary
The meta description “limit” is a guideline derived from display width, not a hard rule, so the most useful habit is to front-load the key message. When you work with lengths, be clear about escaped versus decoded, characters versus bytes, and which stage of processing the string is at. Our recurring “164” was never a broken cap; it was HTML’s safe notation for an apostrophe being counted as five extra characters. Check what a number actually measures before acting on it.