Accuracy & Methodology
How syllable counting and literary analysis actually work on this site, written in enough detail that you can judge the results for yourself.
Reviewed for accuracy · Last updated August 4, 2026
How syllable counting works
Every word passes through up to three layers, in order:
- Dictionary lookup. The word is checked against a pronunciation dictionary derived from the CMU Pronouncing Dictionary (~126,000 entries), which records real, attested pronunciations rather than guessing from spelling. Words with more than one accepted pronunciation (like "fire" or "every") keep every valid count.
- Curated overrides. A small, hand-reviewed table corrects or extends specific entries — for example, adding the common fast-speech pronunciation of "chocolate" that the base dictionary doesn't capture. This table grows from user-submitted corrections after editorial review.
- Rule-based fallback. Words in neither the dictionary nor the override table (proper nouns, coinages, rare compounds) are estimated using vowel-group counting with adjustments for silent final "e," "-ed" and "-es" endings, "-le" endings, hyphenation, and common prefixes.
Numbers and acronyms get dedicated handling: numbers are estimated by how they'd be read aloud ("1998" → "nineteen ninety-eight"), and all-caps acronyms are either matched against a short list of ones commonly said as a word (like "NASA") or spelled out letter by letter.
Every result carries a confidence label — high, medium, low, multiple pronunciations, or manually corrected — so results are never presented as more certain than they are.
How literary analysis works
Structure and syllables are checked deterministically. Literary analysis — seasonal words, imagery, juxtaposition, and the haiku-vs-senryu estimate — works differently:
- Seasonal words (kigo) are matched against a curated reference list tagged by season and confidence. The list is a starting point, not exhaustive — a missing match doesn't mean a poem has no seasonal quality, just that it isn't in the current list.
- Sensory imagery is detected via word lists across sight, sound, smell, touch, taste, movement, and temperature.
- Abstract vs. concrete language flags words and phrases (like "beautiful," "I feel," "very") that tend to state rather than show.
- Cuts (juxtaposition) are detected primarily through end-of-line punctuation (a dash, comma, colon, semicolon, or ellipsis), with a softer fallback that checks whether the first line shares no words with the rest of the poem.
- Haiku vs. senryu is estimated by weighing nature/seasonal word matches against everyday human/social word matches, always reported with a confidence percentage rather than a flat verdict.
None of this is AI-generated — every check is a documented word list or structural rule, which means the same input always produces the same output.
How scoring works
The Haiku Score is a 100-point rubric split across seven categories: structure (25), syllable accuracy (20), imagery (15), seasonal context (10), juxtaposition (10), brevity and clarity (10), and natural language (10). Each category is computed from a documented formula — never a random number — and Modern Haiku mode specifically avoids penalizing a poem for not matching 5-7-5, scoring syllable accuracy against a typical brevity range instead.
The Haiku Rater uses a separate, more literary rubric (originality, sound, ending strength, and more) built the same way — documented rules, not a black box.
Known limitations
- Rare proper nouns, brand names, and invented words rely on an estimate, not a dictionary lookup.
- The pronunciation dictionary reflects North American English; other dialects may pronounce some words differently.
- The seasonal word list reflects common English-language haiku convention and isn't exhaustive or globally universal — seasons and their associations vary by region.
- Literary analysis (imagery, juxtaposition, haiku-vs-senryu) is a structured estimate, not a substitute for human editorial judgment — poetry interpretation is inherently subjective.
- The cut/juxtaposition detector relies heavily on punctuation; a cut created purely through subtle phrasing without any punctuation may be missed.
Review process
Anyone can report a syllable count they believe is wrong, including a suggested count, pronunciation, and reason. Reports are queued for editorial review before being added to the override table — nothing is added to the live dictionary automatically. See Contact to submit a correction.
Update history
- August 4, 2026 — Initial launch: syllable engine, haiku analysis engine, Haiku Checker, Syllable Counter, Haiku Generator, and Haiku Rater.
Frequently asked questions
What's the source of the pronunciation dictionary?
The CMU Pronouncing Dictionary, a long-standing, freely available machine-readable dictionary of North American English pronunciations maintained by Carnegie Mellon University, covering roughly 126,000 words and their common variant pronunciations.
What happens with words the dictionary doesn't have?
They're estimated using a rule-based algorithm that accounts for vowel groups, silent letters, common endings, and hyphenation, and are always labeled with lower confidence so you know it's an estimate.
How is literary analysis (imagery, senryu, etc.) different from syllable counting?
Syllable counting is checked against a dictionary and is close to deterministic. Literary analysis is pattern-based — word lists and structural heuristics — and is inherently more approximate, which is why it's always presented as an estimate.
Does the site use AI to count syllables?
No. Syllable counting never relies on AI — it's dictionary lookup plus a documented rule-based fallback, so results are consistent and explainable.