How the score is measured
What the number is, how it is estimated, how sure it is, and what nobody has measured yet.
The short version
Your SpellingIQ is a skill rating for spelling on a 300–2200 scale. Every word in the corpus carries a difficulty on the same scale, learned from how often players spell it. Your rating rises when you beat words above it and falls when words below it beat you, weighted by the gap, the way chess ratings and ranked games work. A short adaptive test estimates it in about ten words and reports a likely range beside it; rated practice then moves it a little at a time. It is not an intelligence test, and one number cannot say which kind of spelling knowledge you are strong or weak in.
The scale
Ratings run from 300 to 2200. Most adults first place around 900–1200. The levels the product acts on, which are also the Daily Five levels, are Starter 300–500, Learner 500–900, Normal 900–1200, Skilled 1200–1500, Expert 1500–1800, Genius 1800+.
Words and players share one scale on purpose. The distance between a player and a word is the model's prediction of the outcome: a word rated the same as you is one you would spell about half the time; a word 400 points below you, about nine times in ten. That one rule is what lets the same number pick the Daily Five level, aim practice at the edge of your ability, and be updated by every graded attempt.
How word difficulty is set
A new word enters the corpus with a starting difficulty estimated when it was authored. From then on the players decide. Every graded practice attempt moves the word as well as the player: a word that beats strong spellers climbs, a word that everyone gets falls, and a word with few attempts behind it moves faster than one with hundreds. Difficulty is therefore a measurement of this corpus against these players, not a dictionary's opinion, and it keeps moving as the player base changes.
The placement test types the word cold from audio, which is a harder task than practice's letter bank (arranging a given set of letters), and the two are not blurred into one number: responses from the test are folded into a separate free-recall difficulty on a schedule, and the estimator corrects for the format gap when it scores you. The Learn section explains why the two formats are different tests in what makes spelling practice work.
How the placement test estimates your rating
The test uses an item-response model: it assumes that how likely you are to spell a word depends on the gap between your ability and the word's difficulty, on a logistic curve. After each answer the whole estimate is recomputed from every answer so far, so the order the words arrived in makes no difference and the estimate can move either way at any point. The next word is aimed near the current estimate, because a word you would spell about half the time is the one whose outcome says most.
The model allows for slips: a typo or a misheard word is treated as a possibility on every answer, and so is knowing one hard word well above your general level. That bounds how far any single answer can move the estimate, which is what stops one fat-fingered key on an easy word from sinking the score.
The number you are shown is the centre of the model's estimate and the likely range is its 95% interval: the band the model thinks your true rating is very likely to sit in, given the answers so far. The test stops once that range is narrow enough, after at least 8 words and at most 12. Eight well-aimed words leave a range of roughly ±290 points, and each further word narrows it less than the last, which is why the test does not run to twenty.
How practice moves it
Each first attempt at a rated practice word updates your rating the way chess does: the model predicts your chance of spelling the word from the gap, and the rating moves by the difference between what happened and what was predicted. Beating a word you were expected to beat moves it a little; beating one you were not moves it a lot; the same in reverse for a miss.
The size of each move shrinks as rated words accumulate, in steps. The score page names them: a rating is provisional under 12 rated words, settling under 30, steady under 100, and established from 100. A rating claimed from the placement test and never practised is marked placed. Nemesis reviews, second and third attempts at a word, and the Daily Five are all graded but none of them changes the rating; the Daily has its own daily percentile instead.
How accurate it is
Measured by simulation, not yet on real players. The test suite that runs on every change to the estimator simulates thousands of sessions against a corpus shaped like ours and measures the result against the ability each simulated speller was given. Measured in September 2026, the placement estimate's error was about 134 rating points root-mean-square across the ability range, its likely range contained the true value about 96 times in 100, and sessions averaged just under ten words. What the suite enforces is looser than what it measured: it fails the build if the error exceeds 160, if the range covers the truth less than 90 times in 100, or if sessions average eleven words or more. So the figures above are a dated reading, and the bounds are the guarantee. The previous estimator, a bracket search, scored 192 on the same measure, and the model described above replaced it for that reason.
A simulation can only be as honest as its assumptions, and the biggest one is that word difficulties are right. That is what the calibration above is for, and it is the reason the first real-player figures, below, are the ones that will matter.
What the test measures, and what it does not
It measures one thing: producing the spelling of a word you have just heard, by typing it, with no autocorrect and no options to choose from. That is recall, not recognition, and it is closer to what fails people in a document than a multiple-choice test is. It is also typed, and typing adds noise; the slip allowance above is the model's answer to that, and it is a partial one. American English is the canonical spelling, with common variants accepted.
One number cannot separate the kinds of knowledge spelling draws on: sound-to-letter mapping, letter patterns, word parts, and origin. A player can be strong in one and weak in another and get the same rating as a player with the opposite profile. The Learn section covers what those layers are in how people learn to spell. And the rating is not an intelligence measure or a diagnosis: a learner with a suspected reading or spelling difficulty deserves a qualified professional, not a game score.
What we do not know yet
- How ratings are distributed across players. The Daily Five reports a percentile within its level; the rating itself does not yet, because the player base is not large enough for a distribution to mean anything.
- How much a rating moves between two tests taken close together (test-retest reliability) on real players. The simulated range is an estimate of this, not a measurement of it.
- How the rating relates to established spelling assessments. That needs a study, and none has been run.
Each will be published here once the data exists. A number that has not been measured will not be quoted, on this page or anywhere else on the site.
Read more
The principles behind these choices are on the philosophy page; the research they rest on is in the Learn section, with references. For the product itself, start with about, play the Daily Five, or take the placement test.