Part of the gap between two practice scores is real learning, part of it is the form you happened to sit, and part of it is noise that no calculator can remove. On this site's Step 1 models, the same raw performance (140 of 200 correct) converts to anywhere from 206 to 211 depending on which form produced it. Treat any single sitting as one sample from a range, and treat the direction of three sittings as the signal.
The same raw score converts differently on every form
Each form on this site has its own conversion line: a starting point and a cost per wrong answer. Feed identical performance into all seven Step 1 forms and the estimates fan out.
| Step 1 form | 60 wrong of 200 (70% correct) | Cost per wrong answer |
|---|---|---|
| NBME 25 | 210 | 1.11 points |
| NBME 26 | 209 | 1.14 points |
| NBME 27 | 208 | 1.11 points |
| NBME 28 | 211 | 1.05 points |
| NBME 29 | 207 | 1.09 points |
| NBME 30 | 210 | 1.15 points |
| NBME 31 | 206 | 1.08 points |
That is a five-point spread produced by nothing except form choice. If you sat Form 28 in week one and Form 31 in week three, you could have improved slightly and still watched your number fall. The reverse trap is worse: finishing on Form 28 can flatter a result you would not repeat elsewhere.
You can reproduce the whole table yourself by running one raw score against every Step 1 curve side by side.
Cost per wrong answer decides how violently a score moves
The second column above matters more than most students realise. On Form 30, every missed question costs about 1.15 points; on Form 28, about 1.05. Ten careless errors, whether a phone buzzing or a rushed block, move the Form 30 estimate roughly 11.5 points and the Form 28 estimate roughly 10.5.
Steep forms exaggerate both good and bad days. That is a property of the conversion line, not a statement about how hard the questions were.
Put both sittings on one ruler before you compare them
Here is the trap in full, using two real curves from this site.
You sit Form 28 and miss 60 questions. The estimate is 211. Two weeks later you sit Form 31 and miss 58, which is genuinely better, and the estimate comes back 208. Three points lost for two questions gained.
Convert both raw counts on a single curve and the picture inverts. On Form 28's line, 60 wrong is 211 and 58 wrong is 213. On Form 31's line, 60 wrong is 206 and 58 wrong is 208. Either way the movement is +2, not -3. The only thing that changed between your two readings was the ruler.
This is the single most useful habit in this whole article: when two sittings disagree, re-convert both raw counts on one form before you interpret anything. It takes about thirty seconds and it removes the largest avoidable source of confusion in practice-score tracking.
Step 2 CK forms agree with each other, then disagree about the exam
The Step 2 CK picture inverts. Run 70 wrong of 200 through all seven CCSSA forms and the converted scores land within three points of each other, but the range this site attaches to them does not.
| Step 2 CK form | 70 wrong of 200 | Modelled range around the estimate | Realistic exam range |
|---|---|---|---|
| NBME 9 | 222 | ±9 | 222–242 |
| NBME 10 | 223 | ±6 | 228–235 |
| NBME 11 | 224 | ±6 | 229–236 |
| NBME 12 | 224 | ±6 | 229–236 |
| NBME 13 | 224 | ±6 | 229–236 |
| NBME 14 | 224 | ±5 | 221–229 |
| NBME 15 | 224 | ±5 | 221–228 |
Same estimate, very different meaning. Form 9's model treats the converted number as a floor and pushes the realistic range up to 20 points above it. Form 14's model barely shifts it at all. So a student who goes from Form 9 to Form 14 and sees "no improvement" in the three-digit estimate may have moved substantially. The older form was never claiming to be a forecast.
The Step 2 CK side of the site shows each form's band alongside its conversion.
Three causes of movement, and how to tell them apart
The form. Anything explained by the table above. Check this first, because it costs nothing to check and it dissolves a surprising number of panics.
Real change. Knowledge, timing, and test-taking habits genuinely shift over a dedicated period. Real change shows up as a consistent direction across three or more sittings, not as one dramatic jump.
Noise. Question sampling, sleep, whether you were sick, whether two guesses happened to land. Noise is symmetric: it moves scores up as often as down, and it does not repeat.
A useful discipline is to write down which of the three you think is responsible before you look at the next form's result. Predictions made after the fact are worthless.
How large a difference is worth reacting to?
Compare the gap to the band the model already admits to. This site attaches ±9 to a Form 9 estimate and ±5 to a Form 14 estimate. Two results that differ by less than the wider of the two bands are not evidence that anything changed.
That is a modelling convention, not a measured error rate — nobody has published a verifiable error distribution for community conversions of these self-assessments. It is still a better decision rule than reacting to every point. The reasoning behind those bands is set out in the modelling notes for every calculator here.
What a real trend looks like
Movement you can act on has three properties: it holds direction across at least three sittings, each step is larger than the model's own band, and it survives being converted on a single curve.
A clean example, all four results converted on the same Step 1 form:
| Sitting | Wrong of 200 | Estimate | Pass probability |
|---|---|---|---|
| Week 1 | 72 | 197 | 54% |
| Week 3 | 66 | 204 | 81% |
| Week 5 | 60 | 210 | 93% |
| Week 7 | 54 | 217 | 98% |
Six or seven points per fortnight, in the same direction, every time. That is a trend. Notice that it took a 20-point climb in the estimate to move the pass reading from 54% to 98%. The probability column is a summary of the score column, not independent evidence.
Compare that with 204, 199, 206, 201 over four sittings. Nothing there exceeds the band by a convincing margin, and no amount of staring will turn it into a direction. That pattern means "keep working and collect another data point", not "change everything".
What to do when three scores disagree
Start with the most recent one, because it is the only sitting that reflects everything you have studied. Then ask whether the older results argue for a lower or higher reading, and by how much.
Weighting several assessments by hand is error-prone, which is why the multi-assessment forecast tool exists: it takes each result plus its date and returns one number instead of three.
If your last two sittings were on forms with very different costs per wrong answer, convert both raw scores on a single form before you compare them. It is the only way to see whether you moved or the ruler did.
Quick answers
Is a 10-point swing between practice forms normal? It is common, and part of it is structural: on this site's models a 10-point gap can be produced by form choice plus a handful of missed questions, with no change in knowledge required.
Which score should I trust — the highest or the most recent? The most recent, adjusted for which form it came from. The highest score is the one most likely to contain good luck.
Does averaging my forms help? A weighted average is usually more stable than any single sitting, provided recent results count for more. A flat average of a diagnostic taken in week one and a form taken the day before the exam is misleading.
My score dropped on a newer form. Should I move my exam date? Not on that evidence alone. Check the form's cost per wrong answer first, then look at whether the drop exceeds the model's own band.
What this cannot tell you
These conversion lines are built on community-reported data, not NBME's equating. NBME does not publish a wrong-answers-to-three-digit formula for its self-assessments, so every offline conversion, this one included, is an estimate with real uncertainty around it.
None of the models can see how you took the form: untimed, split across two days, or with the reference sheet open. Each of those breaks the comparison with a form you sat under exam conditions, and none of them show up in the number.