A pass probability on a practice-score calculator is arithmetic applied to your estimated score, not an observed outcome frequency. Nobody tracked a cohort of students with your exact result and counted how many passed. On this site the number comes from a curve fitted around the passing standard, and because different assessments carry different curves, the same estimated 220 reads 60% on one Step 2 CK form and 88% on another.
Where the percentage comes from
Every calculator here converts your raw result to a three-digit estimate first. The pass probability is then computed from that estimate alone, using a curve anchored at the passing standard: 196 for Step 1, 218 for Step 2 CK.
Nothing else feeds in. Not your study history, not how many forms you have taken, not how long until your exam date. Two students with identical estimates on the same form always see the same percentage.
Here is the Step 1 curve as it behaves on NBME 25:
| Wrong of 200 | Estimated score | Pass probability |
|---|---|---|
| 62 | 208 | 90% |
| 66 | 204 | 81% |
| 68 | 201 | 71% |
| 71 | 198 | 59% |
| 73 | 196 | 50% |
| 76 | 192 | 33% |
Read the middle row carefully. At the passing standard the curve returns exactly 50%, by construction, on every Step 1 calculator on this site. That is the definition of the anchor, not a finding.
The 50% point is a design decision
The same thing happens on the Step 1 Free 120 model, which takes a percentage rather than a wrong-answer count: 68% correct converts to 196, and 196 returns 50%. On the Step 2 CK forms the anchor sits at 218.
So a 50% reading does not mean the model has weighed your case and found it a coin flip. It means your estimate landed on the line. What the curve encodes is how quickly confidence should grow as you move away from that line, and that rate is a modelling choice.
Two forms, one score, two very different percentages
This is the part no calculator advertises about itself. The Step 2 CK assessments on this site do not share a single probability model. Forms 9 and 11 use a smooth logistic curve; Forms 10, 12, 13, 14 and 15 use a step table that holds a value across a band of scores.
| Wrong of 200 | Estimate | NBME 11 (logistic) | NBME 12 (step table) |
|---|---|---|---|
| 70 | 224 | 77% | 88% |
| 72 | 222 | 69% | 88% |
| 74 | 220 | 60% | 88% |
| 76 | 217 | 45% | 60% |
| 80 | 213 | 27% | 60% |
At 220 the two models disagree by 28 percentage points about identical performance. Neither is lying; they are different summaries of the same uncertain territory. If you have been comparing the percentage across forms and drawing conclusions from the movement, you have been reading noise introduced by the models rather than by your performance. Convert the raw counts on one common Step 2 CK curve instead.
Some curves are deliberately reluctant
The Step 2 CK Free 120 model here is intentionally flat. Its output is capped at 97% and floored at 3%, and it moves slowly:
- 60% correct → estimated 206 → 28%
- 70% correct → estimated 223 → 60%
- 80% correct → estimated 240 → 85%
- 90% correct → estimated 257 → 96%
A 51-point swing in the estimate moves the probability from 28% to 96%. On a steeper curve, an estimate of 240 would have saturated near 99% long before. The flatness is the model admitting that a 120-question sample is thin evidence, whichever direction it points.
The curve is steepest exactly where the decision is hardest
The Step 1 curve is not evenly spaced. Here is what each five points of estimated score buys you in probability:
| Estimated score | Pass probability | Gained over previous row |
|---|---|---|
| 196 | 50% | — |
| 201 | 71% | +21 |
| 206 | 86% | +15 |
| 211 | 94% | +8 |
| 216 | 97% | +3 |
| 221 | 99% | +2 |
Two consequences follow. Near the line, the percentage is hypersensitive: a three-point wobble, well inside the noise of any single sitting, swings the reading more than a dozen points. Far above the line, it is saturated and useless for tracking progress; you can gain fifteen real points of ability and watch the percentage move by two.
So the percentage is at its most volatile precisely where students look at it hardest, and at its least informative where they have stopped worrying. Track the estimate itself if you want to see improvement.
Shifted curves: when 218 does not read 50%
The UWSA 2 Step 2 CK model is the exception on this site. Its probability curve is centred at 223 rather than 218, because the underlying model treats UWSA 2 as running optimistic and corrects it before asking about passing.
The consequence is worth stating plainly: an estimated 218 on UWSA 2 returns 32%, not 50%. An estimated 225 returns 57%. If you have used UWSA 2 as your final assessment and expected the pass reading to line up with the Step 2 CK standard, that gap is the correction, not a bug. The reasoning behind each form's adjustment is documented in the assumptions behind these estimates.
How much weight the number deserves
Use it as a coarse position marker: comfortably clear, near the line, or well below. Those three states are about as much resolution as any offline model can honestly support.
Do not use it as odds. "84%" does not mean 84 of 100 students like you passed, because there is no such measured cohort behind it — not here, and not on the competitor sites that publish tidier-looking figures. Any site quoting a precise accuracy rate for these conversions is quoting something that has not been independently verified.
And do not track it across forms as if it were a single instrument. Track the three-digit estimates, and let a weighted forecast across your sittings handle the combination.
Why you will not find an accuracy figure here
Competing calculators advertise numbers like "accurate within five points" or a correlation coefficient carried to two decimals. None of those figures come with a dataset you can inspect, a methodology you can replicate, or a source you can check.
This site takes the opposite position, and it costs us something in persuasiveness. What can be stated honestly about a pass probability here is:
- the exact input it is computed from (your estimated score, and nothing else)
- the exact shape of the model (a logistic curve or a step table, per form)
- where the anchor sits (196 for Step 1, 218 for Step 2 CK, 223 for the corrected UWSA 2 Step 2 CK curve)
- that no verified outcome dataset sits behind any of it
Everything else, including how often a 78% reading corresponds to a pass in reality, is unknown to us, and would be unknown to any site that has not run a longitudinal study it can publish. Treat a precise-sounding accuracy claim as a marketing decision rather than a measurement.
Quick answers
Does 90% mean I will pass? No. It means your estimated score sits far enough above the passing standard that this model's curve returns 90. The exam itself is the only thing that determines the outcome.
Why did my pass probability fall when my score barely moved? Near the passing standard the curves are at their steepest, so a two- or three-point drop in the estimate can move the percentage a long way. Far from the line the same drop barely registers.
Why is my probability different on two forms with the same score? Because the forms use different probability models: a smooth curve on some, a step table on others, and a shifted curve on UWSA 2 Step 2 CK.
Is the passing standard really 196 and 218? Those are the standards this site uses: 196 for Step 1, 218 for Step 2 CK. The official figures are set by the USMLE and have been revised before, so confirm the current values on usmle.org before you plan around them.
The limits worth repeating
Step 1 is reported as pass/fail, which makes the probability reading the whole story for Step 1 students, and that is exactly why it should not be over-read. A single self-assessment, converted offline, cannot account for exam-day conditions, a different item mix, or the difference between a four-block practice form and a full test day.
If you want the underlying conversion rather than the probability, start from a raw Step 1 result and read the three-digit estimate. The score is the measurement; the percentage is a summary of it.