Skip to main content
All posts
Comparison

Your UWorld Percentage Is Not a Three-Digit Score

Why no honest calculator converts a question-bank average into a USMLE score, demonstrated with this site's own percent-based models, where the same 70% correct produces estimates 38 points apart.

There is no dependable conversion from a question-bank average to a three-digit score, and this site does not publish one. A percentage only means something when it is attached to a specific, fixed, whole assessment. Change the instrument and the same percentage moves by tens of points. On the models here, 70% correct is an estimated 198 on one instrument and 236 on another.

The demonstration, using this site's own models

Three of the calculators here take a percentage as their input. Feeding identical percentages into all three shows how little the percentage carries on its own.

70% correct on…Estimated score
Free 120, Step 1198
Free 120, Step 2 CK223
UWSA 1, Step 2 CK236

A 38-point spread, from one number. The exam differs, the item difficulty differs, and the scale each instrument reports on differs. None of that is visible in "70%".

Now add a fourth instrument that no model here covers, a question bank you work through over months, in blocks you choose, with a subject filter you control, and the spread gets wider, not narrower.

A percentage does not even convert linearly within one instrument

The UWSA 1 Step 2 CK model is piecewise: it interpolates between anchor points rather than following a straight line. The exchange rate changes as you move up the scale.

Percent correctEstimated scoreGain from previous row
55%210
60%215+5
65%225+10
70%236+11
75%246+10

Five percentage points buy you five score points at the bottom of that table and eleven in the middle. Any rule of thumb of the form "each 1% is worth two points" is wrong somewhere on the curve, usually where it matters most.

Why a question-bank average runs on a different footing

Exposure. A self-assessment is a fixed set of unseen items. A question bank on a second pass contains items you have already reasoned through, and recall inflates the percentage without inflating your knowledge.

Mode. Tutor mode with immediate feedback is a different task from a timed block. So is a 20-question set versus a 40-question block at exam pace.

Selection. Subject-filtered blocks measure one system on the day you studied it. Random timed blocks measure retrieval across everything. The two numbers are not comparable to each other, let alone to a scaled score.

Population. A percentage correct is a raw statistic. A three-digit USMLE score is a scaled, equated number whose meaning is set against a reference population. Turning one into the other requires the equating that NBME does not publish for these self-assessments.

What your qbank average is actually good for

It is a workload and consistency signal, and a good one. Tracked properly, it tells you whether random timed blocks are trending upward, whether a system you finished six weeks ago has decayed, and whether your accuracy collapses in the last ten questions of a block.

Track it the way you would track training volume rather than a race result:

  • First pass and second pass are separate series. Never merge them.
  • Random timed blocks only, if you want the number to mean anything week to week.
  • Watch the trend across at least ten blocks. Single-block percentages swing wildly on 40 items.
  • Log accuracy by position within the block. A drop in the final third is a pacing problem, not a knowledge problem.

None of that produces a score prediction, and it does not need to.

One more piece of arithmetic keeps block-to-block percentages in perspective. In a 40-question block, a single item is worth 2.5 percentage points. The difference between a 68% block and a 73% block is two questions: one misread stem and one lucky guess.

That is why a single block tells you almost nothing, and why a ten-block rolling average tells you something real. Judge the series, never the point.

What to use instead when you want a number

Use a whole assessment that was designed to be scored, then convert it. That is the entire premise of the calculators here: a fixed item set, sat in one sitting, mapped by a model fitted for that specific form.

For Step 1, start from a full self-assessment result and read the estimate on the Step 1 conversion tools. For Step 2 CK, the CCSSA form converter does the same across Forms 9 to 15. If you want practice volume without a subscription, the site's free question set is open with no account.

The practical rhythm most students land on: question banks for learning and consistency, self-assessments for measurement, and no attempt to make the first do the job of the second.

Two students, one 68%

Student A works 40-item random timed blocks, first pass, no subject filter, and averages 68%. Student B works 20-item tutor-mode sets filtered to the system studied that week, on a second pass, and averages 68%.

Nothing about those two numbers is comparable. They differ in exposure, timing, breadth, and feedback. Yet both students will type the same query into a search box and both will find sites willing to hand them the same three-digit answer.

The gap matters because percentage differences are amplified by any real conversion. On the UWSA 1 Step 2 CK model, 62% correct estimates 219 and 68% estimates 232: six percentage points, thirteen score points. If your percentage is inflated by a few points of recall or filtering, a mapping like that turns a small measurement artefact into a large fantasy.

What a defensible conversion actually requires

Four conditions have to hold before a percentage can be mapped to a three-digit estimate with a straight face:

  1. A fixed item set. Everyone taking the assessment sees the same questions, so difficulty is constant across users.
  2. A single sitting under known conditions. Timed, full length, no references, because otherwise the raw score describes a different task.
  3. A reported scale to fit against. The assessment must produce its own score that can be regressed onto exam outcomes.
  4. Enough reported pairs of assessment result and exam result for the fit to mean anything.

The self-assessments here satisfy the first three. The fourth is where every offline conversion, this site's included, carries genuine uncertainty, which is why the estimates come with bands rather than single confident numbers.

A question bank fails the first condition outright, and usually the second as well. There is no fix for that; it is what a question bank is.

Quick answers

What score is 65% in UWorld? There is no answerable version of this question. The percentage is not tied to a fixed item set or a scaled reporting population, so no defensible mapping exists.

But other sites publish a UWorld-percentage-to-score chart. They do, and none of them can show you the dataset behind it. Treat any such chart as a guess presented with more confidence than it has earned.

Does a rising qbank average mean my score is rising? Usually the direction is meaningful, especially across first-pass random timed blocks. The magnitude is not: you cannot read points off a percentage change.

Is my qbank percentage useless before the exam? No. It is the best available signal for consistency and pacing. It just is not a score, and treating it as one leads to both false comfort and unnecessary panic.

Why do the calculators here accept a percentage for some assessments? Because for the Free 120 and UWSA 1 Step 2 CK, the percentage refers to a fixed, known item set. That is what makes a model possible. A question bank has no fixed item set.

The honest limits

Every conversion on this site is a line built on community-reported data, not official equating, and each one is specific to the assessment it was fitted for. Applying a Free 120 model to a qbank percentage, or an NBME form model to a set of practice blocks, produces a number with no support behind it.

Where a mapping cannot be built honestly, the right answer is to decline, which is why you will not find a question-bank converter here. How the models that do exist were constructed, and what they assume, is set out in the method behind these estimates.