We Fitted 322 Models of What Gets You an Interview. 145 of Them Found Nothing.
Every applicant is told the same list: publish, get AOA, honor your clerkships, collect research. We modeled each of those against real interview outcomes, one specialty and one signal stratum at a time, holding the program fixed. Nearly half the time the effect was not distinguishable from zero. In dermatology, the pattern is sharper than that: your credentials barely register until you have signaled the program.
TL;DR
- 322 credential models across 23 specialties. 145 (45%) show no measurable independent effect on getting an interview.
- Clerkship honor grades are the most consistent credential, registering in 76% of models. First-author publications register in 46%.
- In dermatology, signaling changes which credentials matter at all. Unsignaled, first-author pubs show an odds ratio of 1.03 (CI 0.73 to 1.36): nothing. Signaled, the same credential is 1.24 (CI 1.16 to 1.37) and worth about +3.8 percentage points.
- The effects that do exist are small. Every measurable derm credential is worth between +1 and +4 percentage points of invite probability. None of them is a strategy on its own.
- This is observational. These are partial associations within program, not causal effects, and we say so at every step.
Finding 1: Nearly half of credential effects are not measurable
For each of 23 specialties we fit two models, one for applicants who signaled a given program and one for applicants who did not, and estimated seven credentials in each. That is 322 estimates. In 145 of them the 95% interval crosses an odds ratio of 1.0, which means the data cannot distinguish the credential’s effect from zero.
A null is not proof of no effect. It means the effect, if any, is small enough that this dataset cannot see it after holding the program, Step 2 band, applicant type, away rotation, geographic tie and cycle year fixed. But a credential you cannot detect at this sample size is also not a credential you should reorganize a year of your life around.
Share of 46 fitted models (23 specialties x 2 signal strata) where the credential’s interval clears an odds ratio of 1.0.
View as table
| Credential | Measurable | Share |
|---|---|---|
| Clerkship honor grades | 35 of 46 | 76% |
| Honor society (AOA/Sigma) | 27 of 46 | 59% |
| Class quartile | 27 of 46 | 59% |
| Gold Humanism | 27 of 46 | 59% |
| Research experiences | 22 of 46 | 48% |
| First-author publications | 21 of 46 | 46% |
| Final honor, home rotation | 18 of 46 | 39% |
Clerkship grades are the most consistent signal in the dataset, registering in 35 of 46 models. First-author publications, the credential applicants most often describe as the thing they are grinding for, register in 21 of 46. The gap between how hard a credential is to obtain and how reliably it shows up in the data is the whole point of this post.
The signal-strategy page runs this model against your actual numbers, and moves your estimate only for the credentials with a measurable effect in your specialty and stratum.
See your own credential rail →Finding 2: In dermatology, credentials only start counting once you have signaled
Dermatology is the cleanest illustration because its two strata behave completely differently. Among applicants who did not signal a program, the invite rate is 7.6%, and almost nothing on the CV moves it: publications sit at an odds ratio of 1.03 with an interval from 0.73 to 1.36, which is a shrug. Among applicants who signaled, the invite rate is 30.6%, and three credentials separate from the null.
Odds ratio for an interview invite with a 95% cluster-bootstrap interval, within program. The line at 1.0 is no effect.
Applicants who did NOT signal the program
15,368 applicant x program rows · 372 applicants · 7.6% invite rate · 141 programs with a fixed effect
Applicants who DID signal the program
8,471 applicant x program rows · 479 applicants · 30.6% invite rate · 110 programs with a fixed effect
View as table
| Stratum | Credential | OR | 95% CI | Marginal |
|---|---|---|---|---|
| No signal | Class quartile | 1.27 | 1.09 to 1.47 | +1.43pp |
| No signal | Research experiences | 1.19 | 1.06 to 1.31 | +1.01pp |
| No signal | Final honor, home rotation | 1.24 | 0.99 to 1.57 | +1.29pp |
| No signal | Honor society (AOA/Sigma) | 1.21 | 0.97 to 1.55 | +1.15pp |
| No signal | Gold Humanism | 1.15 | 0.90 to 1.47 | +0.81pp |
| No signal | First-author publications | 1.03 | 0.73 to 1.36 | +0.19pp |
| No signal | Clerkship honor grades | 0.93 | 0.80 to 1.12 | -0.39pp |
| Signaled | Gold Humanism | 1.25 | 1.01 to 1.52 | +3.87pp |
| Signaled | First-author publications | 1.24 | 1.16 to 1.37 | +3.81pp |
| Signaled | Clerkship honor grades | 1.23 | 1.08 to 1.38 | +3.58pp |
| Signaled | Honor society (AOA/Sigma) | 1.15 | 0.93 to 1.42 | +2.38pp |
| Signaled | Class quartile | 1.10 | 0.94 to 1.33 | +1.64pp |
| Signaled | Final honor, home rotation | 1.02 | 0.77 to 1.38 | +0.41pp |
| Signaled | Research experiences | 1.00 | 0.93 to 1.07 | +0.04pp |
Read the top panel as the answer to “will a better CV get me in the door at a program I did not signal?” and the bottom as “now that I am being looked at, what helps?” Class quartile and research experience carry a small effect in the first case; publications, clerkship honors and Gold Humanism carry a larger one in the second. Note that clerkship grades actually point the wrong way in the unsignaled stratum (0.93) while being the strongest credential in the signaled one (1.23), which is the kind of reversal you get when a credential only becomes readable once someone is actually reading the file.
This lines up with what the signaling data says on its own: in dermatology, an unsignaled application is largely not competitive regardless of what is in it.
The companion piece: unsignaled derm applicants without ties interview at 8%, signaled at 39%, and an away rotation takes it to 86%.
Read the derm signal analysis →Finding 3: What the model holds fixed, and why it matters
Most credential advice comes from comparing applicants who matched against applicants who did not. That comparison is contaminated: the applicants with more publications also tend to have higher Step 2 scores, be at stronger schools, and apply to different programs.
These models are fitted at the applicant-by-program row level with a fixed effect for every program that has at least 30 rows in the stratum, so each credential coefficient is a comparison within the same program. Step 2 band, applicant type, away rotation, geographic tie and cycle year are all held fixed alongside. The credential block is collinear by construction, so a ridge penalty keeps each coefficient conservatively shrunk rather than letting correlated predictors trade places.
Intervals come from a cluster bootstrap that resamples applicants, not rows, because one applicant contributes many rows and those rows are not independent. That is why the intervals here are wider than a naive fit would produce, and why more of them cross the null.
Ridge parameters, the fixed-effect threshold, the bootstrap design, the calibration check and the deliberate exclusions are all documented.
Read the full model spec →Finding 4: What to do with a year
The honest summary is that no credential on this list is worth more than a few percentage points of invite probability, and the two instruments that move interview odds by tens of points are the ones that decide which programs read your file at all: signals and away rotations. In dermatology an away rotation is associated with an 86% interview rate at that program. No credential in this dataset comes close to that.
That does not mean stop doing research. It means: if you are choosing between a fourth first-author submission and an away rotation at a program you want, the data is not ambiguous about which one changes your odds more.
Away applications, response times and community outcomes per program.
Plan your aways →What this cannot tell you
These are associations, not causes. Applicants who signal a program differ from applicants who do not in ways no dataset records, and the signal split is itself contaminated because people signal programs they already fit.
A null is a limit of detection, not a proof of zero. Thin strata produce wide intervals, and a wide interval crossing 1.0 is reported as no measurable effect whether the true effect is zero or merely small.
Interviews are not matches. This models the invite, which is the stage where a CV is read. What happens on interview day is not in here.
First-author publication counts are missing for 2023 and 2024 in the source data, so that credential is fitted on 2025 and 2026 rows only, with no cross-year imputation.
A population-level starting point from published NRMP outcomes, before any of the program-level modeling above.
Estimate your match probability →Related reading
- How long away rotations take to decide: 3,347 applications across 12 specialties.
- Dermatology signal strategy 2026: where the 4.8x signal lift comes from.
- Residency match statistics 2026: the population baseline across 22 specialties.
Methodology & sources
Ridge-penalized logistic regression of the interview outcome at the applicant-by-program row level, fitted separately within the signaled and unsignaled strata for each of 23 specialties. Covariates: Step 2 CK band, applicant type, away rotation, geographic connection, cycle-year dummies, and a fixed effect per program with at least 30 rows in the stratum. Credentials are mean-centered within specialty, so a coefficient reads as the effect of being one unit above that specialty’s norm. Intervals are 2.5 and 97.5 percentiles from a cluster bootstrap resampling applicants. Effect sizes and intervals lead; p-values are not used. Dermatology figures are from a strong-reliability fit over 23,839 rows and 250 bootstrap replicates.
Data as of August 31, 2026. Charts use an emphasis palette (one hue plus a de-emphasis gray) validated for colorblind separation and surface contrast in both light and dark modes; identity is additionally carried by the legend, by whether the interval crosses 1.0, and by the printed odds ratio on every row. Every chart carries a table view.