← Blogs

The offline–online gap is well measured. Its causes aren't.

2026.09.14  ·  evaluation · recommenders  ·  notes

When a model looks better offline and the A/B test says otherwise, you'll often hear: "it's usually not the model." I've said it myself. It sounds right, and it's a comfortable thing to say because it points at the data and the logging instead of at the modelling. But "usually" is a claim about how often something happens. You'd need a base rate to back it. I had one experiment on this and wanted to know if that base rate exists anywhere. So I went and looked. Here's what the experiment showed, what the papers have actually measured, and the one thing I couldn't find anyone measuring.

What my one experiment showed

The experiment is RankShift. It uses KuaiRand, a Kuaishou dataset with something you almost never get: for two weeks the app slipped uniformly random videos into real feeds. So there's a test set where what people saw didn't depend on the recommender. I trained a LightGBM ranker on the normal logs and tested it on both kinds of exposure over the same two weeks.

Three things came out. Ranking held up. Lift at the top 23.9% of traffic went from 3.43× under normal exposure to 3.36× under random — down 2%. Calibration didn't hold up at all. Predicted probability over actual rate went from 0.96× to 2.07×. When I fitted a calibrator on the random slice, it fixed that slice and broke the normal one by about the same factor the other way. And the standard off-policy estimators — IPS, self-normalized IPS, a direct reward model, doubly robust — all missed the random slice's true value by 3 to 3.9× when I gave them item-level propensities. The confidence intervals were tight, so that's bias, not noise. Flip the direction, use exact propensities, and you still only recover a third of the policy's value. The other two-thirds is which user gets which video, and item-level propensities can't see that. All of it held at 322 million interactions.

Now the limits, because they matter. One dataset. One platform. Two weeks. One model, one seed. The features include watch time, which you only have after someone was shown the video — in a real pre-impression ranker every number above would drop. And the shift I studied, normal to uniformly random, is a stress test, not what a typical product change looks like. What the experiment tells you is how big one mechanism can get. It tells you nothing about how often that mechanism is the one biting you.

What the papers have measured

The gap itself is measured well. Here are the numbers I found.

StudyWhat was comparedResult
Booking.com, KDD 2019Offline gain vs. RCT business gain across 23 model pairsPearson −0.10, Spearman −0.18
Criteo, WSDM 2018Off-policy estimate vs. outcome of 39 real A/B testsbest estimator: precision 0.56
Spotify, WSDM 2019Offline vs. online ranking of 12 systems, by estimatorKendall τ 0.33 (n.s.) to 0.64
swissinfo.ch, RecSys 2014Offline replay vs. three-week live testmost-popular best offline, worst online
Docear, TPDL 2015Offline metrics vs. clicks vs. user ratingsoffline–rating r 0.52–0.67

Booking's authors call offline metrics "only a health check." The Spotify result is the one that stuck with me: same twelve systems, and whether offline agrees with the A/B depends on how you set up the estimator. You can get no correlation or a significant one from the same logs. So part of the gap is how the offline number was computed. Not all of it is the world.

The causes people have demonstrated fall into four places. Each one has at least one real measurement behind it.

The log. Offline data comes from whatever policy produced it. Cañamares and Castells showed that plain popularity beats collaborative filtering on MovieLens and Netflix under the usual held-out evaluation, and the winner flips once you collect test data missing-at-random. Schnabel et al. proved the bias of importance-weighted estimators scales with how wrong your propensities are. Dudík, Langford and Li proved doubly robust's bias is the product of reward-model error and propensity error. Get either one right and you're fine; get both wrong and you inherit both. And here's the part I hadn't appreciated: the two best real-log benchmarks for these estimators, Open Bandit and Bottou's Bing work, both used propensities logged at serving time. Their good results don't carry over to the situation most teams are in, where you estimate propensities afterwards. That's the situation my experiment was in. The estimators failed there, and the theorems say why.

The probabilities. Ranking and calibration come apart. Temperature scaling doesn't move the argmax (Guo et al.), and Facebook's ads team wrote back in 2014 that a single global multiplier fixes a 2× over-prediction without touching AUC. Ovadia et al. tested calibration under distribution shift on Criteo ad clicks and found that a calibrator fitted on the validation set made things worse than doing nothing. Google's ads team said in 2013 that the feedback loop between model and traffic makes calibration guarantees impossible in principle. A 2023 ICLR paper gives the mechanism: any policy that serves its top-scored items is selecting on the lucky side of its own errors, so served traffic looks over-predicted even when the model has no bias. My 0.96× → 2.07× is one instance of that. I looked for a published production A/B where AUC held and calibration broke, with numbers. I didn't find one. The closest are a worked example and a synthetic shift.

The plumbing. Google's ML Test Score calls the training/serving skew check "the most important and least implemented." When the Play store's data validation caught features that were always there in training and always missing in serving logs, fixing that was worth 2% on install rate in an A/B. The model was fine. Its inputs weren't. Chapelle measured that on Criteo half the conversions show up more than 24 hours after the click, so a model that treats unconverted clicks as negatives underpredicts by 21% while its offline loss looks normal. Ji et al. found 54% of recent RecSys papers use random splits. Those leak the future into training and swing HR@20 anywhere from −35% to +89% depending on the model — and they reorder the models. Krichene and Rendle showed sampled metrics can flip which algorithm wins, even in expectation.

The online number. Kohavi's team at Bing found that 67% of day-one treatment effects land outside the final confidence interval, and that most suspected novelty effects are just noise. Sometimes offline was right and the A/B hadn't finished yet.

What I couldn't find

All of those studies measure the gap. None of them measure the causes. Booking lists four mechanisms for its 23 cases and doesn't say how many fall under each. Criteo blames estimator bias and variance without splitting them. Garcin, Beel and Rossetti report reversals and call their explanations speculation. I went looking for one sentence of the form "X percent of offline–online gaps come from Y". I checked those papers, the Chen et al. bias survey, and the industry write-ups I could reach. Nothing. That's a bounded search, not proof it doesn't exist. But I looked.

I don't think it's laziness. The causes stack. A random split leaks the future, the leaked data is exposure-biased, and the serving features don't match training — all on the same model at once. When the A/B disappoints, you find the cause you went looking for and stop. To actually apportion a gap you'd need to rule out the others, and that takes the exploration slice, the logged propensities and the skew monitoring that Google's own rubric says almost nobody implements. Not having those is a big part of why the gap exists. And post-mortems mostly don't get published. So "it's usually X" is, on everything I can find, a guess. Maybe a good one. Still a guess, whoever's saying it.

So what do you do

If you can't know the cause in general, you can still order the checks by how much they cost. That's the version of the advice I'd stand behind now. Run an A/A and let the test run long enough, because the online number might be noise. Join the features you logged at serving against the ones you trained on, because skew is cheap to catch and Google keeps finding it. Look at whether it's the ranking metrics or the probabilities that disagree, because those break for different reasons and only one of them is fixed by recalibrating. Check the split and the metric: random or by time, sampled or exact. Then, last, the exposure question, because it's the expensive one.

The one structural change is to make the next gap explainable. Log the propensity when you serve — it's one field in the request log — and keep a small slice of randomized exposure running all the time. With those two, every estimator in the table above becomes usable and every future gap can be attributed. Without them you're stuck where the field is now: a list of plausible mechanisms and a shrug.

What I'd run next

Three things this experiment left hanging. Rerun the estimator comparison with user-conditional propensities, which KuaiRand's per-user random slice makes possible, and see if the missing two-thirds closes. Drop watch time and redo the transfer test as a real pre-impression problem — I expect the ranking result to get weaker and want to know by how much. And take the same decomposition to Open Bandit, where propensities are logged, so the answer can be checked. None of that turns one experiment into a base rate. It turns one measurement into three.

Sources Bernardi, Mavridis, Estevez — 150 Successful ML Models (KDD 2019) · Gilotte et al. — Offline A/B Testing for Recommender Systems (WSDM 2018, arXiv:1801.07030) · Gruson et al. — Offline Evaluation to Make Decisions about Playlist Recommendation (WSDM 2019) · Garcin et al. — Offline and Online Evaluation of News Recommenders at swissinfo.ch (RecSys 2014) · Beel & Langer — Offline Evaluations, Online Evaluations and User Studies (TPDL 2015) · Cañamares & Castells — Should I Follow the Crowd? (SIGIR 2018) · Schnabel et al. — Recommendations as Treatments (ICML 2016, arXiv:1602.05352) · Dudík, Langford, Li — Doubly Robust Policy Evaluation (ICML 2011, arXiv:1103.4601) · Saito et al. — Open Bandit Dataset and Pipeline (NeurIPS 2021, arXiv:2008.07146) · Bottou et al. — Counterfactual Reasoning and Learning Systems (JMLR 2013) · Guo et al. — On Calibration of Modern Neural Networks (ICML 2017, arXiv:1706.04599) · He et al. — Practical Lessons from Predicting Clicks on Ads at Facebook (ADKDD 2014) · McMahan et al. — Ad Click Prediction: a View from the Trenches (KDD 2013) · Ovadia et al. — Can You Trust Your Model's Uncertainty? (NeurIPS 2019, arXiv:1906.02530) · Fan, Si, Zhang — Calibration Matters: Tackling Maximization Bias (ICLR 2023, arXiv:2205.09809) · Breck et al. — The ML Test Score (IEEE Big Data 2017) · Breck, Polyzotis et al. — Data Validation for Machine Learning (MLSys 2019) · Chapelle — Modeling Delayed Feedback in Display Advertising (KDD 2014) · Ji et al. — A Critical Study on Data Leakage in Recommender System Offline Evaluation (TOIS 2023, arXiv:2010.11060) · Krichene & Rendle — On Sampled Metrics for Item Recommendation (KDD 2020, arXiv:1912.02263) · Kohavi et al. — Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained (KDD 2012) · Gao et al. — KuaiRand (CIKM 2022, arXiv:2208.08696). The experiment: PSCRedefine/rankshift, every figure from a committed results/*.json.