files: 151 total, 120 train, 31 test pairs: 2041 train, 1187 test token-space vocabulary: 150 words, e.g. ['fastapi', 'use', 'py', 'docs_src', 'hl', 'path', 'python', 'using', 'https', 'example'] task A zero-shot baseline (context TF-IDF used directly as query, no predictor trained): hit@1=0.005 hit@5=0.127 MRR=0.068 (n=1187, random-guess hit@1 would be 0.0008) task A: embedding-space prediction (Ridge, context TF-IDF -> target TF-IDF) frac n_train hit@1 hit@5 MRR train_s 0.10 204 0.011 0.026 0.025 0.05 0.25 510 0.012 0.051 0.039 0.19 0.50 1020 0.018 0.061 0.049 0.29 1.00 2041 0.020 0.083 0.062 0.47 task B: token-space prediction (LogisticRegression x150 words, context TF-IDF -> word presence) frac n_train micro_F1 train_s 0.10 204 0.001 2.19 0.25 510 0.105 0.64 0.50 1020 0.120 0.97 1.00 2041 0.230 1.64 task B majority-class baseline (predict train-set >50%-frequent words every time): micro_F1=0.000 sample efficiency, 100%->10% of train pairs: task A hit@1: 0.020 -> 0.011 (+45.8% relative change) task B micro_F1: 0.230 -> 0.001 (+99.4% relative change) task B's relative drop is 2.17x steeper than task A's writeup ratios (full precision): data scale 10%->100%: 10.00x task A hit@5 100%/10%: 3.16x task A hit@1 100% vs random baseline: 24.00x zero-shot vs trained(100%) hit@1: trained is 4.00x zero-shot zero-shot vs trained(100%) hit@5: zero-shot is 1.54x trained zero-shot vs trained(100%) MRR: zero-shot is +10.2% relative to trained