files per category: {'advanced': 34, 'deployment': 9, 'how-to': 12, 'reference': 24, 'tutorial': 51} train files: 96, test files: 34 train snippets: 1550, test snippets: 637 train snippets per category: {'advanced': 492, 'deployment': 96, 'how-to': 114, 'reference': 54, 'tutorial': 794} test snippets per category: {'advanced': 109, 'deployment': 177, 'how-to': 28, 'reference': 6, 'tutorial': 317} accuracy=0.600 macro_f1=0.427 (baseline if always predicting majority class 'tutorial': 0.498) confusion matrix (rows=true, cols=predicted), classes: [np.str_('advanced'), np.str_('deployment'), np.str_('how-to'), np.str_('reference'), np.str_('tutorial')] advanc deploy how-to refere tutori advance 40 1 2 1 65 deploym 32 86 0 1 58 how-to 12 0 2 1 13 referen 2 0 0 3 1 tutoria 50 4 7 5 251 calibration (predicted max-probability bucket vs observed accuracy): confidence [0.2,0.4): n= 57 mean_confidence=0.363 observed_accuracy=0.368 confidence [0.4,0.6): n= 249 mean_confidence=0.505 observed_accuracy=0.510 confidence [0.6,0.8): n= 214 mean_confidence=0.698 observed_accuracy=0.673 confidence [0.8,1.0): n= 117 mean_confidence=0.883 observed_accuracy=0.769 median single-snippet latency (tfidf transform + logreg predict_proba): 0.628ms over 300 calls, min=0.575ms max=0.970ms for reference only (not measured, arithmetic on published price): routing these 637 test snippets through Jev at its published input rate would cost about $0.0014 for ~33,595 approx tokens (word-count-based estimate, not a real tokenizer). This local classifier: $0, no network call.