Got a bug report that PCA queries returned different top-k than brute-force cosine. Same collection, same query vector, different rankings. Both were supposed to be cosine similarity. One of them had to be wrong.
Traced one specific vector that ranked 62 in brute-force but 33,220 in PCA. Brute-force cosine: 0.192. PCA-space cosine: -0.077. Negative. The angle between query and target wasn't preserved by the projection. The 32D reduction we'd defaulted to was capturing about 25% of GloVe's 300D variance and the rest of the signal was getting clipped. Vectors that were close in 300 dimensions weren't even pointing the same direction in 32.
Backed PCA out of the default query path until the experiment scheduler validates whether it fits a given collection. Works-on-the-benchmark and works-on-your-data are 2 different claims, and only one of them pays.
