We show that LLM-based e-commerce rankers (Llama, Mistral, Qwen) exhibit systematic geographic bias favoring brands and continents tied to their developers’ home region, with user-location context having the strongest effect on rankings.
We propose iterative reranking, a compute-scaling strategy that repeatedly applies listwise LLM rankers to exploit their non-determinism, yielding consistent nDCG gains across open datasets and LLMs.
We propose NextLevelBERT, a Masked Language Model that operates on text-embedding representations of whole chunks instead of tokens, and show it handles long-document tasks effectively while outperforming much larger embedding models.
We introduce SCAN, a novel analogy dataset spanning dissimilar domains, and show that state-of-the-art pretrained language models struggle with complex analogical reasoning.