Good question! I tested a bunch of model sizes and quantizations against a specific benchmark: correctly identifying Canada geese instead of just labeling them "ducks." Smaller models (~4.5B params) failed that test consistently — even at full precision, no quantization — so it wasn't a quantization issue, it was that they simply don't have enough capacity for fine-grained visual discrimination. Gemma 4 12B at 5-bit was the first one that nailed it reliably, so that became my baseline for TYPh.