EMNLP 2026 · Main Conference · Oral
Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition
I study how neural network architecture shapes the way speech models understand audio. Analyzing 24 pretrained speech models, we found that Conformers categorize acoustic information earlier, while Transformers integrate context later. We introduce Architectural Fingerprinting to systematically uncover these differences.
