PitchBench: Measuring Pitch Hearing in Audio-Language Models
ALMs are moving quickly toward high-level music understanding and captioning, yet lacks a systematic, hierarchical benchmark for the perceptual abilities those applications assume. PitchBench addresses this gap with 28 experiments that disentangle what a model perceives from when it perceives it, progressing from I. atomic pitch to II. contextual, time-grounded perception and III. melodic pitch perceptions. Together, these levels test the absolute and relative pitch hearing needed to reason about high-level intervals, melodies, and harmony.