Research by Song-Ze (Jimmy) Yu

Public work and ongoing research, newest first.

2026 · Preprint BAIR
Click to enlarge

PitchBench: Measuring Pitch Hearing in Audio-Language Models

Milan Liessens Dujardin*, Song-Ze Yu*, Craver Corbyn Thomas-Smith, David M. Chan, Karina Nguyen * Equal contribution

ALMs are moving quickly toward high-level music understanding and captioning, yet lacks a systematic, hierarchical benchmark for the perceptual abilities those applications assume. PitchBench addresses this gap with 28 experiments that disentangle what a model perceives from when it perceives it, progressing from I. atomic pitch to II. contextual, time-grounded perception and III. melodic pitch perceptions. Together, these levels test the absolute and relative pitch hearing needed to reason about high-level intervals, melodies, and harmony.

2026 · DAFx CNMAT
Click to enlarge

InstructFX2FX: A Multi-Turn Text-to-Effect System for Sequential Audio Effect Refinement

Song-Ze Yu, Milan Liessens Dujardin, Yuxuan Cai, Wantong Zhang, Brian Cruz, Jeremy Wagner, Carmine-Emanuele Cella

Introduces sequential FX refinement: given the current effect state and a new instruction, update the sound while preserving what earlier instructions have already achieved. InstructFX2FX combines LLM planning with CLAP-guided optimization for iterative, multi-turn audio effect design.

2026 · Under review CNMAT
Click to enlarge

MuNo-LM: Unifying Score and Performance for Audio-Language Modeling

Milan Liessens Dujardin, Song-Ze Yu

Fine-grained music understanding requires models to represent both what is played and how it is performed, grounded in time. MuNo-LM (Music Notation for Language Models) replaces coarse, weakly time-grounded synthesized captions with an LLM-readable notation that unifies score structure and time-aligned perceptual features, then generates musically informed, time-grounded captions and question-answer pairs.

2026 · Under review BAIR
Click to enlarge

Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation

Song-Ze Yu, Joseph Suh, Serina Chang, David M. Chan

Anamnesis is an interactive system for demographically controllable survey simulation using large language models. Researchers can sample virtual respondents from a validated pool of 30K+ richly detailed narrative backstories pool to conduct multiple choice questions, open-ended or multimodal surveys, yielding opinion distributions that more closely match real-world data than standard persona prompting or LLM-as-a-judge baseline.

Research figure