DeepASMR-NSpeech: A Fine-Grained Benchmark for Non-Speech ASMR Generation
73,829 clips · 202.6 hours · structured SVO annotations · SVO-AQA
Auditory Cognition and Computational Acoustics Lab
Shanghai Jiao Tong University
Abstract
We introduce DeepASMR-NSpeech, a benchmark designed to support text-to-audio generation research in the largely underexplored domain of non-speech ASMR. DeepASMR-NSpeech contains 73,829 ten-second clips totaling 202.6 hours and represents each sound-producing interaction using a structured Subject–Verb–Object schema. Its closed verb vocabulary comprises 37 fine-grained actions grouped into 18 superclasses, while subjects and objects are described using open-set material–entity expressions. We further introduce SVO-AQA, a human-validated multiple-choice benchmark that separately evaluates interaction verbs, subject materials, and object materials, providing finer-grained diagnostics than conventional distributional and caption-level metrics. Listening tests are also used to assess overall generation quality and ASMR comfort. Zero-shot evaluation of six representative text-to-audio models reveals a pronounced domain gap between AudioCaps and DeepASMR-NSpeech, showing that performance on general environmental audio does not reliably transfer to non-speech ASMR. Generated clips preserve interaction verbs more consistently than subject and object materials, indicating that material-dependent acoustic modeling remains a major limitation. Human evaluations further reveal substantial gaps from reference recordings in both semantic alignment and listening comfort. These results show that current TTA systems are inadequate for fine-grained ASMR generation and establish DeepASMR-NSpeech as an important benchmark for advancing interaction-aware and material-faithful audio generation.
Taxonomy
DeepASMR-NSpeech uses mechanism-driven verb superclasses (left) and twelve material superclasses (right). Verbs branch from each superclass box tail. Clay denotes wet sticky modeling pack; foam solid is expanded polymer foam (e.g., floral foam crush), distinct from foam lather (wet shaving foam).