DeepASMR-NSpeech: A Fine-Grained Benchmark for Non-Speech ASMR Generation

73,829 clips · 202.6 hours · structured SVO annotations · SVO-AQA

Ziyi Yang · Leying Zhang · Chenda Li · Yanmin Qian

Auditory Cognition and Computational Acoustics Lab
Shanghai Jiao Tong University

Abstract

We introduce DeepASMR-NSpeech, a benchmark designed to support text-to-audio generation research in the largely underexplored domain of non-speech ASMR. DeepASMR-NSpeech contains 73,829 ten-second clips totaling 202.6 hours and represents each sound-producing interaction using a structured Subject–Verb–Object schema. Its closed verb vocabulary comprises 37 fine-grained actions grouped into 18 superclasses, while subjects and objects are described using open-set material–entity expressions. We further introduce SVO-AQA, a human-validated multiple-choice benchmark that separately evaluates interaction verbs, subject materials, and object materials, providing finer-grained diagnostics than conventional distributional and caption-level metrics. Listening tests are also used to assess overall generation quality and ASMR comfort. Zero-shot evaluation of six representative text-to-audio models reveals a pronounced domain gap between AudioCaps and DeepASMR-NSpeech, showing that performance on general environmental audio does not reliably transfer to non-speech ASMR. Generated clips preserve interaction verbs more consistently than subject and object materials, indicating that material-dependent acoustic modeling remains a major limitation. Human evaluations further reveal substantial gaps from reference recordings in both semantic alignment and listening comfort. These results show that current TTA systems are inadequate for fine-grained ASMR generation and establish DeepASMR-NSpeech as an important benchmark for advancing interaction-aware and material-faithful audio generation.

Taxonomy

DeepASMR-NSpeech uses mechanism-driven verb superclasses (left) and twelve material superclasses (right). Verbs branch from each superclass box tail. Clay denotes wet sticky modeling pack; foam solid is expanded polymer foam (e.g., floral foam crush), distinct from foam lather (wet shaving foam).

Verb superclass tree and material superclass tree
Figure 1: Closed verb taxonomy (37 verbs → 18 superclasses) and material taxonomy (12 superclasses).

Overview

Eight test-set clips with structured captions and reference recordings. We recommend headphones for low-intensity ASMR triggers.

Evaluation

SVO-AQA Multiple-Choice Examples

Six SVO-AQA items (2 per task). Reference audio answers correctly; each row aligns one model’s audio with its multiple-choice prediction. Four finetuned backbones use the matched test-time protocol from the supplementary material. Object-material cases highlight foam solid confusion; subject-material and verb cases show additional failure modes on generated audio.

LLM Evaluator Limitations

Even on human-validated reference audio, the Omni-based AQA scorer can err (e.g., stirring vs. crunching, wax vs. foam solid). These cases illustrate that SVO-AQA should complement—not replace—human listening tests.