Research

Publication

Paper 01 NeurIPS 2026

LAION-Tunes

Accepted at NeurIPS 2026

Abstract

Generative music platforms have reached commercial scale, with outputs approaching human-made quality. Yet music machine learning lags image-text research because no open music dataset matches the scale that catalyzed image-text foundation models: commercial recordings cannot be legally redistributed, and existing corpora are small, single-platform, or both. AI-generated music is the natural substrate to close this gap.

We introduce LAION-Tunes, an open dataset of URLs and metadata for AI-generated music from Suno, Udio, and Mureka. The release ships public CDN URLs, 768-dimensional audio and text embeddings, captions, ASR transcription embeddings—not raw transcripts—five-dimensional aesthetics scores, real and predicted engagement counts, and three-axis NSFW safety labels under Apache 2.0. We also release the models used to annotate the corpus: a 242M-parameter audio model and fingerprint extractor, a calibrated quality scorer, and an engagement predictor trained on platform play and upvote counts.

To demonstrate downstream utility, we construct a perceptual benchmark from AI-generated music and genre-matched human controls. Human listeners and language-model configurations evaluated the same songs. Human listeners detect AI music above chance but modestly, reaching 62.5% accuracy and outperforming every LLM configuration on balanced accuracy. We further identify a quality–authenticity halo effect: songs receiving higher aesthetic ratings are more likely to be judged human-made. Equivalently, songs judged real receive nearly two more aesthetic-quality points than songs judged AI-generated, independent of true provenance.

All dataset artifacts, models, and code are available through the anonymous double-blind review repository, alongside a live LAION-Tunes search interface.

Paper 02

DEFINE

Abstract

Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining.

Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps short accent exemplars into a conditioning space using an exemplar encoder supervised through learned accent prototypes, requiring neither accent labels at inference time nor post-synthesis waveform conversion. On seen accents, increasing accent guidance improves accent-probe accuracy from 6.5% to 19.6%.

More importantly, a single DEFINE model generalizes accent control beyond its training accent set: on seen and out-of-domain accents, though not on held-out accents, it matches the accent-transfer performance of a two-model TTS–voice-conversion cascade while achieving higher speaker similarity and comparable predicted speech quality. These results demonstrate that speaker identity and accent can be independently controlled from audio exemplars within a single zero-shot TTS model, including for accents unseen during training.