# Introducing Project Imagine

A state of the art dataset that is created to tackle scarcity of pre-training data for training artificial intelligence in music models.

In the past, datasets like **LAION-DISCO, Sleeping-DISCO, DISCO-10M**, tried to overcome this void by providing multiple minutes long metadata index by scouring YouTube and third-party sources but those methods faced extreme public backlash.

To overcome those backlash and provide a dataset that is equally high-quality and large enough, we are releasing **Project Imagine**.

Imagine is aimed to create the largest known pre-training dataset for research organisations and non-profit, and academic labs for training state-of-art artificial intelligence in music models.

For tasks like:

1. **LTS (Lyrics to Song)**
2. **Voice Cloning**
3. **Pre-training foundation audio models**

## What are the core benefits of Imagine?

Imagine consists of songs and recordings which are all generated by commercial AI models like **Suno, Udio, Mureka, Sonauto and Lyria**.

Whilst we don’t provide metadata for undisclosed reasons but we provide at the moment of writing **1.2TB (terabytes)** of artificial intelligence generated songs under Project Imagine.

## What countries were involved?

We made this dataset over the period of one year and consisted of **Europe, Southeast Asia and Eastern Europe**.

# LICENCE

Imagine is licensed under **Sleeping-Imagine version 1.0**.

That means Sleeping AI owns all the rights for this dataset and making any derivatives, publishing this dataset on torrents website, redistributing them without explicit permission of Sleeping AI is a crime.

### Project Status
It is an on-going project and we have released version 0 on Huggingface.

**Dataset:** https://huggingface.co/datasets/teamsleeping/imaginev0