Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

48 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

midi-generation

Midi Generation uses a transformer-based VAE to conditionally generate classical music MIDI files based on a given composer. Overall, the current version of the model has weak label conditioning and generates music that sounds similar regardless of composer. However, the actual sound of this music is not too bad. We recommend trying out a view of the .mid files included in the git.

Model Architecture

image We utilize PyTorch's built in Transformer Encoder and Decoder objects to build a Variational Autoencoder to generate MIDI sequences. We then feed that MIDI sequence through a pre-trained GRU classifier to obtain a confidence score. We use CrossEntropy to measure the reconstruction loss, and use KL-Divergence with a mixture of gaussian prior distribution. During inference, all that's required for the decoder is feeding a random gaussian vector with the correct label mean and a label, upon doing this it will generate a full 500 token length MIDI sequence which can be easily saved as a file. The encoder follows traditional VAE architecture, basically appended to the output of a transformer encoder. For the decoder we use both self and cross attention to strengthen the label conditioning.

Comparing Model Accuracy

image The accuracy of our models are almost always below random baseline. We used a subset of 12 classes of our full 34 class dataset.The only model that does beat random baseline is our original model with no LayerNorm. This shows that the generated sequences are very similar regardless of label.

Comparing Model Relative Cosine Similarity

image The accuracy of our models are almost always below random baseline. Each of the models gives a near-zero relative cosine similarity score, illustrating that our generated sequences are not unique for each label, being just as similar to other classes.

Visualization of Latent Space

image Our latent space shows a moderate amount of separation between classes, with a very noisy bunch within the center of the mean. This means that the classes separated from the rest probably have more unique generated sequences.

Comparing Model Silhouette Score

image Our latent space shows a moderate amount of separation between classes, with a 0.2 silhouette score regardless of model. This means that the latent space is closer to being mixed than being cleanly separated, but it is significantly above 0.

Confusion Matrix of Best Model

image Model loves guessing Frederic Chopin. Remember to use git lfs pull to download the MIDI tensor files (and to make sure you have git lfs installed). A regular git pull will not download them.

Using MIDIVAE_new.ipynb (and variants)

To run MIDIVAE_new.ipynb, we strongly suggest following the below proceedure:

  1. Go to the SCC website, click on Files and upload the maestro_new_splits_no_augmentation.csv (click here to download it) and MIDIVAE.ipynb files in your choice of directory (or use our project directory: /projectnb/ec523/projects/proj_MIDIgen, which already has the files in it)
  2. Click on click on Interactive Apps, and click on Jupyter Notebooks
  3. We are going to request and launch an interactive Jupyter Notebook session with the following parameters:
    • List of modules to load (space separated): miniconda academic-ml/spring-2025
    • Pre-Launch Command (optional): conda activate spring-2025-pyt
    • Interface: notebook
    • Working Directory: wherever you uploaded the files in step 1
    • Number of hours: 12
    • Number of cores: 4
    • Number of GPUs: 1
    • GPU compute capability: 8.0
  4. Click launch and wait for your session to begin
  5. Once your session begins, run all cells and a new model will begin training.
  6. Once the model is trained, it will automatically save, install prettyMIDI with pip, then run inference to generate a .MID music file
  7. To listen to this file, download it and run it using your favorite MIDI synthesizer or convert it to a .WAV with a free online tool such as this one. You can then play the WAV on virtually any music playing app such as VLC or Windows Media Player

Warning: We highly recommend AGAINST running MIDIVAE.ipynb in Google Colab. Google Colab has a different configuration for PyTorch than the SCC's academic-ml/spring-2025 they are NOT cross-compatible.

Using GRU Classifier (music_classifier.py)

To run the file, you should be connected to a GPU. The model was trained locally using an NVIDIA 4070 Ti SUPER, but it will work on any CUDA capable device. Training 10 epochs with: Embedding layer size: 128 Hidden size: 128 Layer number: 2 To train the model, you need to:

  1. Load in the dataset (as tokenized sequences)
  2. Run the code.
  3. The model should take about 1.5-2 hours to finish training.
  4. Once done, the model will run the evaluation code to test itself out. It will print the accuracies afterwards.

To run the pre-trained model, you need to:

  1. Load in pre-trained model
  2. Load in the tokenized MIDI sequence to test out.
  3. Run the model with the sequence. After a few seconds, the model will generate an output of likelihoods, and tell you the top 1 and top 3.
  4. You're done.

Generating preprocessed data using MIDI_preprocess.ipynb

  1. Upload two copies of the MAESTRO dataset to a Google Drive account you have access to and make sure the directory is named according to the directory specified in the data_directory variable. Also make sure that the name output .csv file of the initial token generation matches the name of the file in input_csv, and that the directory specified in the pm variable in the final cell matches the second copy of the dataset you uploaded.
  2. Run each cell in order from top to bottom using Google Colab.
  3. The output .csv should be in the content directory.
  4. Download the .csv
  5. That's it. The current token vocabulary used will generate a file that is roughly 6.7GB that needs to be zipped, so this might take a while.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages