Video-Audio Separation through Text (VAST)

Codebase for submission"Language-Guided Audio-Visual Source Separation via Trimodal Consistency".

Environment

The code is developed under the following configurations.

Hardware: 1-4 GPUs (change [--num_gpus NUM_GPUS] accordingly)
Software: Ubuntu 16.04.3 LTS, CUDA>=10.3, Python>=3.5, PyTorch>=1.9.0

Training

Prepare video dataset.

a. Download MUSIC dataset from: https://github.com/roudimit/MUSIC_dataset

b. Download videos.

Preprocess videos. You can do it in your own way as long as the index files are similar.

a. Extract frames at 8fps and waveforms at 11025Hz from videos. We have following directory structure:

data
├── audio
|   ├── acoustic_guitar
│   |   ├── M3dekVSwNjY.mp3
│   |   ├── ...
│   ├── trumpet
│   |   ├── STKXyBGSGyE.mp3
│   |   ├── ...
│   ├── ...
|
└── frames
|   ├── acoustic_guitar
│   |   ├── M3dekVSwNjY.mp4
│   |   |   ├── 000001.jpg
│   |   |   ├── ...
│   |   ├── ...
│   ├── trumpet
│   |   ├── STKXyBGSGyE.mp4
│   |   |   ├── 000001.jpg
│   |   |   ├── ...
│   |   ├── ...
│   ├── ...

b. Make training/validation index files by running:

python scripts/create_index_files.py

It will create index files train.csv/val.csv with the following format:

./data/audio/acoustic_guitar/M3dekVSwNjY.mp3,./data/frames/acoustic_guitar/M3dekVSwNjY.mp4,1580
./data/audio/trumpet/STKXyBGSGyE.mp3,./data/frames/trumpet/STKXyBGSGyE.mp4,493

For each row, it stores the information: AUDIO_PATH,FRAMES_PATH,NUMBER_FRAMES

Train the default model.

./scripts/train_bimodal_cyclic_losses_solos_music.sh

During training, visualizations are saved in HTML format under ckpt/MODEL_ID/visualization/.

Evaluation

(Optional) Download our trained model weights for evaluation.

Updates

We have observed that finetuning the separation model with the latent captions with a very low learning rate further helps to improve performance. More details will come soon.

Name		Name	Last commit message	Last commit date
Latest commit History 124 Commits
dataset		dataset
models		models
scripts		scripts
README.md		README.md
bottleneck_bimodal_mask_cyclic_combined.py		bottleneck_bimodal_mask_cyclic_combined.py
extract_latent_concepts_main.py		extract_latent_concepts_main.py
teaser.png		teaser.png

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Video-Audio Separation through Text (VAST)

Environment

Training

Evaluation

Updates

About

Releases

Packages

Languages

rxtan2/AVSeT

Folders and files

Latest commit

History

Repository files navigation

Video-Audio Separation through Text (VAST)

Environment

Training

Evaluation

Updates

About

Resources

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages