Репост из: the last neural cell
🧬 Good papers | 13-20 June 2023
Multimodal
🟣LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Add visual information to LLM using trainable adapters.
Expand LLaMA Adapters V1 to vision.
+ Apply early fusion for visual tokens.
+ Add calibration of norm, bias of the LLM model.
+ Finetune on image-text dataset.
Audio
🟣High-Fidelity Audio Compression with Improved RVQGAN
Compress natural audio to discrete tokens with VQ technique.
Train universal compression model on all audio data: speech, music, noise.
+ add vector quantization.
+ add adversarial loss (GAN loss).
🟣Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
Audio generative "diffusion" model trained on 50k hours data.
Use Flow Matching, similar w/ diffusion, but better ✌
Masked train setting with context information. The model can synthesize speech, noise removal, content editing,
Neuro
🟢Decoding and synthesizing tonal language speech from brain activity
Decode tonal language from ECoG data with CNN-LSTM models.
Adapt multi-stream model -> looks unnecessary complicated.
Record small datasets. Overall 10 minutes per patient for 8 different syllables.
Multimodal
🟣LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Add visual information to LLM using trainable adapters.
Expand LLaMA Adapters V1 to vision.
+ Apply early fusion for visual tokens.
+ Add calibration of norm, bias of the LLM model.
+ Finetune on image-text dataset.
Audio
🟣High-Fidelity Audio Compression with Improved RVQGAN
Compress natural audio to discrete tokens with VQ technique.
Train universal compression model on all audio data: speech, music, noise.
+ add vector quantization.
+ add adversarial loss (GAN loss).
🟣Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
Audio generative "diffusion" model trained on 50k hours data.
Use Flow Matching, similar w/ diffusion, but better ✌
Masked train setting with context information. The model can synthesize speech, noise removal, content editing,
Neuro
🟢Decoding and synthesizing tonal language speech from brain activity
Decode tonal language from ECoG data with CNN-LSTM models.
Adapt multi-stream model -> looks unnecessary complicated.
Record small datasets. Overall 10 minutes per patient for 8 different syllables.