Репост из: the last neural cell
🧬 Tasty AI papers | 01-31 July 2024
💎Vision models
Genie: Generative Interactive Environments
What: learn latent actions from videos (only) of games.
- predict future frames based on previous and latent actions.
- they trained actions to help model make transition between frames.
- just let’s AI model figures out commands by yourself.
SAM 2: Segment Anything in Images and Videos
What: SAM now works well with videos.
- annotate big dataset of videos.
- add memory block to ensure temporal consistency of predicted mask.
💎 General
Mixture of A Million Experts
What: expand MoE for lots of experts.
- store low rank approx of experts.
- works better than dense FFN.
The Road Less Scheduled
What: propose schedule-free optimizer.
- one more thing that beats AdamW.
- easy to drop in your training pipeline.
🔘 Diffusion
Rolling Diffusion Models
What: incorporating temporal info in generative diffusion process for videos.
- let’s make denoising and predict next frames at the same time.
- hard math, but idea is interesting.
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
What: step into merging local and global planning.
#digest
💎Vision models
Genie: Generative Interactive Environments
What: learn latent actions from videos (only) of games.
- predict future frames based on previous and latent actions.
- they trained actions to help model make transition between frames.
- just let’s AI model figures out commands by yourself.
SAM 2: Segment Anything in Images and Videos
What: SAM now works well with videos.
- annotate big dataset of videos.
- add memory block to ensure temporal consistency of predicted mask.
💎 General
Mixture of A Million Experts
What: expand MoE for lots of experts.
- store low rank approx of experts.
- works better than dense FFN.
The Road Less Scheduled
What: propose schedule-free optimizer.
- one more thing that beats AdamW.
- easy to drop in your training pipeline.
🔘 Diffusion
Rolling Diffusion Models
What: incorporating temporal info in generative diffusion process for videos.
- let’s make denoising and predict next frames at the same time.
- hard math, but idea is interesting.
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
What: step into merging local and global planning.
Our approach is shown to combine the strengths of next-token prediction models, such as variable-length generation, with the strengths of full-sequence diffusion models, such as the ability to guide sampling to desirable trajectories.
#digest