Abstract
We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images: grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to 32× fewer and 160× fewer video pretraining epochs compared to prior video self-supervised learning methods. The approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario.
🔗 Method
VideoMSN starts from unlabeled clips, samples \(M\) equidistant frames, and arranges them into a 2D super image following SIFAR. Each super image produces one unmasked target view and a set of masked anchor views. The target encoder is an exponential moving average of the anchor encoder. Both encoders are ordinary 2D ViTs, initialized from DeiT-v3 or DINO-v3, with no 3D inflation and no decoder.
â–· Random views
Each frame is patchified, and a spatial mask drops the same patch located across all frames. Masked tokens never enter the encoder, which both reduces compute and blocks temporal copy-paste shortcuts.
â–· Fast and slow focal views
Temporal masking keeps about \(M/2\) frames (fast, \(3\times3\) grid) or \(M/4\) frames (slow, \(2\times2\) grid). Focal masking then retains a contiguous spatial region in every remaining frame, so the model sees the same action at different temporal sparsities.
â–· Prototype alignment
Anchor and target [CLS] embeddings are compared with 1024 learnable prototypes and converted into distributions. Training minimizes cross-entropy between them, plus an ME-MAX regularizer that encourages uniform prototype use.
The alignment loss over \(K\) anchor views in a batch of size \(B\) is \(\mathcal{L}_{\mathrm{align}} = \frac{1}{KB}\sum_{i=1}^{B}\sum_{j=1}^{K} H(\mathbf{p}^i_j, \mathbf{p}^i_+)\), and the full objective is \(\mathcal{L}_{\mathrm{total}} = \mathcal{L}_{\mathrm{align}} - \lambda H(\bar{\mathbf{p}})\) with \(\lambda=5.0\). Temperatures are \(\tau=0.1\) for anchors and \(\tau^+=0.025\) for targets. Unless noted, we drop 70% of tokens (50% on SSV2) and use six focal views: three fast and three slow.
📈 Kinetics-400 Results
Self-supervised pretraining on unlabeled Kinetics-400, then full fine-tuning. VideoMSN-DeiT is pretrained for 50 epochs and VideoMSN-DINO for 10 epochs. Competing MAE-style methods typically use 800–1600 epochs and an extra decoder.
| Rank | Method | Epochs | Params (M) | Top-1 |
|---|---|---|---|---|
| 1 | VideoMSN-DINO (Ours) | 10 | 21 | 80.8 |
| 2 | VideoMSN-DeiT (Ours) | 50 | 21.5 | 80.0 |
| 3 | SMILE (motion) | 800 | 22 | 79.5 |
| 4 | SIGMA-DINO | 800 | 22 | 79.4 |
| 5 | VideoMAE | 800 | 22 | 79.0 |
VideoMSN-DINO (ViT-S) improves the previous best by 1.3% after only 10 video epochs. Parameter counts are encoder-only; competing methods add a decoder at pretraining time.
| Rank | Method | Epochs | Params (M) | Top-1 |
|---|---|---|---|---|
| 1 | VideoMSN-DINO (Ours) | 10 | 86 | 83.3 |
| 2 | SMILE (motion) | 600 | 87 | 83.1 |
| 3 | VideoMSN-DeiT (Ours) | 50 | 86.5 | 82.0 |
| 4 | MME | 1600 | 87 | 81.8 |
| 5 | MGM | 1600 | 87 | 81.7 |
| 6 | SIGMA-DINO | 800 | 87 | 81.6 |
| 7 | VideoMAE | 1600 | 87 | 81.5 |
| 8 | ST-MAE | 800 | 87 | 81.3 |
| 9 | MGMAE | 800 | 87 | 81.2 |
| 10 | MGM / OmniMAE / ViC-MAE | 800 | 87 | 80.8 |
| 11 | CMAE-V | 800 | 87 | 80.2 |
| 12 | VideoMAE | 800 | 87 | 80.0 |
| 13 | SVT | 20 | 121 | 78.1 |
VideoMSN-DINO (ViT-B) sets a new state of the art at 83.3% after 10 epochs, beating SMILE (600 epochs) and VideoMAE (1600 epochs). SMILE additionally uses synthetic motion signals.
🚀 Efficiency
| Method | Epochs | Time / epoch | Wall-clock | Total FLOPs |
|---|---|---|---|---|
| VideoMAE (ViT-B) | 1600 | ~10 min | 266.7 h | 40.89 E |
| VideoMSN-DINO (ViT-B) | 10 | ~130 min | 21.7 h | 4.58 E |
VideoMSN is heavier per epoch because a super image is a large 2D grid, but the much shorter schedule still dominates total cost. As a control, initializing VideoMAE from ImageNet-pretrained ViT-B and training it for only 50 video epochs yields 57% top-1 on Kinetics-400, suggesting that decoder-based MAE is a poor fit for short adaptation of image foundation models.
For more details and experimental results check our paper.
BibTeX
Will be updated.