Image Classifiers are Efficient Self-Supervised
Video Representation Learners

1 Indian Institute of Technology Kharagpur 2 Indian Institute of Science Bangalore
3 École de technologie supérieure, Montreal
🎉 BMVC 2026

Can image classifiers learn video representations efficiently?

Kinetics-400 top-1 accuracy versus pretraining epochs for VideoMSN and prior self-supervised video methods
Comparison of top-1 accuracy on Kinetics-400 across state-of-the-art self-supervised video representation learning methods. Each point is a method, with bubble size proportional to its pretraining epochs. VideoMSN with a DINO-v3 backbone reaches state-of-the-art accuracy with 160× and 60× fewer training epochs than VideoMAE and SMILE, respectively.

Abstract

We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images: grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to 32× fewer and 160× fewer video pretraining epochs compared to prior video self-supervised learning methods. The approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario.

🔗 Method

VideoMSN starts from unlabeled clips, samples \(M\) equidistant frames, and arranges them into a 2D super image following SIFAR. Each super image produces one unmasked target view and a set of masked anchor views. The target encoder is an exponential moving average of the anchor encoder. Both encoders are ordinary 2D ViTs, initialized from DeiT-v3 or DINO-v3, with no 3D inflation and no decoder.

â–· Random views

Each frame is patchified, and a spatial mask drops the same patch located across all frames. Masked tokens never enter the encoder, which both reduces compute and blocks temporal copy-paste shortcuts.

â–· Fast and slow focal views

Temporal masking keeps about \(M/2\) frames (fast, \(3\times3\) grid) or \(M/4\) frames (slow, \(2\times2\) grid). Focal masking then retains a contiguous spatial region in every remaining frame, so the model sees the same action at different temporal sparsities.

â–· Prototype alignment

Anchor and target [CLS] embeddings are compared with 1024 learnable prototypes and converted into distributions. Training minimizes cross-entropy between them, plus an ME-MAX regularizer that encourages uniform prototype use.

The alignment loss over \(K\) anchor views in a batch of size \(B\) is \(\mathcal{L}_{\mathrm{align}} = \frac{1}{KB}\sum_{i=1}^{B}\sum_{j=1}^{K} H(\mathbf{p}^i_j, \mathbf{p}^i_+)\), and the full objective is \(\mathcal{L}_{\mathrm{total}} = \mathcal{L}_{\mathrm{align}} - \lambda H(\bar{\mathbf{p}})\) with \(\lambda=5.0\). Temperatures are \(\tau=0.1\) for anchors and \(\tau^+=0.025\) for targets. Unless noted, we drop 70% of tokens (50% on SSV2) and use six focal views: three fast and three slow.

📈 Kinetics-400 Results

Self-supervised pretraining on unlabeled Kinetics-400, then full fine-tuning. VideoMSN-DeiT is pretrained for 50 epochs and VideoMSN-DINO for 10 epochs. Competing MAE-style methods typically use 800–1600 epochs and an extra decoder.

RankMethodEpochsParams (M)Top-1
1VideoMSN-DINO (Ours)102180.8
2VideoMSN-DeiT (Ours)5021.580.0
3SMILE (motion)8002279.5
4SIGMA-DINO8002279.4
5VideoMAE8002279.0

VideoMSN-DINO (ViT-S) improves the previous best by 1.3% after only 10 video epochs. Parameter counts are encoder-only; competing methods add a decoder at pretraining time.

🚀 Efficiency

160×fewer video pretraining epochs than VideoMAE (10 vs 1600)
12.3×lower wall-clock time on 4x H100 GPUs (21.7 h vs 266.7 h)
8.9×fewer total pretraining FLOPs (4.58 E vs 40.89 E)
MethodEpochsTime / epochWall-clockTotal FLOPs
VideoMAE (ViT-B)1600~10 min266.7 h40.89 E
VideoMSN-DINO (ViT-B)10~130 min21.7 h4.58 E

VideoMSN is heavier per epoch because a super image is a large 2D grid, but the much shorter schedule still dominates total cost. As a control, initializing VideoMAE from ImageNet-pretrained ViT-B and training it for only 50 video epochs yields 57% top-1 on Kinetics-400, suggesting that decoder-based MAE is a poor fit for short adaptation of image foundation models.

For more details and experimental results check our paper.

BibTeX


        Will be updated.