Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

arXiv cs.CV / 5/6/2026

📰 NewsModels & Research

共有:

Key Points

Traditional video object-centric learning enforces temporal consistency by training learned dynamics modules that predict future object slots, but the work argues these predictors are effectively costly approximations of discrete correspondence.
The paper shows that modern self-supervised vision backbones already provide instance-discriminative features, allowing temporal prediction to be unnecessary for identity consistency.
It proposes Grounded Correspondence, which maintains frame-to-frame identity using deterministic bipartite matching (Hungarian matching) over slot representations instead of learned transition functions.
Slots are initialized from salient regions using frozen backbone features, and the method uses zero learnable parameters for temporal modeling while still achieving competitive results on MOVi-D, MOVi-E, and YouTube-VIS.

Abstract

The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We demonstrate that these predictors function as expensive approximations of discrete correspondence problems. Modern self-supervised vision backbones already encode instance-discriminative features that distinguish objects reliably. Exploiting these features eliminates the need for learned temporal prediction. We introduce Grounded Correspondence, a framework that replaces learned transition functions with deterministic bipartite matching. Slots initialize from salient regions in frozen backbone features. Frame-to-frame identity is maintained through Hungarian matching on slot representations. The approach requires zero learnable parameters for temporal modeling yet achieves competitive performance on MOVi-D, MOVi-E, and YouTube-VIS. Project page: https://magenta-sherbet-85b101.netlify.app/

Google AI Releases Multi-Token Prediction (MTP) Drafters for Gemma 4: Delivering Up to 3x Faster Inference Without Quality Loss

MarkTechPost

Solidity LM surpasses Opus

Reddit r/LocalLLaMA

Quality comparison between Qwen 3.6 27B quantizations (BF16, Q8_0, Q6_K, Q5_K_XL, Q4_K_XL, IQ4_XS, IQ3_XXS,...)

Reddit r/LocalLLaMA

We measured the real cost of running a GPT-5.4 chatbot on live websites

Reddit r/artificial

AI ecosystems in China and US grow apart amid tech war

SCMP Tech

Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

Key Points

Abstract

Related Articles

Google AI Releases Multi-Token Prediction (MTP) Drafters for Gemma 4: Delivering Up to 3x Faster Inference Without Quality Loss

Solidity LM surpasses Opus

Quality comparison between Qwen 3.6 27B quantizations (BF16, Q8_0, Q6_K, Q5_K_XL, Q4_K_XL, IQ4_XS, IQ3_XXS,...)

We measured the real cost of running a GPT-5.4 chatbot on live websites

AI ecosystems in China and US grow apart amid tech war

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer