Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization

arXiv cs.LG / 2026/3/25

💬 オピニオンIdeas & Deep AnalysisModels & Research

共有:

要点

The paper identifies a key weakness in existing vector quantization (VQ) approaches for multimodal tokenization: “Premature Discretization,” where discrete quantization is applied before the encoder has learned the data manifold.
It introduces Progressive Quantization (ProVQ), treating quantization hardness as a training curriculum that gradually anneals from continuous latents to discrete tokens.
Experiments show ProVQ improves reconstruction and generative performance on ImageNet-1K and ImageNet-100, indicating benefits for image generative modeling.
The method also performs strongly on complex biological sequence modeling, setting a new state-of-the-art performance ceiling for protein structure tokenization on StrutTokenBench.

Abstract

Vector Quantization (VQ) has become the cornerstone of tokenization for many multimodal Large Language Models and diffusion synthesis. However, existing VQ paradigms suffer from a fundamental conflict: they enforce discretization before the encoder has captured the underlying data manifold. We term this phenomenon Premature Discretization. To resolve this, we propose Progressive Quantization (ProVQ), which incorporates the dynamics of quantization hardness as a fundamental yet previously overlooked axis in VQ training. By treating quantization as a curriculum that smoothly anneals from a continuous latent space to a discrete one, ProVQ effectively guides the codebook toward the well-expanded manifolds. Extensive experimental results demonstrate the broad effectiveness of ProVQ across diverse modalities. We report improved reconstruction and generative performance on the ImageNet-1K and ImageNet-100 benchmarks, highlighting the ProVQ's boost for generative modeling. Furthermore, ProVQ proves highly effective for modeling complex biological sequences, establishing a new performance ceiling for protein structure tokenization on the StrutTokenBench leaderboard.

AIとロゴス

note

Speculative Decodingで27Bが逆に遅くなった

Qiita

信号処理の視点で見るデータ分析：共通点の整理と記事まとめ

Qiita

言語処理学会第32回年次大会(NLP2026) 参加報告

Qiita

AIと「ズッ友」になる魔法！─心をピタッと合わせるコツ

note

Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization

要点

Abstract

関連記事

AIとロゴス

Speculative Decodingで27Bが逆に遅くなった

信号処理の視点で見るデータ分析：共通点の整理と記事まとめ

言語処理学会第32回年次大会(NLP2026) 参加報告

AIと「ズッ友」になる魔法！─心をピタッと合わせるコツ

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer

要点

Abstract

関連記事

AIとロゴス

Speculative Decodingで27Bが逆に遅くなった

信号処理の視点で見るデータ分析：共通点の整理と記事まとめ

言語処理学会第32回年次大会(NLP2026) 参加報告

​AIと「ズッ友」になる魔法！─心をピタッと合わせるコツ

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer

AIと「ズッ友」になる魔法！─心をピタッと合わせるコツ