PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding

arXiv cs.CV / 4/2/2026

📰 NewsSignals & Early TrendsIdeas & Deep AnalysisModels & Research

共有:

Key Points

The paper introduces PixelPrune, a training-free pixel-level method that removes redundant image patches before the ViT encoder in vision-language pipelines.
It leverages the observation that only 22–71% of patches are pixel-unique in document and GUI benchmarks, enabling predictive-coding-based compression in pixel space.
PixelPrune accelerates both the ViT encoder and the downstream LLM by reducing visual token counts early in the inference pipeline.
The approach supports pixel-lossless compression (τ=0) as well as controlled lossy modes (τ>0) without learnable parameters.
Experiments on three model scales across document and GUI benchmarks report up to 4.2× inference speedup and up to 1.9× training acceleration while maintaining competitive accuracy.

Abstract

Document understanding and GUI interaction are among the highest-value applications of Vision-Language Models (VLMs), yet they impose exceptionally heavy computational burden: fine-grained text and small UI elements demand high-resolution inputs that produce tens of thousands of visual tokens. We observe that this cost is largely wasteful -- across document and GUI benchmarks, only 22--71\% of image patches are pixel-unique, the rest being exact duplicates of another patch in the same image. We propose \textbf{PixelPrune}, which exploits this pixel-level redundancy through predictive-coding-based compression, pruning redundant patches \emph{before} the Vision Transformer (ViT) encoder. Because it operates in pixel space prior to any neural computation, PixelPrune accelerates both the ViT encoder and the downstream LLM, covering the full inference pipeline. The method is training-free, requires no learnable parameters, and supports pixel-lossless compression (

\tau{=}0

) as well as controlled lossy compression (

\tau{>}0

). Experiments across three model scales and document and GUI benchmarks show that PixelPrune maintains competitive task accuracy while delivering up to 4.2

\times

inference speedup and 1.9

\times

training acceleration. Code is available at https://github.com/OPPO-Mente-Lab/PixelPrune.

Black Hat Asia

AI Business

v5.5.0

Transformers（HuggingFace）Releases

Bonsai (PrismML's 1 bit version of Qwen3 8B 4B 1.7B) was not an aprils fools joke

Reddit r/LocalLLaMA

Big Tech firms are accelerating AI investments and integration, while regulators and companies focus on safety and responsible adoption.

Dev.to

Inference Engines - A visual deep dive into the layers of an LLM

Dev.to

PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding

Key Points

Abstract

Related Articles

Black Hat Asia

v5.5.0

Bonsai (PrismML's 1 bit version of Qwen3 8B 4B 1.7B) was not an aprils fools joke

Big Tech firms are accelerating AI investments and integration, while regulators and companies focus on safety and responsible adoption.

Inference Engines - A visual deep dive into the layers of an LLM

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer