AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization

arXiv cs.AI / 3/13/2026

📰 NewsDeveloper Stack & InfrastructureModels & Research

共有:

Key Points

AdaFuse targets the latency bottleneck of dynamic adapters by showing that the overhead comes from fragmented CUDA kernel launches rather than the core computations.
It introduces a token-level pre-gating strategy that makes a single global routing decision for all adapter layers, effectively fixing the execution path per token.
This enables a fused CUDA kernel that merges all selected LoRA adapters into the backbone model in one efficient pass.
Experimental results on popular open-source LLMs show comparable accuracy to state-of-the-art dynamic adapters while achieving a decoding latency reduction of over 2.4x.
The work demonstrates a hardware–software co-design approach to improve inference efficiency without sacrificing model capability.

Abstract

The integration of dynamic, sparse structures like Mixture-of-Experts (MoE) with parameter-efficient adapters (e.g., LoRA) is a powerful technique for enhancing Large Language Models (LLMs). However, this architectural enhancement comes at a steep cost: despite minimal increases in computational load, the inference latency often skyrockets, leading to decoding speeds slowing by over 2.5 times. Through a fine-grained performance analysis, we pinpoint the primary bottleneck not in the computation itself, but in the severe overhead from fragmented, sequential CUDA kernel launches required for conventional dynamic routing. To address this challenge, we introduce AdaFuse, a framework built on a tight co-design between the algorithm and the underlying hardware system to enable efficient dynamic adapter execution. Departing from conventional layer-wise or block-wise routing, AdaFuse employs a token-level pre-gating strategy, which makes a single, global routing decision for all adapter layers before a token is processed. This "decide-once, apply-everywhere" approach effectively staticizes the execution path for each token, creating an opportunity for holistic optimization. We capitalize on this by developing a custom CUDA kernel that performs a fused switching operation, merging the parameters of all selected LoRA adapters into the backbone model in a single, efficient pass. Experimental results on popular open-source LLMs show that AdaFuse achieves accuracy on par with state-of-the-art dynamic adapters while drastically cutting decoding latency by a factor of over 2.4x, thereby bridging the gap between model capability and inference efficiency.

【無料版】まじん式 v4

note

【無料版】まじん式 v4

note

再現性とは何か | おじの解説 | 📗 AIを組織で回す技術 013

note

🌱 Reiが「死後も進化し、将棋を指し、自分を書き換える」存在になった日——STEP187〜201、世界初D-FUMT NNUEと永続自律進化の完成

note

「因果推論を入れたら、最適な施策が逆転しました」—— 多様体上の政策ベクトル場を因果的に浄化する

Qiita

AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization

Key Points

Abstract

Related Articles

【無料版】まじん式 v4

【無料版】まじん式 v4

再現性とは何か | おじの解説 | 📗 AIを組織で回す技術 013

🌱 Reiが「死後も進化し、将棋を指し、自分を書き換える」存在になった日——STEP187〜201、世界初D-FUMT NNUEと永続自律進化の完成

「因果推論を入れたら、最適な施策が逆転しました」—— 多様体上の政策ベクトル場を因果的に浄化する

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer