DeCode: Decoupling Content and Delivery for Medical QA

arXiv cs.CL / 3/16/2026

📰 NewsModels & Research

共有:

Key Points

DeCode is a training-free, model-agnostic framework that decouples content and delivery to tailor LLM answers to individual clinical contexts.
It evaluates on OpenAI HealthBench and reports a zero-shot performance rise from 28.4% to 49.8%, achieving new state-of-the-art among existing methods.
The approach enables contextualized clinical QA without additional fine-tuning, facilitating deployment across existing LLMs in healthcare settings.
Experimental results suggest DeCode improves clinical relevance and validity of LLM responses, with practical benefits for patient-centered care.

Abstract

Large language models (LLMs) exhibit strong medical knowledge and can generate factually accurate responses. However, existing models often fail to account for individual patient contexts, producing answers that are clinically correct yet poorly aligned with patients' needs. In this work, we introduce DeCode (Decoupling Content and Delivery), a training-free, model-agnostic framework that adapts existing LLMs to produce contextualized answers in clinical settings. We evaluate DeCode on OpenAI HealthBench, a comprehensive and challenging benchmark designed to assess clinical relevance and validity of LLM responses. DeCode boosts zero-shot performance from 28.4% to 49.8% and achieves new state-of-the-art compared to existing methods. Experimental results suggest the effectiveness of DeCode in improving clinical question answering of LLMs.

14 Best Self-Hosted Claude Alternatives for AI and Coding in 2026

Dev.to

[P] Finetuned small LMs to VLM adapters locally and wrote a short article about it

Reddit r/MachineLearning

Experiment: How far can a 28M model go in business email generation?

Reddit r/LocalLLaMA

Qwen 3.5 397b (180gb) scores 93% on MMLU

Reddit r/LocalLLaMA

Qwen 3.5 27B - quantize KV cache or not?

Reddit r/LocalLLaMA

DeCode: Decoupling Content and Delivery for Medical QA

Key Points

Abstract

Related Articles

14 Best Self-Hosted Claude Alternatives for AI and Coding in 2026

[P] Finetuned small LMs to VLM adapters locally and wrote a short article about it

Experiment: How far can a 28M model go in business email generation?

Qwen 3.5 397b (180gb) scores 93% on MMLU

Qwen 3.5 27B - quantize KV cache or not?

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer