Soft-Label Governance for Distributional Safety in Multi-Agent Systems

arXiv cs.AI / 4/23/2026

💬 OpinionSignals & Early TrendsIdeas & Deep AnalysisModels & Research

共有:

Key Points

The paper introduces SWARM, a simulation framework for multi-agent safety that replaces binary “good/bad” behavior labels with soft probabilistic labels for more faithful risk and uncertainty modeling.
SWARM’s governance engine can be configured with intervention levers such as transaction taxes, circuit breakers, reputation decay, and random audits, and it evaluates outcomes using continuous probabilistic metrics (e.g., expected toxicity and quality gap).
Experiments across seven scenarios (with five-seed replication) find that strict governance can significantly reduce overall welfare by 40%+ without improving safety, highlighting the need for balanced policy design.
The study also shows that some governance settings—like circuit breakers—must be carefully calibrated: overly strict thresholds sharply reduce system value, while an intermediate optimal threshold can better trade off welfare and toxicity.
Companion experiments indicate that soft metrics can detect “proxy gaming” by agents that exploit conventional binary evaluations, and the approach can be applied to live LLM-backed agents without modification.

Abstract

Multi-agent AI systems exhibit emergent risks that no single agent produces in isolation. Existing safety frameworks rely on binary classifications of agent behavior, discarding the uncertainty inherent in proxy-based evaluation. We introduce SWARM (\textbf{S}ystem-\textbf{W}ide \textbf{A}ssessment of \textbf{R}isk in \textbf{M}ulti-agent systems), a simulation framework that replaces binary good/bad labels with \emph{soft probabilistic labels}

p = P(v{=}+1) \in [0,1]

, enabling continuous-valued payoff computation, toxicity measurement, and governance intervention. SWARM implements a modular governance engine with configurable levers (transaction taxes, circuit breakers, reputation decay, and random audits) and quantifies their effects through probabilistic metrics including expected toxicity

\mathbb{E}[1{-}p \mid \text{accepted}]

and quality gap

\mathbb{E}[p \mid \text{accepted}] - \mathbb{E}[p \mid \text{rejected}]

. Across seven scenarios with five-seed replication, strict governance reduces welfare by over 40\% without improving safety. In parallel, aggressively internalizing system externalities collapses total welfare from a baseline of

+262

down to

-67

, while toxicity remains invariant. Circuit breakers require careful calibration; overly restrictive thresholds severely diminish system value, whereas an optimal threshold balances moderate welfare with minimized toxicity. Companion experiments show soft metrics detect proxy gaming by self-optimizing agents passing conventional binary evaluations. This basic governance layer applies to live LLM-backed agents (Concordia entities, Claude, GPT-4o Mini) without modification. Results show distributional safety requires \emph{continuous} risk metrics and governance lever calibration involves quantifiable safety-welfare tradeoffs. Source code and project resources are publicly available at https://www.swarm-ai.org/.

I’m working on an AGI and human council system that could make the world better and keep checks and balances in place to prevent catastrophes. It could change the world. Really. Im trying to get ahead of the game before an AGI is developed by someone who only has their best interest in mind.

Reddit r/artificial

Deepseek V4 Flash and Non-Flash Out on HuggingFace

Reddit r/LocalLLaMA

DeepSeek V4 Flash & Pro Now out on API

Reddit r/LocalLLaMA

I’m building a post-SaaS app catalog on Base, and here’s what that actually means

Dev.to

r/LocalLLaMa Rule Updates

Reddit r/LocalLLaMA

Soft-Label Governance for Distributional Safety in Multi-Agent Systems

Key Points

Abstract

Related Articles

I’m working on an AGI and human council system that could make the world better and keep checks and balances in place to prevent catastrophes. It could change the world. Really. Im trying to get ahead of the game before an AGI is developed by someone who only has their best interest in mind.

Deepseek V4 Flash and Non-Flash Out on HuggingFace

DeepSeek V4 Flash & Pro Now out on API

I’m building a post-SaaS app catalog on Base, and here’s what that actually means

r/LocalLLaMa Rule Updates

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer