More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration

arXiv cs.CL / 4/10/2026

💬 OpinionSignals & Early TrendsIdeas & Deep AnalysisModels & Research

共有:

Key Points

The paper investigates when LLM agents fail to cooperate in a “zero-cost collaboration” setting where helping others has no direct personal cost, focusing on cooperation failures separate from competence issues.
It shows that higher capability does not reliably translate to better cooperative outcomes: OpenAI o3 attains only 17% of optimal collective performance while o3-mini reaches 50% under identical group-revenue maximizing instructions.
Using causal decomposition with automated analysis of agent communication, the authors disentangle cooperation failures from competence failures and trace the causes to agents’ reasoning and interaction dynamics.
Targeted interventions reveal that explicit cooperative protocols can roughly double performance for lower-competence models, while small sharing incentives can improve cooperation for models with weak cooperative tendencies.
The study concludes that scaling intelligence alone is unlikely to eliminate coordination problems in multi-agent systems, emphasizing the need for deliberate cooperative design and alignment of interaction mechanisms.

Abstract

Large language model (LLM) agents increasingly coordinate in multi-agent systems, yet we lack an understanding of where and why cooperation failures may arise. In many real-world coordination problems, from knowledge sharing in organizations to code documentation, helping others carries negligible personal cost while generating substantial collective benefits. However, whether LLM agents cooperate when helping neither benefits nor harms the helper, while being given explicit instructions to do so, remains unknown. We build a multi-agent setup designed to study cooperative behavior in a frictionless environment, removing all strategic complexity from cooperation. We find that capability does not predict cooperation: OpenAI o3 achieves only 17% of optimal collective performance while OpenAI o3-mini reaches 50%, despite identical instructions to maximize group revenue. Through a causal decomposition that automates one side of agent communication, we separate cooperation failures from competence failures, tracing their origins through agent reasoning analysis. Testing targeted interventions, we find that explicit protocols double performance for low-competence models, and tiny sharing incentives improve models with weak cooperation. Our findings suggest that scaling intelligence alone will not solve coordination problems in multi-agent systems and will require deliberate cooperative design, even when helping others costs nothing.

Black Hat Asia

AI Business

GLM 5.1 tops the code arena rankings for open models

Reddit r/LocalLLaMA

Big Tech firms are accelerating AI investments and integration, while regulators and companies focus on safety and responsible adoption.

Dev.to

My Bestie Built a Free MCP Server for Job Search — Here's How It Works

Dev.to

can we talk about how AI has gotten really good at lying to you?

Reddit r/artificial

More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration

Key Points

Abstract

Related Articles

Black Hat Asia

GLM 5.1 tops the code arena rankings for open models

Big Tech firms are accelerating AI investments and integration, while regulators and companies focus on safety and responsible adoption.

My Bestie Built a Free MCP Server for Job Search — Here's How It Works

can we talk about how AI has gotten really good at lying to you?

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer