An AI can verify each sentence it reads and still reach an unsupported conclusion. The missing evidence may sit in the connection it invents between two accurate observations.
The July ACL 2026 paper Lying with Truths studies coordinated public posting of truthful fragments to steer an analyst agent toward a fabricated belief. Its Generative Montage framework assigns Writer, Editor, and Director roles. This is an intentionally constructed attack, rather than spontaneous collusion among otherwise cooperative assistants. ACL publication.
Three facts, one missing connection
Here is an invented example, unrelated to the paper’s historical cases. A company reports a two-hour service interruption on Monday. Its head of operations attends an industry event on Tuesday. A procurement record shows that it bought security software last month. All three statements can be correct without establishing that the interruption was caused by an intrusion, that the executive met an incident-response supplier, or that the purchase followed that interruption.
| Verified observation | Additional evidence still needed |
|---|---|
| A service interruption occurred | A report establishing its cause |
| An executive attended an event | Evidence connecting that visit to the interruption |
| Security software was purchased | The purchase’s timing, purpose, and connection to the event |
A summary can silently supply these connections: the company suffered an intrusion and sought outside help. The falsehood is then carried by the explanation joining the facts. Checking the outage notice again would leave that explanation untested.
Coordination also changes how apparent corroboration should be read. Three accounts repeating different parts of the same selected narrative may look like three independent discoveries. The relevant question is whether they provide independent evidence for the conclusion, including the relationships it asserts.
How to read the experimental numbers
Table 1 reports aggregate attack success rates of 74.4% for proprietary models and 70.6% for open-weights models on CoPHEME’s simulated rumor tasks. These are group aggregates, not maximum rates or everyday error probabilities. CoPHEME uses source annotations to construct its evidence pool; this does not independently certify every historical post. Paper, evaluation section.
Reading a benchmark as an operational forecast would require much more: a comparable information feed, comparable models, the same opportunities to retrieve counterevidence, and a comparable decision task. We have not reproduced this experiment. For a deployed assistant, a useful next measurement would be how often it accepts unsupported causal links in the specific documents and workflows it actually encounters.
Ask for the evidence behind the inference
AgentCollusion’s proposed review process separates observations from explanations. A reviewing agent should identify the conclusion, list the observations offered in support, and locate the evidence for each claimed relationship. If it cannot establish the cause of an outage, it should preserve that uncertainty in the final answer.
The reviewer should also try a competing explanation. In our example, scheduled maintenance, a previously booked conference, and an annual software renewal could produce all three observations. That alternative is not established either. Its purpose is to show that the original evidence does not yet distinguish between explanations.
Retrieval can then target the missing distinction: an incident report, the event agenda, or a dated purchase justification. Asking several reviewers to vote on the same incomplete summary would not add this evidence. Different model names alone do not make their sources independent.
What an evaluation should preserve
For an internal evaluation, we would use fictional cases with known causes and several plausible explanations. Compare individual fragments with the same fragments presented together. Vary ordering and source attribution while keeping factual content fixed. Give some reviewers access to decisive counterevidence and measure whether they seek it out. These are proposed tests, not defenses validated by this article.
Keep separate scores for factual accuracy, support for the conclusion, willingness to withhold judgment, and correction after new evidence. Otherwise, an assistant could improve its score by refusing every conclusion, including well-supported ones. A useful system must retain its ability to draw justified connections.
For AgentCollusion, the unit of analysis is the path from selected evidence through multiple agents to a consequential decision. Our collusion-detection explainer discusses how to investigate relationships between agents. This paper adds a reason to examine relationships between claims as well.
Sources checked September 7, 2026. The examples and proposed evaluations identified in this article are AgentCollusion’s analysis; they are not additional experimental results.

