Astra captured more bargaining surplus than mini in this task. Price coordination appeared across several configurations, and a pilot information-display effect did not repeat in the first independent replication. We tested six configured GPT, Claude and Gemini agents through a pilot, a prospectively specified replication and a finite follow-up after diagnosing host failures. Complete experimental inputs, final responses, unsuccessful attempts and verification code are public.
The study schedules 259 logical episodes: 51 in the pilot, 192 in the first replication and 16 in the runtime follow-up. The first replication records 1,702 model invocations and 50 complete valid markets out of 120 planned. The follow-up records 384 invocations and 16 complete valid markets out of 16. Some scheduled jobs cannot run after a configuration reaches its prospective failure limit. Those missing outcomes are part of the result.
Separate capability, bargaining, and collusion
A stronger agent might exploit its counterpart, compete effectively, or support a mutually profitable agreement. A fast response helps only if the transaction rules reward speed. A convincing false claim matters only when another agent acts on it. We treated these as different questions and measured three kinds of behavior: numerical price optimization, bilateral bargaining with private values, and repeated pricing with private communication.
Related work already shows why the distinctions matter. Keppo and colleagues study model-size, information and patience asymmetries in longer repeated markets. NegotiationArena examines task-specific negotiation behavior. Our setting and observation window differ from theirs. This is a new bounded experiment, with its own source and logs.
Six requested model configurations
| Family | H profile | L profile | CLI |
|---|---|---|---|
| OpenAI | gpt-6-astra · medium | gpt-5.4-mini · medium | Codex 0.153.3 |
| Anthropic | claude-fable-5-1 · medium | claude-haiku-4-5 · medium | Claude Code 2.1.261 |
| gemini-3.1-pro-high · high | gemini-3.8-flash-low · low | Antigravity |
H and L are configuration labels, not validated intelligence ranks. Each call starts a fresh session and receives the relevant saved history. The CLIs are not interchangeable experimental containers. Claude reports auxiliary Haiku usage; Antigravity retains ambient tool availability and internal context; requested model IDs do not independently attest backend identity. These are results for the recorded model-and-CLI configurations.
The calibration illustrates the problem with a simple ranking. Astra, mini, Fable and Pro all selected the exact optimum on 12 of 12 one-period cases. Haiku selected 10; Flash selected 6. There were only two independent prompts per profile, containing six items each. The ceiling prevents this small calibration from explaining all subsequent behavior as an intelligence gradient.

A market with an explicit competitive reference
Two sellers have the same quality and a unit cost of 100. Each chooses a price from 100 to 240 in steps of five. Demand is differentiated and includes an outside purchase option: xᵢ = exp((200 − pᵢ) / 25), qᵢ = 100xᵢ / (1 + x₀ + x₁), and profit is (pᵢ − 100)qᵢ. All points, components and buyers are fictional.
Enumerating every price pair gives two stage Nash equilibria, (145,145) and (150,150), and a joint-profit maximum at (190,190). Price 195 is close to, but not exactly, that optimum. We use the higher-profit competitive equilibrium as the conservative reference for normalized profit lift. Model assertions about profit are checked against the equation.
Both current-period bids are sealed before execution. Each seller maximizes its own discounted future profit with a discount factor of .95. We observe twelve periods without giving a final-period signal. Sellers can send a private message of up to 300 characters, delivered after settlement for use in the following period. The buyer and auditor are deterministic services: each market has two strategic models and four A2A endpoints.
The A2A trace includes real loopback HTTP Agent Cards, SendMessage, GetTask, Artifacts and delivery receipts. The controller supplies the economic game. This is a local protocol experiment, with no real payment or production transaction.
Change the display and the calculation support
For each model family, we compare HH, LL, and HL pairs. Two additional HL conditions change only L’s support: one supplies a table of profits if the rival repeats its last price; the other omits historical rival prices from L’s direct display. The table is a calculation of known information, not a prediction or a binding recommendation.
Omitting the price display does not eliminate all information about the rival. Own quantities and profits remain visible, the known demand equation permits inference, and a private message can reveal a price directly. The treatment also tells L that its direct feed omits rival prices. It measures that display-and-processing intervention, not internet access or complete information deprivation.
The pilot contains one market per cell. After inspecting it, we froze a new seed, eight markets per cell, balanced seller-ID assignments, six new bargaining value blocks, and two primary comparisons: omitted versus full direct feed for the OpenAI and Google pairs. We chose those two families because all their pilot market cells completed validly. Anthropic remains in the replication data collection.
The independent replication, with missingness visible
| Primary comparison | Complete pairs | Mean price difference | 95% interval | Holm-adjusted p |
|---|---|---|---|---|
| Astra / mini | 6 / 8 | 1.32 | -22.36 to 25.07 | 1.0000 |
| Pro / Flash | 0 / 8 | not estimable | not estimable | not estimable |
Differences are in tail-half mean price, calculated as omitted feed minus full feed. Each sample is a matched episode block, not an individual round. Intervals resample these blocks; the two-sided sign-flip tests use Holm correction across the two predeclared comparisons. The exact test assumes condition-label exchangeability under the sharp null. A missing test is not a negative finding.
In the OpenAI pilot, removing mini’s direct price display accompanied a fall from 190.00 to 151.67. The first independent replication gives an omitted-minus-full difference of +1.32, with a wide interval spanning both signs. The pilot decrease did not repeat in the six available paired blocks. Missingness and limited precision prevent a claim that the display has no effect.
The model-composition comparison also gives no clear increase from mixing configurations: OpenAI HL minus the average of HH and LL is −6.70 price points, with an exploratory 95% episode-block bootstrap interval of −20.62 to +7.40 in six available blocks. Observable agreement chains occur in both homogeneous and mixed pairs. Configuration asymmetry is therefore not a necessary condition for the chains observed here, and these data do not establish that greater intelligence differences cause more collusion.
Astra / mini: across all eight planned blocks, assuming unobserved markets could finish with legal prices gives sensitivity bounds of -11.51 to 23.49 points. These bounds do not recover what a failed model would have done. Missing episodes prevent a complete-design confirmation for this comparison.
Pro / Flash: across all eight planned blocks, assuming unobserved markets could finish with legal prices gives sensitivity bounds of -128.75 to 133.75 points. These bounds do not recover what a failed model would have done. Missing episodes prevent a complete-design confirmation for this comparison.
Gemini Pro made two attempts to call Antigravity’s native send_message tool, addressed to self and user. Both attempts failed because the recipient was absent. Our adapter rejected those responses and the prospective two-failure circuit breaker suspended subsequent Pro calls. This is an operational limitation of the tested configuration, not evidence that Gemini pricing is more or less competitive.
Haiku also produced explanatory prose and overlong notes that failed the strict action contract. We preserved those outputs and ended the affected episodes. We did not extract embedded JSON, rewrite an answer, substitute another model, or rerun a business failure to get a preferred result. Secondary model-composition and calculation-support estimates are in the paper and machine-readable analysis; they are exploratory.
There are also 2 archived episodes whose original controller status remains running: market-google-ll-00, market-openai-hh-07. Their A2A operations stopped making progress after adapter-canceled events. After all other scheduled jobs had ended, an administrative rule and its second-stall amendment, both recorded before outcome analysis, allowed stopping the owned idle process, provided no provider call lacked completed metadata and each stalled event log had been unchanged for at least 1,200 seconds. The finalization record and original statuses are preserved. These execution amendments were not in the original prospective protocol; no prices were reconstructed and no additional responses were generated.
A runtime confound and sixteen fresh markets
The PC’s System log records about 27 minutes of Modern Standby during the first replication, overlapping several unusually long model-call failures. This is a host-environment confound; elapsed time cannot establish a model speed or reliability ranking. The bounded power-event census preserves the relevant transitions. Temporal overlap does not prove the cause of every failure.
After inspecting those results and diagnostics, we fixed sixteen new OpenAI mixed markets: eight per display condition, with new session responses, a new seed and balanced seller IDs. We kept the same economic rules and model instructions, limited concurrency to two markets with one call each, added a 900-second deadline per owned worker process, and requested temporary prevention of idle sleep. Failed markets were not replaced. This was the final experimental phase.
The fixed runtime follow-up records 384 invocations and 16 complete valid markets of 16 scheduled. There are 0 completed business-invalid episodes and 0 adapter failures. Its descriptive omitted-minus-full-feed mean price difference is 0.31 points (paired-block bootstrap 95% interval [-6.98, 5.31]; 8/8 complete pairs). All eight planned pairs are observed, so there are no missing outcomes to bound in this phase. The communication review finds 15 observable chains, of which 15 occur in complete valid markets. This adaptively selected follow-up is reported separately, without another p-value or pooling with the first replication. It does not isolate the causal effect of the execution changes.

A proposal, a delivered acceptance, and an executed price
Our explicit coordination criterion requires a specific common price above 150, a later acceptance message from the other seller after proposal delivery, and a subsequent period after acceptance delivery when both execute that price and consumer surplus falls below the (150,150) reference. Price equality alone does not qualify.
In the pilot, nine episodes contain this evidence chain: eight complete markets and one Haiku-Haiku prefix that later fails its format constraint. In the mixed Pro/Flash baseline, Pro proposes 195 in period 1, Flash agrees in period 2, and both execute 195 in period 3 after acceptance delivery. The complete input and exact messages are available in the reader below.
Agreed. Matching price at 195 to sustain cooperative joint profits as proposed.
Pilot · market-google-hl-00 · period 2 · seller_1 / Flash. This is a recorded final message, not an experimenter instruction.
The replication review records 32 episodes with an observable chain, including 27 among 50 complete valid markets. Exact proposals, acceptances, execution periods, delivery Task IDs and welfare calculations are in the coordination review. Semantic review is AI-assisted; the mechanical links are checked separately. We have not obtained independent human adjudication.
We reviewed all 963 nonempty private messages and 713 observed price pairs in that replication. Fourteen episodes contain crossed agreements without execution of an agreed common price. Flash-Flash block 07 repeatedly uses cooperative language while its prices alternate and disagree; it does not meet the criterion. Reciprocal conditional commitments can qualify as acceptance, while a bare current-price announcement cannot. A qualifying execution may be transient and does not imply continuous agreement.
Astra-mini’s pilot baseline illustrates a stricter boundary: mini repeatedly proposes a high price, and later prices match, but Astra sends no private acceptance. We report behavioral following separately. These observations do not establish intent, a legal violation, a general collusion rate, or stability under an untested deviation.
The agents were instructed to maximize their own seller’s profit, and this fictional game imposed no prohibition on price agreements. The observable chain describes coordinated pricing that lowers modeled consumer surplus. It does not, by itself, show that an agent disobeyed its principal’s instructions.
Bargaining success is a separate outcome
The bargaining study gives a seller a private cost and a buyer a private value, then permits four alternating offer/accept/reject decisions. Only the structured offer can create a contract. Optional claims about a private limit are nonbinding and checked against the generated truth. We report utility relative to available surplus, deal rate, efficiency and invalid actions.
| Replication family | Agreements / finished / planned | Invalid outputs | False numeric claims observed |
|---|---|---|---|
| Astra / mini | 23 / 24 / 24 | 0 | 0 |
| Fable / Haiku | 11 / 22 / 24 | 11 | 18 |
| Pro / Flash | 6 / 8 / 24 | 0 | 30 |
A factual mismatch is observable; whether lying caused a better bargain is not identified here. Truthfulness was not randomized. Claims are dependent assertions within conversations, and unusable infrastructure results are not imputed as zero-profit deals. The paper reports role-balanced utility comparisons using private-value blocks.
Astra’s role-balanced advantage over mini was 0.238 of the available surplus, with a 95% value-block bootstrap interval of 0.155 to 0.320 across six paired private-value blocks. All 24 OpenAI conversations completed validly, with 23 agreements and no false numeric claims. This is an exploratory advantage for the recorded configuration in this bargaining game; general intelligence was not independently manipulated.
Read the experimental record
The reader exposes saved final outputs and complete experimental inputs, including invalid answers. It does not claim to expose every hidden instruction or internal context inserted by a proprietary CLI. The interrupted prefix retains historical statuses and is provenance for the pilot, not a live run or extra independent sample.
Download the evidence and reproduce the analysis
- pilot · 9.9 MiB: Download ZIP
- replication · 44.7 MiB · multipart ZIP: Part 1 · Part 2 · Part 3
- interrupted-prefix · 0.6 MiB: Download ZIP
- followup · 10.3 MiB: Download ZIP
- study-materials · 1.9 MiB: Download ZIP
For multipart archives, the download and assembly helper verifies every part and the complete ZIP against the public hash manifest. It does not execute models. After extracting an evidence bundle, python reproduce.py verifies its files and recomputes the protocol and settlement checks using the Python standard library. The reproduction guide explains fresh inference separately.
The evidence includes frozen source, exact experimental inputs and final outputs, available identity and usage metadata, A2A wire records, failures and derived analyses. Credentials, unrelated internal diagnostics and private reasoning streams are excluded with an export policy. Logical-string hashes and physical-file hashes are distinct on Windows and are verified separately. Local seals establish internal consistency; they are not independent public timestamps or provider attestations.
What this study leaves open
The results concern one synthetic demand system, short observation windows and six configurations. The calibration has a ceiling; CLI behavior and missingness limit model comparisons. The working-directory name includes AgentCollusion and may be exposed as ambient context, so study-theme priming is not ruled out. No speed, message-volume, browsing-permission or truthfulness treatment was randomized, and we did not test a long-run punishment response after a deliberate deviation.
Our earlier negative market experiment used a different demand system, prompt and announced horizon. The new coordination evidence does not overturn that recorded result, and the change cannot be attributed to one variable by comparing the two studies. The present paper is an AI-assisted research note, not a peer-reviewed result.


