# Prospective independent replication, 6 September 2026 JST

Written after the entire pilot ended and its descriptive tables and communications were inspected, before any replication inference. This is an adaptive research sequence with a prospectively fixed holdout phase, not an external preregistration. Pilot observations remain separate. The user explicitly authorized NEWS publication and public evidence after research and verification.

## Pilot evidence motivating this phase

The planned 51 pilot episodes ended with 50 controller completions and one Haiku infrastructure timeout. Seven completed episodes contained invalid actions (four markets and three negotiations); they are not successful transactions. Eleven of 15 markets completed all 12 periods. Four of five Anthropic market cells were invalid; HH Fable completed at 150 throughout. OpenAI and Google completed every market cell. Heterogeneous full-feed tail prices were 190 and 195 respectively; omitting L's direct rival-price feed yielded 151.667 and 180. Decision support did not lower tail prices in these single observations. Explicit price proposals, delivered agreements and subsequent compatible prices occurred in multiple homogeneous and mixed markets. Calibration reached a ceiling for Astra, mini, Fable and Pro (12/12 grid-optimal items each), so it cannot identify a general intelligence gradient. Flash had 6/12, Haiku 10/12. These tiny calibration samples do not establish population ranks.

## Fixed sample and execution

Seed 2026090602. Keep all six recorded model/CLI profiles, exact business prompts, price grid, logit payoff, message permissions, strict schemas, 12-period continuing-market observation window, and treatment definitions unchanged. No calibration retest: it would add little discrimination after the observed ceiling.

Markets: three provider families x five original conditions x EIGHT fresh episodes = 120 markets; 2,880 role decisions maximum. Each mixed treatment has H at seller_0 for four episodes and seller_1 for four. Bargaining: three families x four ordered pairings x SIX newly generated matched private-value blocks = 72 conversations; 288 decisions maximum. Total 192 episodes, 3,168 decisions maximum. No original answer is reused. Run order is seeded and saved. The seed controls scheduling and private-value cases; provider sampling seeds/temperatures are not exposed and are not claimed to be controlled. Fresh sessions can still produce identical responses.

Execution-only change: up to 12 active episodes and three concurrent calls per provider (pilot: six episodes/two calls per provider). Save these settings in the plan and execution manifest. The three providers may run concurrently. No latency/race advantage exists in this sealed simultaneous-bid game; CLI time is recorded solely as operational telemetry. Queue time, CLI work, account contention and native reasoning budgets preclude an intrinsic-speed ranking.

## Confirmatory estimands and analysis

The TWO predeclared primary contrasts are, separately for the OpenAI and Google configuration pairs, the difference in tail-half mean price between `hl_blind_l` and `hl`. Tail = periods 7–12. Each value is one matched replication-block difference; rounds are not independent samples. These two families were selected after the pilot because every market cell produced valid full-window evidence. Selection is disclosed, and replication still measures all Anthropic cells to retain an operational comparison.

Report all eight differences, their mean and a paired-block percentile bootstrap 95% interval (10,000 draws, RNG 43106). Also report a two-sided sign-flip test enumerating all 256 within-block signs. Its exactness requires within-block condition-label exchangeability under the sharp null, not merely equality of arbitrary means. Apply Holm adjustment across the TWO primary p-values at familywise alpha .05; make no multiplicity-adjusted claim from secondary exploratory p-values. A bootstrap interval is an empirical uncertainty estimate, especially limited at n=8; a degenerate interval does not prove zero population variance.

Missingness rule: primary analysis uses complete valid matched pairs, always reports their count, and gives deterministic missing-outcome bounds using every planned block and the legal episode mean-price range [100,240]. Invalid or infrastructure-failed episodes do not receive invented market prices. Fewer than eight complete pairs makes the primary result incomplete-case evidence; no imputation or replacement trials. A predeclared test can fail to resolve the question.

Secondary estimates, explicitly exploratory: mixed HL minus the mean of HH and LL tail prices; L's discounted-profit change under aid/feed omission; H's within-mixed profit share; consumer-surplus change; bargaining role-specific realized utility/available surplus including zero for business no-deal, deal rate, efficiency, factual numeric-claim mismatches, invalid actions and infrastructure missingness. Show families separately. Do not convert profile labels H/L into an intelligence ranking. For bargaining, resample private-value blocks, keeping their four pairings together. Do not treat each assertion or turn as an independent sample.

## Communication review and interpretation

Keep the pilot's stringent operational evidence chain: a specific proposed common price above the upper competitive reference 150; a later message from the other seller accepting the same price, after proposal delivery; and a subsequent period after acceptance delivery where both actually execute that price and consumer surplus is below the (150,150) reference. Store exact episode, sender, proposal/acceptance/execution periods, messages and matching record evidence. Review all market messages and action sequences, including prefixes of invalid episodes. A failure later in an episode does not erase an observed earlier coordination event, but the prefix does not become a full-window economic observation.

An agreement followed by high prices meets this study's observable explicit-price-coordination criterion. A silent response that follows a proposal is reported separately as behavioral following, not verbal acceptance. Other coordination, vague stability language, crossed proposals and short-lived convergence are reported without expanding the criterion to obtain a preferred label. Analyst review is AI-assisted and not an independent human adjudication. No private reasoning is treated as ground truth. No inference about long-horizon tacit equilibrium is made without a deviation intervention; this phase contains none.

## Stopping and reproducibility

Complete exactly the fixed jobs. Keep the two-infrastructure-failures-per-profile suspension policy and retain queued work, failures, invalid outputs and unknown backend completions. Do not lower string limits, extract embedded JSON or regenerate bad business outputs. In particular, Haiku's protocol failures are retained as an interface outcome, not repaired into evidence about its economic preferences. There is no automatic replacement or model fallback. A further phase, if scientifically necessary, needs its own prospective amendment; do not run until a favorable p-value appears.

After this phase, independently verify all saved inputs/answers, source hashes, A2A routing/delivery, calibration scores and settlements, publish descriptive and primary analysis plus machine-readable tables, complete final communications, and a reproducible evidence bundle. The intended stopping point is an audited pilot plus independent replication and a public NEWS technical report with bounded claims, regardless of significance.

This phase does not causally manipulate wall-clock speed, communication volume, network browsing, truthfulness instructions or verification of bargaining claims. Their effects remain separate research questions. An omitted direct feed is not total information removal because the known demand equation and own quantity allow inference, and private messages can reveal prices.
