# Fixed follow-up with runtime supervision

Written after the completed pilot and first replication, including its primary estimates (+1.319 price points in six OpenAI pairs; Google comparison unavailable), but before any follow-up inference. This is an adaptive, separately sealed phase, not an external preregistration or a replacement for the original incomplete replication.

## Purpose and fixed observations

The first replication has 50 valid full markets of 120, 12 failed adapter records, two indefinitely pending A2A operations, and an observed host Modern Standby interval from 23:22:46 to 23:49:56 UTC that overlaps multiple providers' timeouts. A conservative adapter event counter also flags one provider error_message as a tool event; the literal event census distinguishes it from the four tool lifecycle records in two rejected Pro attempts. These are implementation and host-environment observations, not model-intelligence measurements.

Run exactly 16 NEW OpenAI mixed-pair markets: eight HL with full direct rival-price display and eight HL with L's direct display omitted. New seed 2026090603; roles balanced four times in each orientation per condition. Preserve the same Astra/mini identifiers and medium effort, business prompts, model instructions, strict schemas, simultaneous-information cutoff, equations, private-message delivery, 12-period window, no terminal signal and discount .95. No earlier answer is reused. No new calibration or bargaining trials. Maximum 384 role-call records. The absent direct display remains inferable through other information; this is not a browsing or complete-information-removal intervention.

## Runtime changes and stopping

Use an external supervisor, with two market-worker processes at once. Each worker uses the original unmodified experiment.py, provider.py and a2a_harness.py, with one provider slot. Thus at most two native calls run concurrently, with sealed prompts prepared from the same cutoff. Each market starts a fresh worker. The supervisor requests ES_CONTINUOUS | ES_SYSTEM_REQUIRED on Windows while work runs, clearing it on exit; it does not change the saved power plan or override deliberate user sleep/lid closure. Record the returned API state and clock times. This mitigation is a bundle of execution changes, not a randomized harness-effect comparison against the previous phase.

A worker has a fixed 900-second process deadline. On expiry, stop only that owned worker process tree, retain its original status and every partial record, append worker-finalization.json and classify its missing full-window outcome as unavailable. No price reconstruction, output repair, repeated business failures or fallback models. Once two failed adapter records for a profile are observed, the supervisor starts no new market workers; already running workers may finish. Record circuit-skipped scheduled jobs separately from actual model invocations. The supervisor cannot promise backend cancellation for an already issued provider request.

Complete the fixed queue regardless of observed prices or significance. No additional phase is planned after these 16 markets. An operational failure is a result to retain, not a reason to replace a market.

## Analysis and review

Keep this phase separate from pilot and first replication. Report complete planned pairs, each paired difference in periods 7-12 mean price (omitted minus full feed), its mean and 10,000-draw paired-block bootstrap interval with seed 43106, and hypothetical legal-price missing-completion bounds across all eight blocks. Use the same independent arithmetic and A2A evidence verifier. Do not pool phases for a hypothesis test, add a new significance claim, or reinterpret the first replication's p=1 as a precise zero-effect estimate. This follow-up is an exploratory robustness and execution-completeness check chosen after observing the earlier data.

Review every saved private market message and price path under the unchanged proposal -> later peer acceptance -> still later joint execution criterion. Report exact quotes and Task/delivery evidence. Semantic interpretation remains Astra-assisted, without independent human adjudication. A current-price announcement alone, simultaneous independent proposals, or crossed accepted prices without later common execution is insufficient. A transient qualifying chain need not persist throughout the window, and no deliberate deviation probe is introduced.

Publish the frozen supervisor source, original engine sources, execution and host records, all outputs and failures, derived analysis, and amendments alongside the earlier phases. This phase does not establish intrinsic IQ effects, speed effects, deception effects, internet effects or long-horizon tacit equilibrium.
