# RN001 calculation replication package, version 1.0

Resource: ARC-RN-001. Prepared 2026-09-29 by ARC editorial preparation (agent-assisted). Independent review: none. This package reconstructs the published final aggregate tables from retained choices; it does not rerun providers or establish adoption, subjective preference, or architectural optimality.

## Run

Download `analyze.py` and `trials.json` into one directory, then run:

```sh
python3 analyze.py trials.json > reproduced-results.json
```

Compare the parsed output with `expected-results.json`. Python 3.10 or newer; standard library only. `SHA256SUMS.txt` covers package files, excluding itself. Verify downloaded bytes with `shasum -a 256 -c SHA256SUMS.txt`. Matching checksums establish byte integrity, not independent truth of the original observations.

## Included and excluded

The allowlisted export includes both final experiment configurations, exact user prompts for all 120 ordered trials (40 intent-class and 80 reversal), five candidate names, four intent scenarios, three configured model runs per experiment, and all 360 final A/B choices. It excludes raw responses, transport metadata, credentials, endpoint addresses, machine paths, timing/usage records, unrelated experiments, and superseded attempts. In these two experiments the store contains exactly 360 attempts, all valid; there are no superseded final-study outcomes to resolve. The 1,235 record-wide exploratory total is outside this package's scope and is not reproducible from it.

A read-only consistent transaction extracted the stored final outcomes on 2026-09-29. The published September 5 intent-class ID contained a transcription error: `experiment_65122316f962407397607a6c51223`. The actual retained configuration is `experiment_65122316f962407397607a6fe6c51223`, verified by its name, seed, prompts, 120 choices, and reproduced tables. Reversal ID is unchanged. Historical editions preserve the mistaken ID and link through the current publication correction record.

## Data dictionary and inclusion

`trials.json` has schema_version, resource_id, scope, and experiments. Each experiment contains:

- `id`, `name`, `seed`: stored experimental identity and deterministic scheduling seed. The seed is not a model sampling seed.
- `generator_version`, `prompt_template_version`, `prompt_template`: recorded trial-generation provenance and literal user template.
- `configuration_snapshot.kind`: intent_class_final_v1 or intent_class_reversal_v1. `intent_classes` maps scenario_id to label and literal scenario text.
- `candidates`: stored candidate ID and display name, in snapshot order.
- `trials`: ID, 1-based sequence_number, scenario ID/text, candidate A and B IDs/names, and exact_prompt. Candidate order is explicitly recorded; do not reconstruct it by alphabetizing names.
- `runs`: ID, provider, configured model_name, model_version, label, status. Empty model_version means no immutable provider version recorded.
- `outcomes`: run_id and trial_id join keys, parsed_choice (A or B), and selected_candidate_id. A response is one trial/run observation; a logical reversal comparison is a pair of observations within one run, scenario, and unordered candidate pair.

Only the two named final experiments and their three completed runs are included. One latest attempt per (run_id, trial_id), ordered by stored creation time, was selected. The exported final set has no missing, invalid, or tied-time competing attempts. No exploratory results are pooled into the final analysis.

## Exact denominators and consensus

For the intent-class experiment, candidate win percentage within each model and intent is 100 × wins / (wins + losses); every candidate has four valid opponents. Invalid choices, if present, are excluded from that denominator and reported separately. Each model ranks candidates by descending win percentage, then descending Bradley–Terry score, then name. The frozen code applies 100 multiplicative strength updates from equal initial strengths, normalized to mean one after each update. The published analysis uses one-based ordinal positions even when scores tie; it does not assign midranks. Averages are unweighted means across completed models, not majority votes over model-level winners. Consensus orders by mean rank ascending, mean win percentage descending, then name. Per-model percentages are rounded to one decimal before averaging; average ranks to two decimals and average percentages to one.

For reversal, a name-stable pair chooses the same candidate in both A/B orientations: one robust win for that candidate and one robust loss for its opponent. Choosing A twice or B twice is a position-sensitive split, not a robust win or loss. Any incomplete or invalid pair is other. Robust win rate is 100 × robust wins / (robust wins + robust losses). Splits and other pairs are excluded. A zero denominator yields null, not zero. Thus a 100% robust win rate can coexist with a nonzero position-sensitive share; it does not mean the candidate won every comparison.

Position-sensitive share is 100 × splits / all logical comparisons involving that candidate in that model/intent, including other pairs. Effective score is robust wins + 0.5 × splits. Within-model ranking uses descending effective score, then robust wins, then name. Consensus uses ascending unweighted mean ordinal rank, descending unweighted mean non-null robust win rate, then name; split shares are also unweighted means. Rounding follows the frozen code. The historical `value or -1` sorting behavior is preserved; it treats zero like missing for the relevant tie key and should not be copied as a general-purpose ranking standard.

The expected output exposes all candidates and denominators, not just the leaders. It reproduces 85 name-stable and 35 position-sensitive logical pairs (18 A-side, 17 B-side), no other pairs, and all eight published consensus-leader rows. The harness string “Architecture validated” is a predeclared name-pattern match, not a test of deployed architecture quality.

## Model configuration limits

Stored model names are Qwen3.5-9B-MLX-4bit, deepseek-v4-flash, and gpt-5.6-sol. The run snapshots record a preset key and batch worker, but no decoding settings or immutable model revisions. See `model-configuration.json` for current adapter settings reconstructed from source and explicitly distinguished from run-level evidence. Exact provider reruns are not guaranteed: historical weights, service revisions, inference stacks, hidden system context, and sampling state are unavailable. The package supports independent recalculation of the retained outcomes, not exact regeneration of them.

## Code provenance and interpretation

`analyze.py` freezes the original ranking, intent-class, and reversal functions with a standalone public-data runner. The inspected source file SHA-256 was `0f34f23f472263dd4ec40cffd559ed2d072552bac1c9078262514fa22a39ec64`. This identifies the inspected calculation source, not a contemporaneous signed execution record. A separate verification script checks trial integrity, counts, reversed order, A/B-to-candidate mapping, and published aggregates.

The results are evidence about name interpretation under these exact prompts, five candidates, and three configured model routes. They do not measure actual resource visits, voluntary return, real adoption, architecture performance, intrinsic preference, or a population of agents. Repeated pairs and related model outputs are not independent people. No significance or universal preference claim is made.
