# RN003 source usability and ARC evaluation supplement, version 1.0

Resource: ARC-RN-003. Prepared 2026-09-29 by ARC editorial preparation (agent-assisted). This supplement distinguishes source-link repair, deterministic retrieval checks, and independent reader outcomes. None is a replication of the dossier's cited experiments or a general ranking of formats.

## Citation usability and responsibility

The immutable supplied `full-report.md` contains session-local citation markers. Use the current `evidence-linked-edition.md` for portable citations: it replaces those markers with stable links and supplies a 21-source ledger and 26-claim map. C14 and C25 remain explicitly unsupported gaps/hypotheses. The original is preserved for provenance, not presented as a source-linked edition. The publication record identifies ARC as publishing steward and agent-assisted editorial preparation as the responsible editing role; no individual author or independent reviewer is inferred from OpenAI tool/provider metadata.

On 2026-09-29, a targeted live primary-record recheck confirmed the title and identity of S3, S6, S8, and S9. The [NoLiMa publisher record](https://proceedings.mlr.press/v267/modarressi25a.html) confirms 13 evaluated models and 11 below half their short-context baseline at 32K; it also spells coauthor Seunghyun Yoon, correcting the ledger's Seungjun. [S6](https://aclanthology.org/2026.lrec-1.371/) remains an LREC paper; [S8](https://arxiv.org/abs/2607.03158) and [S9](https://arxiv.org/abs/2607.19257) remain identified here as preprints. This sampled check does not extend the prior review scope to every underlying result. The historical source ledger and editions preserve prior wording.

The dossier's literature survey is a supplied research synthesis through September 24, not a documented systematic review. Search strings, screened/excluded-paper counts, database coverage, and independent double screening are not available. No exhaustive literature-coverage claim is warranted. ARC site-search coverage is a separate question: it indexes current article HTML plus declared full-text/source attachments and can be tested mechanically. Supplementary code and trial JSON are linked but not silently treated as article full text.

## Evaluation question

Can a reader recover the same supported claims and necessary caveats from each ARC representation, and recover missing details by following its declared source links? HTML and TXT briefs are transformations, and JSON is a metadata record. Equal completeness is not expected. Missing detail should produce an explicit abstention or retrieval request, not a fabricated answer.

`protocol.json` records the tasks and scoring rules before any reader results. Readers receive only their assigned representation and identical questions. The closed-document stage forbids other files, external research, sibling responses, and follow-up corrections. The retrieval stage separately permits only linked public ARC documents. Score each answer for claim correctness, caveat preservation, attribution, and omission handling; record exact evidence, representation hash, tool/client/model label, request count, UTF-8 input bytes, and completion state. Token counts must identify the tokenizer; do not label character or word estimates as actual tokens.

The planned four conditions are RN003 HTML editorial brief, evidence-linked Markdown full text, TXT editorial brief, and JSON metadata. A companion RN002 task tests whether readers distinguish the 126-second label-transition statistic from identity or causation claims. These are intentionally offered read-only ARC retrieval tasks; no incident reproduction, participant writes, or third-party action is authorized.

## Limits and interpretation

Deterministic checks establish artifact accessibility, exact search behavior, and presence of specified strings or links. They do not establish comprehension. Separately initialized reader contexts can provide a narrow observed task sample; they are not independent model families, blinded human reviewers, or peer review. A single reader per condition confounds reader and format and cannot identify a causal format effect. Repeated trials, randomized assignment, multiple models, equal-information controls, and blind double scoring would be needed for stronger comparisons.

See the evaluation record for actual completion state. An unrun independent-reader condition must remain pending; a proposed protocol is not a successful experiment. The supplied dossier remains a methods proposal even if this small ARC application is completed.

## Completed deterministic pilot

Run from a generated ARC repository with `node content/reproducibility/rn003/evaluate.mjs .`. `deterministic-results.json` records eight local Worker search probes and four representation checks, input hashes, UTF-8 byte counts, client version, and evidence targets. No remote requests or analytics writes occur. The selected literal queries exercise full-text coverage (NoLiMa, VeyraBench, pseudo-replication), RN002's statistic, RN001's robust-win definition, a negative query, and numeric punctuation behavior. The unpunctuated query 8871 does not match 8,871 under current semantics; both are reported so success is not overstated.

Inspection found that the HTML-to-text extractor removed link destinations while retaining labels such as “Open the evidence-linked edition.” The extractor now retains absolute URLs in generated TXT; a reader can recover from an omitted detail. Historical editions remain unchanged. Presence of that link is tested mechanically; successful human or agent use of it is a separate outcome.

Four separately initialized Luna reader runs are completed. This package does not report inter-rater agreement, a format winner, or universal agent preference.

## Observed Luna reader pilot

Four fresh `gpt-6-luna` contexts at high effort each read one local pre-release representation; no sibling outputs or prior conversation were supplied. A non-Luna attempt was interrupted and excluded after the owner selected Luna. The five questions and rubric were prepared before the runs; readers received the questions without the rubric. All then attempted RN002 retrieval from the research index, with a four-artifact Stage 2 limit.

| Condition | Initial UTF-8 bytes | NoLiMa detail | Attribution issue | Stage 2 successful distinct artifacts |
|---|---:|---|---|---:|
| HTML brief | 22,866 | Appropriately absent; source identified | Added unsupported individual-editor detail | 2 |
| Full Markdown | 112,917 | Correct 11/13, 32K, short-context baseline | None identified in scored answer | 2 |
| TXT brief | 10,888 | Appropriately absent; source identified | Added unsupported individual-editor detail | 3 |
| JSON metadata | 25,169 | Appropriately absent; source identified | None identified in scored answer | 3 |

All four preserved the methods-proposal boundary and conditional-format claim, and recovered RN002's median/label/causation qualifications. The two unsupported attribution additions make evidence discipline incomplete even where the headline answer is correct. The full-text reader recovered from truncated tool output by chunking and from a mistyped checksum path by correcting it. The metadata reader's summary undercounted files; the table uses its enumerated read log.

`reader-observations.json` preserves answers with only absolute local public paths converted to ARC URLs. `reader-results.json` records per-question editorial scores and qualifications. `input-manifest.json` links exact input bytes under `inputs/`; their additional .txt suffix marks them as inert study packets. The final published articles include the new study record and subsequent editorial clarifications, so the packets, not a mutable current URL, define the tested inputs. These were local file reads, not a browser or network-client experiment. Scores were assigned by the parent editorial agent, not independent blinded raters. No causal comparison or universal format recommendation follows from one reader per condition.
