How evidence is represented can shape what survives.

The dossier surveys methods and findings relevant to the intake, representation, retrieval, and transmission of evidence by language models. It argues that there is no format that is universally best: apparent fidelity depends on the receiving model, task, context size and position, token budget, and the evidence properties one wants to preserve. Existing work offers useful foundations in long-context retrieval, factual consistency, and citation evaluation, while leaving important gaps around provenance, epistemic status, uncertainty, scope, relations, and multi-step transfer. The report proposes a controlled benchmark using canonical evidence graphs, several renderings, multiple models and budgets, and a multidimensional scoring scheme. This is a methods dossier and research synthesis—not a report of a newly conducted controlled experiment.

Representation quality is conditional, not a single ranking.

Performance depends on the interaction among the representation, the model receiving it, the question being asked, how much evidence fits, and where that evidence appears in context. A compact encoding can help when it preserves the distinctions needed for the current query and reduces retrieval burden. But a short packet may silently discard details needed for an unknown later question. Conversely, a verbose or highly structured format is not automatically clearer: it can consume context, add parsing overhead, or interact poorly with a particular model.

Accordingly, “best format” should not be treated as a property of a serialization in isolation. Comparisons need to specify the model, task, budget, position, and scoring target. Token efficiency and evidence fidelity should be reported together rather than collapsed into one unqualified score.

Measure what the receiving system can recover—and what it changes.

The surveyed literature provides methods for testing long-context retrieval, factual consistency, attribution, and citation behavior. These are valuable foundations, but they do not by themselves establish that a system preserves the structure of evidence. Correctly repeating a proposition is different from preserving who asserted it, whether it was observed or inferred, its scope, confidence, caveats, or relation to other claims.

The dossier identifies comparatively thin direct evidence for fidelity to epistemic status, provenance edges, uncertainty, scope, contradiction, and relationships. It also highlights an open question about recursive handoffs: when one model summarizes or reformats evidence for another, which distinctions survive each transfer, and can later stages recover the original resolution? The report treats these as research questions, not settled findings. Evidence about a model’s performance in one setting should not be generalized to all formats or agents.

Use a hidden evidence graph and test several dimensions independently.

The proposed benchmark begins with a canonical, hidden evidence graph containing claims, sources, relations, uncertainty, scope, and epistemic labels. It renders equivalent evidence in multiple formats, including controlled isomorphic forms and native, naturally authored forms. Tests vary the receiving model, task, evidence budget, context length, and the position of relevant material. Multiple packets and model configurations are needed; many questions over a single packet must not be mistaken for many independent evidence samples.

Evaluation should separate factual recovery from source attribution, relation recovery, epistemic classification, uncertainty and scope preservation, omissions, unsupported inferences, contradiction handling, and survival across recursive transmission. Closed-book baselines and counterfactual twins can help detect contamination or answers supplied from prior model knowledge instead of the evidence packet. Token counts should use the relevant tokenizer and be reported alongside quality measures; comparisons should expose trade-offs or Pareto frontiers rather than declare a single universal winner. The proposal recommends deterministic scoring where possible, blinded diverse judgments for harder criteria, human audits, paired item-level comparisons, and explicit controls against pseudo-replication.

Treat the dossier as a measurement agenda.

The most actionable contribution is a design for asking better empirical questions—not a ready-made prescription that ARC or other systems should adopt one serialization. Before choosing a format, define what a recipient must be able to recover, what may be compressed, and what must remain attached to its source and uncertainty. Then compare candidate representations under the intended receiver and task, including constrained budgets and multi-stage transfer if those matter in practice.

For ARC, this suggests preserving source attribution, limitations, and claims about evidence strength when presenting the report, while keeping the portal’s abstract and brief distinct from the original full text. Any future benchmark based on these recommendations should publish its materials, tokenizer and model configurations, scoring rubric, exclusions, and uncertainty. Results should be described as bounded to the tested conditions.

Scope and provenance matter.

This publication reproduces a supplied OpenAI Deep Research dossier and provides an ARC-authored navigation and summary layer. It does not independently validate every source or citation marker embedded in that file, and the report is not a newly run controlled experiment. The proposal’s metrics and hypotheses require operational definitions, reliability checks, and replication before strong conclusions follow. The token counts shown here are rough word-count-based estimates, not counts from the resolved model’s tokenizer; use the file itself for exact bytes and words.

Original supplied Markdown

Open or download the complete report — approximately 20,000 tokens (estimate). The Markdown is served as a separate text resource and is unchanged from the supplied file.

Exact citation label: OpenAI. (2026). Machine Evidence Intake and Representation Fidelity: Methods Dossier. Generated by OpenAI Deep Research. Resolved model: gpt-5-thinking; model reasoning generation/version metadata: v5. Research completed 24 September 2026.