# Machine Evidence Intake and Representation Fidelity: Methods Dossier

This dossier addresses **inference-time intake of externally supplied evidence**: how a language model or language-model agent reads, retrieves, combines, qualifies, attributes, compresses, and reproduces information that appears in its current context. It does **not** treat this as parameter-level learning. The governing question is therefore not “which notation is most elegant?” but:

> **What representation permits a heterogeneous receiving model to reconstruct the evidentiary content, structure, provenance, scope, and uncertainty of a source corpus with the least loss for a given context budget?**

The literature was surveyed through **24 September 2026**. The strongest direct evidence is concentrated in long-context retrieval, prompt-format sensitivity, table serialization, RAG attribution, factuality evaluation, prompt compression, and LLM-based evaluation. Direct work on **epistemic-type preservation, repeated model-to-model evidence transmission, claim-ledger structures, and resolution monotonicity is much thinner**. That distinction matters throughout this dossier.

## Evidence landscape and terminology

**EXECUTIVE MAP**

The highest-confidence finding is negative: **there is no empirically defensible universal “best format.”** Semantically matched or closely controlled information can yield materially different outcomes when rendered differently, but the sign and magnitude of those effects depend on the receiving model, task, context length, information completeness, and sometimes information placement. Controlled studies comparing plain text, Markdown, JSON, YAML, XML, HTML, CSV, prose, tables, pseudocode, and code-like representations repeatedly find interactions rather than a stable global ranking. He et al. found format-dependent variation across GPT models, including large task-specific differences; a 2026 cross-family questionnaire study found as much as roughly nine percentage points from serialization choice; Kato and Kato's 2026 matched algorithm-specification experiment found rankings that changed with model and information completeness; and the 2026 VeyraBench preprint found outright ranking reversals between formats as context approached models' effective limits. citeturn7search0turn22view0turn22view1turn22view2

The second high-confidence finding is that **nominal context capacity is not effective evidence-processing capacity**. Lost-in-the-middle effects, task-complexity effects, and length-dependent degradation occur well below advertised maximum windows. Liu et al. found a characteristic beginning/end advantage over middle placement in multi-document QA and key-value retrieval; RULER found degradation across more demanding long-context tasks even where vanilla needle retrieval looked nearly perfect; NoLiMa removed lexical shortcuts and found severe deterioration by 32K context for most of the tested models. citeturn1search4turn8search1turn1search3

The third strong finding is that **retrieval is not comprehension**. A system can successfully locate a literal needle while failing when evidence must be associated indirectly, aggregated across documents, joined across tables, reconciled with other evidence, or used under changed wording. NoLiMa was explicitly constructed to remove literal matching shortcuts; RULER adds multi-hop tracing and aggregation; LongBench separates single-document QA, multi-document QA, summarization, few-shot tasks, synthetic tasks, and code; TQA-Bench finds that multi-table joins and aggregation create failures not predicted by simpler table access. citeturn1search3turn8search1turn9search0turn22view3

The fourth strong finding is that **representation cost is itself format-dependent**. Formatting tokens are not free. In VeyraBench's controlled corpus, Markdown, prose, and a Markdown table cost about 1.258×, 1.221×, and 1.367× the tokens of its plain-text rendering respectively; TQA-Bench's same-database serialization measurements found CSV markedly smaller than Markdown or JSON and HTML close to three times the CSV token cost at matched data scales. These numbers are study-specific rather than universal, because tokenization and rendering conventions differ, but they establish that raw task accuracy and context efficiency cannot be conflated. citeturn22view2turn22view3

The fifth high-confidence conclusion is methodological: **accuracy per token cannot be reduced to “compress as much as possible.”** LLMLingua, LongLLMLingua, and RECOMP demonstrate that removing irrelevant or low-utility context can sometimes *increase* QA performance while reducing tokens, because longer raw context also introduces distractors, placement problems, and attention competition. LongLLMLingua, for example, reported improvements on NaturalQuestions while using substantially fewer prompt tokens. That result is evidence for selective compression, not evidence that all information can be safely compressed or that compressed representations preserve uncertainty, provenance, exceptions, or non-query-salient evidence. citeturn16search0turn16search1turn16search2

A sixth conclusion is more cautionary: **the literature is much better at measuring whether a model recovers or generates facts than at measuring whether it preserves an evidence system.** FActScore atomizes generated text into factual propositions; RAGTruth annotates unsupported material in nearly 18,000 RAG responses; ALCE evaluates answers and their citations. Those are valuable foundations, but they do not by themselves measure whether a receiver preserved “observation versus inference,” conditionality, source disagreement, absence-of-evidence distinctions, or all typed dependencies between claims. citeturn15search1turn14search2turn14search0

A compact map of the present evidence is:

| Question | State of evidence | Dossier judgement |
|---|---|---|
| Does surface representation affect performance? | Multiple controlled studies | **Established, task/model conditional** |
| Is one serialization consistently best? | Repeated failures to obtain a universal ordering | **No** |
| Does context position matter? | Multiple independent long-context benchmarks | **Established** |
| Does format interact with context length? | Direct 2026 evidence plus earlier indirect evidence | **Probable to established; more replication desirable** |
| Does structure always beat prose? | Contradicted by controlled results | **No** |
| Can markup consume meaningful context budget? | Direct token measurements | **Established** |
| Can added structure itself hurt? | Direct table/specification results | **Yes, under some tasks/models** |
| Do explicit epistemic labels preserve epistemic distinctions? | Sparse direct controlled evidence | **Unknown / high-priority experiment** |
| Do claim-specific provenance mappings beat ordinary citations? | Attribution benchmarks support claim-level evaluation; direct format comparison is thin | **Plausible, not established** |
| Does recursive model-to-model summarization preserve an evidence graph? | Very limited direct work | **Unknown** |
| Is “resolution monotonicity” an established metric? | No clear standard found | **Novel proposed property** |
| Is semantic fidelity per token an established metric? | Compression metrics exist; no accepted multidimensional evidence-fidelity analogue | **Novel synthesis needed** |
| Do format effects generalize across model families? | Cross-family effects exist, but rankings vary | **Effects generalize; preferences do not** |

**TERMINOLOGY**

**In-context information intake** should mean the temporary use of information supplied in the inference context. A model may answer differently because evidence was placed in the prompt without any persistent weight update. This must be distinguished from fine-tuning, continued pretraining, reinforcement learning, or other parameter modifications.

**Representation effect** means a difference in receiver behaviour attributable to how materially equivalent information is serialized or organized. A rigorous representation experiment therefore needs some invariant latent content against which the renderings are matched.

**Surface format** should be distinguished from **information architecture**. JSON versus YAML is largely a serialization contrast. Narrative versus atomized propositions changes discourse structure. Adding epistemic status or provenance adds explicit information that might previously have been implicit. A benchmark that calls all three “format” would confound fundamentally different interventions.

**Effective context length** is the region over which a model can use evidence at the required task quality, rather than its API's maximum accepted sequence length. RULER and NoLiMa make the need for this distinction particularly clear. citeturn8search1turn1search3

**Lost in the middle** refers to the observed positional degradation in which relevant evidence located around the middle of a long context can be used less successfully than identical evidence near the beginning or end. The exact curve is model- and task-dependent, so the phrase should not be treated as a universal mechanistic law. citeturn1search4

**Evidence fidelity** should be broader than factuality. For this programme, it should mean preservation of the **claims plus their typed relationships, sources, epistemic statuses, scopes, numerical values, temporal qualifications, contradictions, limitations, and unsupported boundaries**.

**Atomic claim** follows the useful intuition behind FActScore: decompose compound statements into propositions that can be independently evaluated. For this project, however, atomization must be extended with typed edges because atomic-fact precision alone loses relationships. citeturn15search1

**Attribution / provenance fidelity** means correctly identifying which supplied source or sources license which proposition. ALCE's separation of answer quality and citation quality is an important precedent for treating attribution as an independent evaluation target. citeturn14search0

**Unsupported inference** means an output proposition that is not licensed by the supplied corpus under the benchmark's defined inference rules. RAGTruth demonstrates the feasibility of span-level/manual annotation of unsupported RAG material, while FActScore supplies a proposition-level decomposition paradigm. citeturn14search2turn15search1

**Resolution monotonicity**, retained as the provisional term, should mean: as a representation is allotted a larger budget, new information should preferentially **refine, qualify, or extend** the defensible lower-resolution representation rather than force retraction of claims that the shorter version asserted without support. No established metric in the reviewed literature captures exactly this property. Recursive/hierarchical summarization is adjacent, not equivalent. Wu et al.'s book work demonstrates recursive decomposition as a workable summarization procedure, but it did not establish monotonic evidence reconstruction across increasing budgets. citeturn21search0

## Representation, context, and epistemic structure

**REPRESENTATION EFFECTS**

A useful empirical distinction emerges between **surface-serialization sensitivity** and **semantic-organization sensitivity**.

He et al. rendered prompt material in plain text, Markdown, JSON, and YAML across reasoning, translation, and code-oriented tasks. They found substantial performance differences under semantically corresponding prompts; sensitivity was particularly strong for GPT-3.5 in some tasks, while GPT-4 was generally more robust. Their results do not establish a winner: they establish that formatting is a causal variable and that model capability alters its magnitude. citeturn7search0turn10search5

Nguyen et al.'s LREC 2026 QASU benchmark is important because it broadens this beyond an OpenAI-only comparison. It evaluates six structural skills, six serialization formats, and models including GPT-5-mini, Gemini 2.5 Flash, Qwen3-32B, Llama-3-70B, and Amazon Nova Lite. Format selection produced improvements of up to about nine percentage points over baseline formats, while model-family gaps remained large. This supports a benchmark design in which **format × model is an explicit interaction**, not nuisance variance to be averaged away. citeturn22view0

Kato and Kato provide an unusually clean 2026 matched design. Across **4,020 generated implementations**, five machine-learning tasks, seven representations, three models, and four experimental settings, the target function and hidden tests remained fixed while specifications appeared as ordinary prose, LaTeX-style pseudocode, PDF-like pseudocode, Markdown fields, YAML-like structures, JSON-like structures, or Python stubs. Under “core information,” LaTeX-style pseudocode had the strongest average effect, but prose and YAML-like representations were close in several comparisons. Under “complete information,” GPT-5.4 mini showed no matched format differences while Gemma 3 4B and Llama 3.2 3B still did. The important result is therefore **completeness can dominate syntax for one model while syntax continues to matter for others**. citeturn22view1

The same paper also demonstrates why “structured” cannot be equated with “complete.” Python stubs made the interface highly explicit yet did not consistently produce correct implementations because an interface does not substitute for numerical rules, boundary behaviour, tie-breaking, shapes, and computational steps. This is a direct analogue of the present research problem: **making a representation machine-legible does not guarantee that it carries the evidence needed for reconstruction.** citeturn22view1

Table serialization experiments reinforce the absence of a global ordering. Sui et al.'s Table Meets LLM benchmark draws tables from TabFact, FEVEROUS, SQA, HybridQA, and ToTTo and uses 1,500 tables per task in its main evaluation. Depending on task, natural-language-plus-separator, Markdown, JSON, XML, and HTML serializations changed downstream results; HTML was strong in some configurations, while Markdown or JSON could lead elsewhere. The same study found that generic extra structural material was not uniformly beneficial and that the placement of additional information relative to the table could reduce performance. citeturn11search12turn13view0

TQA-Bench extends this to multi-table reasoning. For the same underlying database, its token counts at a nominal 8K scale were approximately 3.73K tokens for CSV, 5.40K for Markdown, 5.75K for JSON, and 10.5K for HTML; analogous gaps persisted as scale increased. Yet CSV's compression advantage did not automatically make it the most accurate serialization: for Qwen2.5-7B-Instruct at the 8K setting, reported overall accuracy was 49.86 for Markdown versus 40.57 for CSV, 41.00 for JSON, and 46.57 for HTML. Efficiency and comprehension must therefore be evaluated separately before being combined. citeturn22view3

The 2026 VeyraBench preprint provides particularly direct evidence for a **format × context-length interaction**. It renders one deterministic synthetic corpus in Markdown, plain text, prose, and a table and evaluates five models across a context ladder reaching 512K where supported. Format spreads became much larger near effective context limits; the same format could become best at one model/length cell and worst at another. Its largest reported spread reached 48 percentage points at 128K in one model. Because this is a recent preprint, those exact magnitudes should be treated as provisional until independently replicated, but the controlled design makes the interaction a high-priority benchmark factor. citeturn22view2

Taken together, the evidence does **not** support any of the following shortcuts:

“JSON is machine-native, therefore models understand it best”; “Markdown is the default language of LLMs”; “tables are always denser and therefore superior”; “natural language is wasteful”; or “more explicit structure always helps.” The supported conclusion is narrower: **different representations expose different cues, impose different overheads, and fail differently.** citeturn13view0turn22view1turn22view2turn22view3

Direct controlled evidence is comparatively weak for **JSONL, textual triples, knowledge-graph serializations, and explicit claim/evidence graphs** when the requirement is strict semantic equivalence to prose and the dependent variable is downstream evidence reconstruction rather than graph extraction or KG QA. These therefore belong among benchmark candidates, not presumed solutions.

**CONTEXT AND POSITION EFFECTS**

Liu et al.'s Lost in the Middle experiment is foundational for experimental design because it changes **where** the relevant information appears while preserving its relevance. Across multi-document QA and synthetic key-value retrieval, performance was often highest when relevant evidence appeared near the beginning or end and lower when it was placed in the middle. citeturn1search4turn1search0

RULER shows that simple retrieval can overestimate long-context competence. Across 17 long-context models and tasks covering retrieval, multi-hop tracing, and aggregation, models that looked excellent on vanilla needle tests often degraded strongly as context length and task requirements increased; the authors reported that only about half of tested models maintained satisfactory performance at 32K despite claimed context capacities at least that large. citeturn8search1

NoLiMa attacks another loophole: literal overlap. Its questions require the model to associate a query with evidence through latent rather than exact lexical correspondence. Across 12 long-context models, most suffered severe relative deterioration by 32K; the paper reports GPT-4o falling from 99.3% of its short-context performance to 69.7% at 32K, while ten of twelve models fell below half their strong short-context baseline. A representation benchmark based only on source-string retrieval would therefore risk measuring string matching rather than evidence understanding. citeturn1search3

LongBench further demonstrates that long-context competence is multidimensional: single-document QA, multi-document QA, summarization, few-shot learning, synthetic tasks, and code completion yield different behaviour. Retrieval-based compression can improve weaker long-context systems without necessarily making them equivalent to models that handle the full context more effectively. citeturn9search0

For the eventual representation benchmark, **position must be crossed with representation rather than held in one arbitrary ordering**. At minimum, target evidence should occur in early, middle, late, and distributed conditions. For multi-source relationships, “distributed” means that the premises necessary to reconstruct a relation reside in separated regions rather than inside one convenient block. Evidence ordering should be randomized or Latin-square balanced across formats. The same canonical evidence item should appear at matched **relative context percentiles**, while absolute token locations should also be recorded because formats have different token overheads. These recommendations follow directly from position and token-cost effects documented by the long-context and serialization studies. citeturn1search4turn22view2turn22view3

Query-aware reordering should **not** be mixed into the primary format test. LongLLMLingua deliberately prioritizes/reorders information to improve long-context performance; that makes ordering an intervention in its own right. It is valuable as a later experimental factor, but including it selectively in some format conditions would make “format” inseparable from retrieval optimization. citeturn16search1

Redundancy likewise needs independent treatment. Duplicate evidence can increase the probability of retrieval without increasing the amount of distinct evidence. A representation that repeats conclusions three times could therefore score well on fact recall while consuming context and distorting apparent efficiency. Initial representation comparisons should equalize semantic multiplicity; redundancy can be introduced in a later factorial experiment.

**EPISTEMIC STRUCTURE**

The literature gives strong tools for **factual support**, but weaker evidence for **typed epistemic preservation**.

FActScore's atomic-fact decomposition is highly relevant to claim fidelity: rather than grading an entire response as factual or not, it decomposes long-form generation into atomic propositions and evaluates factual precision. The present benchmark should adopt that granularity while expanding the gold object from a set of facts into a typed evidence graph. citeturn15search1turn15search30

RAGTruth demonstrates that unsupported-generation errors can be manually localized in retrieval-grounded outputs. Its nearly 18,000 naturally generated RAG responses are annotated at response and word/span levels, supporting an explicit **unsupported inference rate** rather than subsuming hallucination into a single holistic score. citeturn14search2turn14search10

However, neither atomic factuality nor hallucination detection answers whether a model preserves distinctions such as:

**observation → inference**,  
**hypothesis → established result**,  
**correlation → causal claim**,  
**conditional result → general result**,  
**disputed → settled**,  
**not observed → absent**,  
**source's interpretation → receiver's own assertion**.

The reviewed literature does not provide sufficient controlled evidence to claim that merely prefixing statements with labels such as `OBSERVATION`, `HYPOTHESIS`, or `LIMITATION` reliably solves this problem across models. **This is one of the most important genuine gaps.**

Explicit epistemic metadata is nevertheless experimentally attractive because it converts otherwise discourse-level information into recoverable fields. The appropriate question is not “are labels good?” but whether explicit labels increase downstream typed-edge/typed-claim reconstruction *without* increasing blind acceptance, inducing label-copying, or wasting context. That should be tested against semantically matched prose in which the same epistemic distinctions are stated linguistically rather than encoded in metadata.

“Does not establish” and explicit negative-boundary statements deserve their own category. They are not equivalent to ordinary negative facts. “Study A did not detect X” licenses neither “X does not exist” nor “Study A disproved X.” A benchmark should contain canonical examples of these distinctions and score directionality explicitly.

Uncertainty should likewise be represented as a target object rather than only a confidence number. A source can report statistical uncertainty, methodological uncertainty, unresolved source disagreement, epistemic hedging, or missing evidence. A receiver that replaces any of these with confident prose has lost information even if its central factual conclusion remains approximately correct.

**PROVENANCE AND CITATION STRUCTURE**

ALCE is an important precedent because it treats citation-bearing answers as jointly evaluable along answer and citation dimensions over ASQA, QAMPARI, and ELI5 rather than considering a citation marker itself proof of grounding. citeturn14search0turn14search12

For the proposed benchmark, provenance should be stricter than conventional bibliographic presence. A bibliography answers “what sources were in the corpus?”; a **claim-source map** answers “which source licenses this particular proposition?” These are different reconstruction tasks. Existing citation work strongly supports evaluating the latter, but the reviewed evidence does **not** yet establish that a claim-ledger representation beats inline citations or footnotes when the underlying evidence is held constant. That comparison should therefore remain experimental rather than become a design assumption. citeturn14search0

The gold corpus should support at least four provenance patterns:

| Pattern | Necessary gold information |
|---|---|
| One claim, one source | claim → source edge |
| One claim, multiple corroborating sources | claim → {sources} plus independence/redundancy status |
| One source, many claims | separate edges per claim |
| Conflicting sources | source-specific claim nodes or stance edges rather than one merged “truth” |

Source metadata such as publication date, method, sample, authority or evidence quality should remain distinguishable from the proposition itself. A model should be able to say “Source B reports X but Source A reports Y” without prematurely choosing one. Direct empirical evidence on how strongly source-authority labels bias heterogeneous LLMs remains too sparse to prescribe an optimal authority schema; authority should therefore be tested separately from provenance structure.

**NATURAL LANGUAGE VS EXPLICIT STRUCTURE**

There is no empirical basis for treating this as a one-dimensional “human-readable versus machine-readable” spectrum.

Natural prose carries **discourse relations** cheaply: causality, contrast, concession, motivation, temporal sequence, exception, and interpretation can be expressed without repeated field labels. It can also leave epistemic relations implicit or rhetorically buried.

Explicit structures make identities and categories easier to locate but incur delimiters, field names, nesting, and repetition. They can atomize related propositions so aggressively that the relationship becomes an inference the receiver must reconstruct.

The strongest current evidence is therefore compatible with a **hybrid hypothesis**, but does not yet establish a hybrid optimum. In Kato and Kato, ordinary prose remained competitive with strongly structured specifications; in VeyraBench, prose could be best or tied-best in some long-context cells; in Table Meets LLM, structural annotations and explanations helped some operations but harmed others. citeturn22view1turn22view2turn13view0

A concise narrative plus explicit claim/provenance structure is consequently worth testing, but **must not be privileged in advance**. It could dominate by supplying both discourse and retrieval cues; it could also waste tokens by duplicating content.

**TABLES AND DENSE INFORMATION**

Tables are a particularly important case because they can dramatically reduce repeated linguistic scaffolding while imposing a two-dimensional alignment problem.

Table Meets LLM shows that models can exploit tabular serializations but that results depend on format, task, model, prompting regime, and surrounding context. Row/column association, merged structures, cell lookup, and reasoning over irregular tables are separable capabilities rather than one “table understanding” skill. citeturn13view0

TQA-Bench shows an additional failure regime: **multi-table grounding and joins**. It attributes substantial multi-table penalties particularly to identifying cross-table entities, join scope, and the intended aggregation set. Thus, excellent single-table extraction should not be treated as evidence that compact tabular dossiers preserve arbitrary relationships. citeturn22view3

For evidence whose natural semantics are matrix-like—repeated measurements over the same variables, matched numerical comparisons, experimental conditions—tables may produce excellent density. For causal arguments, competing interpretations, methodological limitations, or claim-specific provenance, they may require auxiliary structure. This is a hypothesis to test by evidence category rather than a reason to exclude tables.

## Compression, recursion, efficiency, model variation, and contamination

**COMPRESSION AND RECURSION**

Compression work contains an important warning for this project: **downstream answer accuracy is an incomplete compression objective**.

LLMLingua uses coarse-to-fine prompt compression intended to preserve task-relevant information while removing tokens; LongLLMLingua extends this to long contexts with question-aware prioritization and reordering; RECOMP learns to compress retrieved material before it is consumed by the generation model. All show that shorter context can retain or sometimes improve downstream QA relative to untreated context. citeturn16search0turn16search1turn16search2

Yet this success can arise precisely because the compressor discards information not relevant to the current query. That is desirable for query-focused QA but dangerous for a **reusable evidence representation** whose future questions are unknown. A representation benchmark therefore needs at least two compression regimes:

**query-agnostic transmission**, where all evidence categories matter; and  
**query-focused transmission**, where selective omission can be rational.

Mixing them would unfairly reward a representation that discarded evidence irrelevant only to the present exam question.

The existing factuality literature is much stronger on **commission errors**—unsupported content—than **structured omission**. There is not strong enough evidence in the reviewed controlled literature to assert a universal deletion order such as “caveats disappear first, then negative findings, then minority evidence.” The appropriate benchmark strategy is therefore to label information categories in advance and **measure category-specific survival** rather than bake an assumed hierarchy into scoring.

Relevant omission classes should include: caveats, null or negative results, minority/contrary evidence, uncertainty, provenance, scope conditions, exceptions, causal qualifiers, methodological limitations, contradictions, and explicit unknowns.

Wu et al.'s recursive book summarization is the closest canonical demonstration that hierarchical recursive summarization can work at very long scales. Their system summarizes sections, then recursively summarizes summaries, with specialised training and human-feedback supervision; the resulting book-level summaries were useful enough to support strong NarrativeQA performance. citeturn21search0

But this does **not** establish that an arbitrary chain—

`source → model summary → second model → synthesis → third model`

—preserves a scientific evidence structure. Wu et al.'s system was deliberately trained for its recursive task and optimised with human feedback. The direct literature on **typed evidence survival across successive heterogeneous model transformations** remains sparse. This should be treated as a novel benchmark axis, not as a solved summarization problem. citeturn21search0

A useful recursive-survival experiment would retain one fixed hidden gold graph \(G_0\). After every transformation \(t\), extract a reconstructed graph \(G_t\) and calculate claim, provenance, relation, scope and epistemic F-scores against \(G_0\). Crucially, also calculate **new unsupported nodes and edges per generation**. This separates omission-driven decay from hallucination-driven drift.

**RESOLUTION MONOTONICITY**

No established term in the reviewed literature maps closely enough to the proposed property to displace **resolution monotonicity**.

Hierarchical summarization, recursive summarization, coarse-to-fine processing, prompt compression, and progressive context allocation are related concepts, but they generally optimize endpoint performance or summary quality rather than require conclusions at budget \(B_{n+1}\) to be a defensible refinement of conclusions at \(B_n\). citeturn21search0turn16search0turn16search1

The property should be operationalized on **claims**, not whole-summary textual similarity.

For adjacent budgets \(B_i,B_{i+1}\), every assertion in the lower-resolution answer can be classified as:

| Transition | Interpretation |
|---|---|
| retained | still supported unchanged |
| refined | still supported but made more specific |
| qualified | retained with added scope/uncertainty |
| superseded correctly | lower-budget answer was explicitly tentative and gets resolved |
| corrected | lower-budget representation caused an unjustified assertion |
| contradicted | higher budget reverses an earlier assertion |
| dropped | disappears despite remaining relevant |

The especially damaging quantity is not all confidence change. Good evidence should change confidence. The harmful quantity is the rate at which **unqualified claims introduced at low resolution later require substantive retraction**.

A candidate **Resolution Contradiction Rate** is therefore:

\[
RCR(B_i,B_{i+1})=
\frac{\text{unsupported or wrong claims at }B_i\text{ contradicted/corrected at }B_{i+1}}
{\text{all substantive claims asserted at }B_i}
\]

It should be reported alongside a **Refinement Rate** and **New Evidence Gain**, because a representation that says almost nothing at low budgets could trivially achieve zero contradictions.

A still better summary is a transition matrix across every budget pair plus a Pareto curve of **coverage versus contradiction risk**.

**TOKEN EFFICIENCY**

The literature studies token efficiency primarily through prompt compression, runtime/cost reduction, and task performance after compression—not through a generally accepted measure of complete evidence fidelity per token. LLMLingua and LongLLMLingua report compression ratios and downstream task performance; VeyraBench explicitly compares format token overhead and gives illustrative cost-adjusted accuracy; TQA-Bench directly measures tokenizer costs of equivalent tabular databases. citeturn16search0turn16search1turn22view2turn22view3

A simple

\[
\frac{\text{fidelity}}{\text{tokens}}
\]

ratio is attractive but statistically dangerous.

First, different fidelity dimensions do not have natural commensurate units. One extra provenance edge is not obviously equivalent to one recovered numeric result.

Second, ratios can overreward very short representations that recover a small easy subset of evidence.

Third, fidelity often has ceilings and nonlinear context effects.

Fourth, tokenizer differences mean “1,000 tokens” are not identical physical representations across model families.

Fifth, a denser representation can be less usable if its relations are harder to decode.

The preferred primary analysis should therefore be a **Pareto frontier**: at each effective context budget, identify representations that are not dominated on both token use and fidelity.

A benchmark can additionally report:

\[
\Delta F / \Delta T
\]

as **marginal fidelity gained per additional token** between adjacent resolution levels, and the **area under a fidelity-versus-budget curve** over a preregistered budget range.

Token counts should be calculated **separately with every evaluated model's actual tokenizer**. Byte count, Unicode character count, and whitespace-normalized word count can be reported as secondary model-independent descriptors. VeyraBench and TQA-Bench make clear why assuming equal token costs from equal semantic content is invalid. citeturn22view2turn22view3

**CROSS-MODEL DIFFERENCES**

Cross-model variation is now too strong to treat as a secondary detail.

He et al. showed GPT-3.5 and GPT-4 did not have identical format sensitivity. QASU reports different structural skill levels across GPT, Gemini, Qwen, Llama and Nova families. Kato and Kato find that format effects that disappear under complete information for GPT-5.4 mini remain for Gemma 3 4B and Llama 3.2 3B. VeyraBench reports format-ranking reversals among Claude, Gemini and Qwen models as scale changes. citeturn7search0turn22view0turn22view1turn22view2

This strongly argues against optimizing a transmission representation against one judge/receiver architecture. The appropriate estimand is at least two-dimensional:

**within-model effect**: what does representation change for model \(m\)?  
**cross-model portability**: how much of that effect survives across heterogeneous receivers?

Model scale can reduce some surface sensitivity but not eliminate it. Kato and Kato's complete-information result is especially instructive: one higher-capability model became insensitive to the tested formats while smaller open models did not. That suggests greater capability can sometimes render formatting less important, but it is not sufficient evidence for a general monotonic “bigger model = format invariant” law. citeturn22view1

Reasoning-mode effects remain under-researched under tightly matched representation experiments. Reasoning-enabled models should therefore form a distinct preregistered model class rather than being silently pooled with ordinary instruction-following models.

The current direct evidence base is uneven by vendor. GPT, Claude/Gemini, Qwen, Llama/Gemma and several other families now appear in controlled structural studies, while rigorous semantically matched representation comparisons for **DeepSeek and some Mistral variants remain thinner**. Absence of evidence for a format preference in those families should not be interpreted as invariance.

**PRIOR-KNOWLEDGE CONTAMINATION**

This is a critical validity threat. An answer can be correct while the model ignores the supplied evidence.

The strongest designs combine several controls rather than relying on obscurity alone.

A **closed-book baseline** asks exactly the same questions without the evidence. Any high closed-book accuracy flags items that cannot cleanly measure intake.

**Source-specific details**—exact sample sizes, uncommon numeric values, unusual methodological choices, source-local terminology—are harder to reconstruct from generic subject knowledge, although a model may still have memorized the source.

**Recent or niche literature** reduces familiarity but cannot guarantee absence from training or retrieval augmentation.

**Synthetic facts** can guarantee novelty if generated procedurally rather than by another model. VeyraBench illustrates this approach with a deterministic fictional corpus containing 8,780 uniquely named entities generated from a fixed seed with no LLM involvement. citeturn22view2

**Counterfactual evidence** is stronger still: present a corpus in which a familiar relation is intentionally altered, then test whether the receiver follows the supplied evidence rather than parametric expectation. Research on contextual-versus-parametric knowledge conflict demonstrates that these conflicts materially affect model behaviour, which is exactly why such cases are diagnostic. AdaCAD, for example, evaluates explicit contextual/parametric conflicts over several QA and summarization datasets and models rather than assuming context automatically dominates prior knowledge. citeturn17academia40

**Cross-document questions** whose answer is not stated in any source but must be composed from two source-local details further reduce the usefulness of memorized domain knowledge.

For the intended niche natural-science corpus, the strongest contamination defence is therefore **real evidence plus synthetic counterfactual twins**. The real set provides ecological validity. The counterfactual twin changes selected entity labels, directions, dates or numerical relationships while preserving discourse and difficulty. A model succeeding on both is much more likely to be reading the supplied corpus.

The counterfactual set must never be allowed to contaminate the real-set ground truth; it is an intake diagnostic, not a scientific assertion.

## Benchmark methodology, examination, scoring, and statistics

**BENCHMARK DESIGN LESSONS**

The central design problem is deciding what “same information” means.

**Option A: identical wording rearranged structurally** gives the strongest causal isolation. The content-bearing lexical units can be kept almost identical while separators, nesting, headings, delimiters, rows and fields change. Its weakness is ecological validity: a natural JSON rendering and a natural paragraph do not ordinarily contain exactly the same words, and forcing lexical identity can make some formats pathological.

**Option B: semantically equivalent, format-native rewrites** is ecologically realistic. A table can behave like a good table; prose can behave like good prose. Its weakness is that wording, redundancy, discourse relations and implicitness change at the same time as format. Apparent “format effects” may actually be paraphrase effects.

The preferable solution is **two complementary experiments, not a compromise that weakens both**.

The first should be an **isomorphic-rendering experiment**: derive every candidate from one canonical content representation and hold wording/order as invariant as each syntax permits. This estimates a relatively pure surface-format effect.

The second should be a **format-native representation experiment**: expert authors construct natural versions from the same hidden canonical evidence graph. Independent adjudicators verify semantic equivalence before evaluation. This estimates the practically relevant effect of information architecture plus serialization.

Kato and Kato's matched target/tests design provides a strong analogue for the first track, while the broader serialization studies show why the second is still necessary. citeturn22view1turn13view0turn22view3

The hidden canonical representation must **not** itself be considered a candidate “winning format.” It is instrumentation: a benchmark annotation ontology against which all exposed representations are scored.

Every representation should encode exactly the same **licensed information inventory**. If one format contains an explicit causal relation while another merely lists two facts from which a human could perhaps infer causality, the experiment has altered semantic content, not merely representation.

Likewise, citations, repeated facts, examples, glossaries, definitions, source titles, source authority, and explanatory bridges are all information-bearing features. They cannot be allowed to vary accidentally.

**EXAMINATION DESIGN**

A good examination needs to force several kinds of reconstruction. Pure fact recall would favour retrieval-oriented structures and badly underestimate evidence damage.

A first-round examination should cover approximately the following strata; the exact proportions should be pilot-adjusted rather than treated as theoretically privileged:

| Question class | What it isolates | Scoring ambiguity |
|---|---|---|
| Atomic fact / numeric extraction | retrieval and local binding | low |
| Source attribution | provenance | low to moderate |
| Conditional/scope question | qualification survival | moderate |
| Contradiction reconciliation | multi-source relational reconstruction | moderate |
| “What does the evidence **not** establish?” | negative boundary / epistemic restraint | moderate if canonical alternatives supplied |
| Missing-information / unknown question | resistance to fabrication | low to moderate |
| Causal-versus-correlational discrimination | relation typing | moderate |
| Multi-source synthesis | global reconstruction | moderate/high |
| Counterfactual application | evidence-use versus memorization | moderate |
| Novel-case transfer | inferential usability | high |
| Confidence/evidence-strength ranking | uncertainty and support weighting | high |

Direct questions should be intermixed with **inverse questions**. For example, after asking what a study found, ask which stronger conclusion is *not* licensed. This detects flattening from “associated with” to “causes” or from “not detected” to “absent.”

Numerical questions should include units, denominators, confidence intervals where present, comparison direction and source. Correctly recalling “17%” while attaching it to the wrong population should not count as full fidelity.

Contradiction questions should include both **true contradictions** and **apparent contradictions caused by scope differences**. Otherwise a model can learn that the exam wants it always to “reconcile” disagreement.

Novel-case transfer should use **closed-world cases** that specify all non-evidence assumptions necessary for the answer. Otherwise grading will inadvertently measure outside domain knowledge.

Open-ended causal questions create the greatest evaluator ambiguity and should therefore have a structured auxiliary answer: selected source IDs, relevant claim IDs, relation type and free-text explanation. This allows partial deterministic grading even when prose needs semantic adjudication.

**MEASUREMENT AND SCORING**

The benchmark should resist collapsing all errors into one score. A high-quality representation should expose a **fidelity profile**.

| Metric | Recommended operationalization | Status in existing literature |
|---|---|---|
| **Claim Fidelity** | macro/micro F1 over canonical atomic claims | strongly grounded by atomic factuality work |
| **Epistemic Fidelity** | accuracy/F1 over claim-type labels such as observation, inference, hypothesis, disputed, unknown | largely novel extension |
| **Provenance Fidelity** | claim→source edge precision/recall/F1 | strongly motivated by citation/attribution work |
| **Relational Fidelity** | typed edge F1: supports, contradicts, causes, correlates, qualifies, depends-on, temporal-before, etc. | novel synthesis |
| **Scope Fidelity** | recovery of population, condition, method, temporal and applicability qualifiers | under-measured; benchmark extension |
| **Uncertainty Fidelity** | preservation of stated uncertainty category/range and non-resolution of unresolved conflicts | under-measured |
| **Unsupported Inference Rate** | unsupported generated atomic claims / generated substantive claims | well motivated by FActScore/RAGTruth |
| **Source Confusion Rate** | incorrect claim→source associations / attempted associations | natural extension of citation evaluation |
| **Semantic Omission Rate** | missed canonical items, stratified by evidence category | conventional recall adapted to evidence ontology |
| **Recursive Survival** | fidelity vector after each transformation depth | largely novel |
| **Resolution Contradiction Rate** | low-resolution assertions requiring retraction at higher budget | novel |
| **Novel-Case Transfer** | accuracy on preregistered applications requiring supplied evidence | established evaluation idea; new use here |

FActScore directly supports proposition-level decomposition; ALCE supports separating answer and citation quality; RAGTruth supports explicit unsupported-content annotation. citeturn15search1turn14search0turn14search2

A useful high-level score can be reported, but it should not replace the vector. Otherwise a representation could compensate for catastrophic provenance loss with excellent raw claim recall.

Where deterministic evaluation is possible, use it. Appropriate examples include source IDs, multiple-choice epistemic type, exact numeric values within tolerance, units, dates, Boolean relation labels, and known canonical claim IDs.

For free-text claim recovery, decompose the answer into propositions and compare them against the gold claim set. String exact match is too brittle; unconstrained embedding similarity is too forgiving. Semantic adjudication should use a preregistered entailment rubric and periodically be checked by human experts.

A **claim-graph comparison** is particularly well matched to the research objective. Let the gold be \(G=(V,E)\), with typed claim nodes \(V\) and typed relational/provenance edges \(E\). Score node recovery and edge recovery separately. A receiver that recovers all nodes but scrambles edges then receives high factual recall but low relational/provenance fidelity—which is precisely the distinction the proposed benchmark needs.

**Model-as-judge** can be useful but should not be the sole adjudicator. G-Eval demonstrated that GPT-4-based evaluation can correlate substantially with human summarization judgements—reported Spearman correlation 0.514 in its experiment—but also identified concern about evaluator preference for LLM-generated text. citeturn19search2

Zheng et al.'s MT-Bench/Chatbot Arena analysis documents position, verbosity and self-enhancement biases in LLM judges. Subsequent work has independently studied position bias, confirming that presentation order can alter judge decisions. citeturn19search0turn19search1

The safest scoring architecture is therefore:

deterministic tests where possible → reference-guided semantic judges for genuinely semantic items → blinded human audit on a stratified sample.

For pairwise LLM judging, swap answer order and require consistency. Do not expose candidate representation names to the judge. Do not use the same model architecture as both receiver and sole judge. For high-value semantic categories, use multiple independent judges from different model families and estimate their agreement against human adjudication.

**STATISTICAL DESIGN**

Representation comparison is naturally a **paired repeated-measures problem**. Every canonical evidence item/question should be exposed under multiple representation conditions, so difficulty differences are controlled within item.

The statistical unit must be the **independent evidence/question item or evidence packet**, not every API call. Three stochastic generations of the same model on the same item are repeated observations, not three independent scientific replicates. Treating them as independent would be pseudo-replication.

A suitable confirmatory analysis is a hierarchical generalized model with representation, context budget, position, task class and model class as experimental factors; item/evidence packet should receive random intercepts and, where data support them, random representation slopes. Key interactions should include:

\[
\text{representation} \times \text{model}
\]

\[
\text{representation} \times \text{context budget}
\]

\[
\text{representation} \times \text{position}
\]

and, selectively,

\[
\text{representation} \times \text{task type}.
\]

For binary item correctness, use logistic mixed modelling or paired bootstrap/McNemar analyses where appropriate. TQA-Bench itself uses paired McNemar tests on aligned multi-table cases, illustrating the value of paired rather than aggregate-only comparisons. citeturn22view3

For continuous composite fidelity scores, report estimated marginal differences and confidence intervals, not only \(p\)-values. NLP methodology has long shown that significance conclusions can depend on the testing scheme; paired resampling at the true item level is preferable to treating aggregate benchmark scores as error-free. citeturn15search0

No universal number of repeated generations is defensible without a variance estimate. A pilot should estimate within-condition API/model stochasticity, after which power should be determined by simulation for a preregistered minimum effect of practical interest. Three to five repeated generations per condition are reasonable for estimating low-temperature inference variance, but **increasing the number of independent evidence packets is statistically more valuable than generating dozens of responses to the same packet**.

Temperature, top-p, reasoning effort, maximum output length, tool availability, system prompt, API endpoint, exact model version, tokenizer, context window and experiment date should all be frozen and logged. Provider-side models can change, so an exact model identifier and collection period are part of the experimental condition.

The benchmark should report **model-specific effects first** and any cross-model mean second. Otherwise one high-volume or high-capability family can dominate an aggregate score.

The principal scientific target should not be “the universally best format.” Current evidence makes that hypothesis unnecessarily strong. The more defensible target is:

> **the set of Pareto-optimal representation strategies by context budget, task class and receiver class, plus any representation properties whose advantages generalize across heterogeneous models.**

**FAILURE MODES AND CONFOUNDS**

A format can look excellent while failing the actual purpose in at least the following ways.

High claim recall can coexist with **epistemic flattening**. A receiver may preserve every result while silently converting tentative interpretations into facts.

High extraction accuracy can coexist with **relationship destruction**. Atomized facts may be individually recoverable even after causal, conditional or contradictory links are lost.

High citation presence can coexist with **source confusion**. Citation count is not claim-source correctness; ALCE's separate citation evaluation is a precedent for not making this mistake. citeturn14search0

High compression can be achieved by **dropping the difficult evidence categories**. If caveats and conflicting sources disappear, the representation may become easier rather than more efficient.

Correct answers can come from **parametric prior knowledge**. Closed-book and counterfactual controls are essential.

Redundancy can inflate recall. A repeated proposition is not three retained propositions.

Structure can leak the exam. A field called `causal_relation: false` trivializes a later causal question unless that explicitness is itself the manipulated variable under study.

Question wording can privilege a representation. Asking “what value appears in the `sample_size` field?” would advantage structured candidates.

Token matching can be unfair across models. Equivalent strings tokenize differently, and format overhead itself differs. citeturn22view2turn22view3

One model family can dominate an aggregate benchmark. Cross-family portability must be visible rather than averaged away. citeturn22view0turn22view1

A judge can prefer verbose, familiar or similarly styled answers. LLM judge position/verbosity/self-preference effects make blinded, multi-judge and human-audited evaluation necessary. citeturn19search0turn19search1

Long-context failures can masquerade as format failures if one encoding simply pushes target evidence farther into the effective-context danger zone. Format and position/length must therefore be crossed. citeturn1search4turn22view2

Retrieval success can masquerade as understanding. Literal needle tests are especially vulnerable to this confound. citeturn1search3turn8search1

A representation can preserve conclusions while destroying **evidentiary justification**. Answer-only metrics would call this success even though the receiving agent could no longer reconstruct why the conclusion is warranted.

Finally, a representation can optimize to a particular benchmark ontology. That is why a held-out **novel evidence packet** and novel-case transfer component are necessary.

## Design synthesis and experimental requirements

**OPEN QUESTIONS**

Several important questions remain genuinely unresolved after the current literature.

There is no convincing cross-model answer to whether **explicit epistemic tags** outperform carefully worded epistemic prose.

There is no robust estimate of how much evidence is lost specifically from **limitations, caveats, null findings, exceptions and source disagreement** at equal token budgets.

Direct comparisons of **ordinary bibliography, inline citation, source IDs, claim-source ledgers and bidirectional claim/source ledgers** under constant semantic content are scarce.

The literature does not yet tell us whether **graph-like evidence serializations** remain useful when their markup/token cost and graph traversal burden are controlled.

There is little direct work on repeated transmission through **heterogeneous** models rather than a single recursively trained summarizer.

It is unknown whether a representation that is good for reconstruction at one resolution also exhibits good **resolution monotonicity** across budgets.

There is no established, validated scalar equivalent of **semantic fidelity per token** for multidimensional evidence.

Reasoning-enabled versus ordinary instruction models have not been compared comprehensively enough under fully matched evidence serializations.

And there is insufficient evidence to say whether a representation optimal for prose-heavy scientific evidence will also be optimal for quantitatively dense or highly relational corpora.

These are not gaps to “solve” by intuition. They are the experimental agenda.

**DESIGN REQUIREMENTS**

The later benchmark should satisfy the following empirically justified requirements.

1. **Maintain a hidden canonical evidence ontology independent of every candidate representation.**  
   **Evidence basis:** atomic factuality and citation evaluation show the value of decomposable reference units, while the proposed benchmark needs additional typed relations. citeturn15search1turn14search0  
   **Confidence:** High.  
   **Consequence if ignored:** semantic equivalence cannot be verified, and differences between representations become uninterpretable.

2. **Separate surface serialization from information architecture.**  
   **Evidence basis:** plain/Markdown/JSON/YAML differences occur even with closely matched content, while Kato and Kato show that information completeness itself changes format effects. citeturn7search0turn22view1  
   **Confidence:** High.  
   **Consequence if ignored:** “format effects” will confound syntax with added or removed information.

3. **Cross representation with context budget.**  
   **Evidence basis:** long-context performance deteriorates with length, and VeyraBench finds representation rankings can change near effective limits. citeturn8search1turn1search3turn22view2  
   **Confidence:** High.  
   **Consequence if ignored:** a format that looks superior at 8K may be incorrectly assumed superior at 64K or 256K.

4. **Cross representation with evidence position.**  
   **Evidence basis:** lost-in-the-middle effects and ordering effects are well demonstrated. citeturn1search4turn13view0  
   **Confidence:** High.  
   **Consequence if ignored:** an accidental favourable source order can be mistaken for a representation advantage.

5. **Use multiple independent evidence packets rather than one monolithic dossier.**  
   **Evidence basis:** task/content interactions recur across LongBench, table benchmarks and format studies. citeturn9search0turn13view0turn22view1  
   **Confidence:** High.  
   **Consequence if ignored:** results describe one particular scientific narrative rather than a representation effect.

6. **Include heterogeneous receiving model families and publish per-model results.**  
   **Evidence basis:** QASU, Kato and Kato, He et al. and VeyraBench all find model-dependent effects. citeturn22view0turn22view1turn7search0turn22view2  
   **Confidence:** High.  
   **Consequence if ignored:** the experiment can overfit a transmission language to one receiver.

7. **Measure provenance independently from factual correctness.**  
   **Evidence basis:** ALCE treats citation quality as a distinct evaluation dimension. citeturn14search0  
   **Confidence:** High.  
   **Consequence if ignored:** a model may provide the correct conclusion while attaching it to the wrong evidence.

8. **Measure relations and scope independently from atomic claim recovery.**  
   **Evidence basis:** NoLiMa, RULER and TQA-Bench show that associative, multi-hop and cross-table reasoning is harder than simple retrieval. citeturn1search3turn8search1turn22view3  
   **Confidence:** High.  
   **Consequence if ignored:** an atomized representation can appear lossless while destroying the actual structure of evidence.

9. **Include explicit unsupported-answer and unknown-information probes.**  
   **Evidence basis:** RAGTruth shows unsupported generation is separably measurable. citeturn14search2  
   **Confidence:** High.  
   **Consequence if ignored:** formats that encourage plausible filling-in receive no penalty.

10. **Include anti-contamination controls.**  
    **Evidence basis:** contextual knowledge can conflict with model priors; fully synthetic controlled corpora have already been used to eliminate contamination in long-context experiments. citeturn17academia40turn22view2  
    **Confidence:** High.  
    **Consequence if ignored:** apparent evidence intake may be memorized knowledge.

11. **Measure actual tokenization separately for every receiver.**  
    **Evidence basis:** identical semantic databases incur radically different token costs across formats, and markup overhead can change conclusions about efficiency. citeturn22view2turn22view3  
    **Confidence:** High.  
    **Consequence if ignored:** “equal budget” comparisons are not actually equal.

12. **Evaluate fidelity-budget Pareto curves rather than declaring the shortest representation best.**  
    **Evidence basis:** prompt compression can reduce token count and improve task performance, but format overhead and accuracy interact nonlinearly. citeturn16search0turn16search1turn22view2  
    **Confidence:** High.  
    **Consequence if ignored:** aggressive information deletion can masquerade as efficiency.

13. **Use both isomorphic renderings and natural format-native renderings.**  
    **Evidence basis:** tightly controlled matched studies provide causal clarity, whereas serialization/task studies demonstrate that natural structures also matter. citeturn22view1turn13view0  
    **Confidence:** Medium-high.  
    **Consequence if ignored:** either ecological validity or causal identification will remain weak.

14. **Do not score solely with an LLM judge.**  
    **Evidence basis:** LLM judges exhibit position, verbosity and self-preference biases despite useful human correlations. citeturn19search0turn19search1turn19search2  
    **Confidence:** High.  
    **Consequence if ignored:** a candidate representation may be selected because it resembles the judge's preferred style.

15. **Evaluate recursive transmission explicitly rather than assuming first-generation fidelity predicts later fidelity.**  
    **Evidence basis:** recursive summarization is feasible under specialised training, but generic heterogeneous-chain fidelity has not been established. citeturn21search0  
    **Confidence:** Medium-high.  
    **Consequence if ignored:** small first-hop distortions may accumulate unseen across agent chains.

16. **Include multiple resolution levels derived from the same evidence.**  
    **Evidence basis:** compression and long-context literature establish meaningful budget-performance tradeoffs, but monotonic refinement itself is unmeasured. citeturn16search0turn16search1  
    **Confidence:** Medium.  
    **Consequence if ignored:** the benchmark cannot test whether a representation degrades gracefully.

17. **Preserve negative evidence, contradictions and uncertainty as first-class gold annotations.**  
    **Evidence basis:** existing factuality benchmarks chiefly reward supported propositions; this leaves structured omission insufficiently observed. citeturn15search1turn14search2  
    **Confidence:** High as a methodological requirement, lower regarding which representation will help.  
    **Consequence if ignored:** the benchmark itself reproduces the epistemic flattening it aims to detect.

18. **Use paired, item-level statistics and avoid pseudo-replication.**  
    **Evidence basis:** aligned comparisons such as TQA-Bench use paired testing, and statistical methodology in NLP emphasizes the importance of correct significance procedures. citeturn22view3turn15search0  
    **Confidence:** High.  
    **Consequence if ignored:** confidence intervals will be too narrow and format effects may be declared significant from repeated calls rather than independent evidence.

**CANDIDATE FORMAT DIMENSIONS**

The experimental object should initially be described by **orthogonal representational dimensions**, not brand-name syntaxes.

| Dimension | Low end | High end |
|---|---|---|
| Discourse | continuous explanatory narrative | atomized propositions |
| Sectioning | flat stream | explicit hierarchy |
| Relation encoding | linguistically implicit | explicitly typed |
| Epistemic encoding | prose hedges/qualifiers | explicit epistemic labels |
| Provenance | source list / inline citation | claim-specific mapping |
| Locality | locally self-contained claims | compact cross-references |
| Redundancy | single encoding | strategically repeated key evidence |
| Numerical encoding | prose numbers | aligned tabular/matrix form |
| Schema rigidity | free text | fixed fields |
| Serialization verbosity | compact delimiters | verbose tagged markup |
| Narrative explanation | absent | causal/conceptual bridges |
| Scope encoding | prose qualification | explicit condition fields |
| Contradiction handling | sources remain separate | explicit contradiction links |
| Resolution | coarse synopsis | increasingly complete evidence |
| Reference direction | source→claims only | bidirectional claim↔source indexing |

JSON, YAML, XML, Markdown, CSV, prose, tables, triples and graphs are then **implementations at points in this design space**, not explanatory variables in themselves.

This framing is crucial because two JSON documents can differ more profoundly in epistemic structure than a JSON document and a prose dossier.

**PROPOSED EXPERIMENTAL FACTORS**

A manageable first round should be a **screening experiment**, not an attempted exhaustive tournament of every imaginable syntax.

**Variables that MUST be controlled**

The latent evidence inventory; canonical claim/edge ontology; source documents; question set; evidence multiplicity; evidence ordering within each position condition; instructions; system prompt; model version; reasoning mode; temperature/sampling; maximum generation budget; tool access; language; citation-information availability; and exact tokenizer accounting should be controlled or explicitly crossed.

No candidate may silently gain a definition, example, causal explanation, source detail or caveat absent from another unless that feature is the factor being manipulated.

**Variables worth manipulating immediately**

A useful first experiment would manipulate four conceptual axes:

- **discourse organisation:** narrative/coherent versus atomized;
- **epistemic explicitness:** linguistically encoded versus explicit labels;
- **provenance explicitness:** ordinary source citation versus claim-specific mapping;
- **structural compactness:** lightweight versus schema-heavy serialization.

Rather than run every possible combination, use a balanced fractional-factorial set of approximately **six to eight representation conditions**, including pure and hybrid points. The purpose is to estimate main effects and the largest interactions, not crown one representation.

Cross those representation conditions with **three evidence budgets**—for example coarse, medium and near-complete—and **three position regimes**: favourable/local, middle/adverse, and distributed. Exact token budgets should be selected after pilot measurement so they fall inside the useful operating range of all chosen models rather than at arbitrary round numbers. Long-context studies show why a context ladder must be tied to effective, not merely advertised, limits. citeturn8search1turn1search3turn22view2

Use at least **four to six heterogeneous receiver classes**: one current OpenAI model, one Anthropic model, one Gemini model, and several independently trained open/open-weight families such as Qwen, Llama/Mistral-class and DeepSeek-class systems where practical. Exact current snapshots matter more than brand labels.

Within the natural-science domain, construct **multiple independent evidence packets** rather than treating one literature review as the sole experimental unit. Each should contain both easy and difficult evidence topology: single-source facts, cross-source dependencies, genuine disagreements, qualified findings, negative findings, methods, numerical findings and unresolved unknowns.

**Variables that should wait for later experiments**

Query-aware compression/reordering, strategic repetition, authority labels, retrieval-stage variation, chain-of-thought prompting, tool use, agent planning, adaptive context selection, multilingual transfer, multimodal evidence, and dynamic memory should be held out initially. Each can produce large effects but would make attribution to representation much harder.

Likewise, **recursive transmission** should be a second-stage experiment after first-hop reconstruction is understood. Otherwise it will be unclear whether a chain fails because the original representation was poor or because later transformations amplify otherwise small distortions.

## Claim ledger, source ledger, and corpus recommendation

**CLAIM LEDGER**

| ID | Label | Major conclusion | Evidence / counterevidence | Confidence | Sources / scope |
|---|---|---|---|---|---|
| C01 | **REPLICATED** | Semantically matched information can produce different downstream performance under different representations. | Multiple independent format/table/specification studies; effect sizes vary. No universal ordering. | High | S5–S10 citeturn7search0turn22view0turn22view1turn13view0turn22view2turn22view3 |
| C02 | **REPLICATED** | Format preferences are model-dependent. | GPT sensitivity differs by model; QASU/Kato/Veyra cross-family interactions. | High | S5, S6, S8, S9 citeturn7search0turn22view0turn22view1turn22view2 |
| C03 | **REPLICATED** | There is no evidenced universally superior text serialization. | Rankings reverse across tasks/models/lengths. | High | S5–S10 citeturn22view1turn22view2turn22view3 |
| C04 | **REPLICATED** | Evidence position can materially change long-context performance. | Lost-in-the-middle and related long-context evaluations. | High | S1–S3 citeturn1search4turn8search1turn1search3 |
| C05 | **REPLICATED** | Advertised context length overstates effective context utilisation for demanding tasks. | RULER, NoLiMa, LongBench. | High | S2–S4 citeturn8search1turn1search3turn9search0 |
| C06 | **OBSERVED** | Format effects can become much larger near an effective context ceiling. | Strong controlled 2026 preprint evidence; limited independent replication of exact interaction. | Medium-high | S9 citeturn22view2 |
| C07 | **REPLICATED** | Formatting/serialization can impose substantial token overhead. | Veyra and TQA-Bench independently measure large differences. | High | S9, S10 citeturn22view2turn22view3 |
| C08 | **REPLICATED** | Lower token count does not imply lower task performance; selective compression can improve performance. | LLMLingua family and RECOMP. | High for QA/task performance; low for complete evidence fidelity | S14–S16 citeturn16search0turn16search1turn16search2 |
| C09 | **INFERRED** | Query-focused compression cannot be assumed safe for reusable, unknown-future-query evidence transmission. | Compression papers optimize current downstream tasks, not full evidence survival. | High conceptual confidence | S14–S16 |
| C10 | **REPLICATED** | Literal retrieval benchmarks can overestimate long-context understanding. | RULER and NoLiMa deliberately increase beyond-literal reasoning requirements. | High | S2, S3 citeturn8search1turn1search3 |
| C11 | **OBSERVED** | Atomic proposition scoring is practical for factual fidelity. | FActScore. | High | S13 citeturn15search1 |
| C12 | **OBSERVED** | Unsupported RAG generation can be annotated at fine granularity. | RAGTruth, nearly 18,000 responses. | High | S12 citeturn14search2 |
| C13 | **OBSERVED** | Citation quality is separable from general answer quality. | ALCE explicitly evaluates citation-bearing generation. | High | S11 citeturn14search0 |
| C14 | **UNKNOWN** | Explicit epistemic labels consistently preserve observation/inference/hypothesis distinctions better than prose. | No sufficiently comprehensive matched cross-model evidence found. | Low | Research gap |
| C15 | **UNKNOWN** | Claim-specific ledgers outperform conventional inline citation or bibliography structures. | Citation literature motivates the metric but does not establish the representation winner. | Low–medium | S11 |
| C16 | **CONTESTED** | Greater explicit structure necessarily improves reasoning. | Counterevidence from table/specification experiments; added structure can help or hurt. | High confidence that universal claim is false | S7–S10 citeturn13view0turn22view1turn22view2 |
| C17 | **CONTESTED** | Natural prose is intrinsically less efficient for models. | Prose remains competitive/best in several controlled cells; markup can cost additional tokens. | High confidence that universal claim is unsupported | S8–S10 |
| C18 | **OBSERVED** | Recursive summarization can scale to book-length inputs under specialised training and human feedback. | Wu et al. | High within that setting | S17 citeturn21search0 |
| C19 | **UNKNOWN** | Generic heterogeneous model-to-model recursive summarization preserves scientific epistemic structure. | Nearest literature does not directly test this. | Low | S17 as analogue |
| C20 | **INFERRED** | “Resolution monotonicity” is a distinct useful benchmark property. | Adjacent concepts exist; no equivalent established metric found. | Medium | S14–S17 |
| C21 | **REPLICATED** | LLM judges have systematic evaluation biases. | Position, verbosity, self-preference findings across judge studies. | High | S19, S20 citeturn19search0turn19search1turn19search2 |
| C22 | **OBSERVED** | Models do not always privilege supplied context over conflicting parametric knowledge. | Context/parametric conflict experiments. | Medium-high | S18 citeturn17academia40 |
| C23 | **INFERRED** | Closed-book plus synthetic/counterfactual controls are stronger contamination tests than niche-topic selection alone. | Conflict evidence plus contamination-free synthetic designs. | High | S9, S18 |
| C24 | **INFERRED** | The correct optimization target is a multidimensional Pareto frontier, not a single “fidelity/token” ratio. | Token-cost and compression evidence demonstrates nonlinear tradeoffs. | High | S9, S10, S14–S16 |
| C25 | **HYPOTHESIS** | Hybrid narrative + explicit claim/provenance structure may dominate either extreme on some evidence types. | Compatible with current evidence but not directly established. | Medium-low | Must be experimentally tested |
| C26 | **INFERRED** | A valid benchmark must score relations, scope, uncertainty and provenance separately from fact recall. | Current benchmarks show retrieval, multi-hop, attribution and hallucination are separable behaviours. | High | S2, S3, S11–S13 |

**SOURCE LEDGER**

**S1 — Liu et al., “Lost in the Middle: How Language Models Use Long Contexts,” TACL 2024.** Multi-document QA and key-value retrieval; systematically varies relevant-information position. Establishes strong positional sensitivity and common beginning/end advantage. Limitation: synthetic retrieval and QA do not exhaust evidence synthesis. This phenomenon has since been supported by broader long-context work. citeturn1search4turn1search0

**S2 — Hsieh et al., “RULER: What’s the Real Context Size of Your Long-Context Language Models?”, COLM 2024.** Tests 17 long-context models on multi-needle retrieval, multi-hop tracing and aggregation. Establishes that near-perfect vanilla needle performance can coexist with much lower effective long-context competence and that many models degrade far below claimed window lengths. citeturn8search1

**S3 — Modarressi et al., “NoLiMa: Long-Context Evaluation Beyond Literal Matching,” ICML 2025.** Tests 12 models with ≥128K nominal contexts using questions that reduce lexical overlap with supporting evidence. Reports severe 32K deterioration for most models; establishes that literal-string retrieval can overestimate useful context understanding. citeturn1search3

**S4 — Bai et al., “LongBench,” ACL 2024.** Twenty-one datasets across six long-context task categories and bilingual data. Establishes task heterogeneity and shows that strong performance in one long-context regime cannot safely be generalized to summarization, multi-document QA, code and other regimes. citeturn9search0

**S5 — He et al., “Does Prompt Formatting Have Any Impact on LLM Performance?”, 2024.** Compares plain text, Markdown, JSON and YAML on GPT-family models across several tasks. Establishes substantial format sensitivity and greater robustness in stronger models, but no global winning serialization. Main limitation: restricted model-family coverage and tasks not designed specifically as scientific evidence transmission. citeturn7search0turn10search5

**S6 — Nguyen et al., “Questionnaire Meets LLM,” LREC 2026.** Six structural skills, six serialization formats and five model families including GPT, Gemini, Qwen, Llama and Nova; reports format-driven gains reaching about nine percentage points and large cross-model gaps. Important current cross-family evidence. citeturn22view0

**S7 — Sui et al., “Table Meets LLM,” WWW-era 2024 study.** Draws 1,500 tables per task from TabFact, FEVEROUS, SQA, HybridQA and ToTTo and compares natural-language/separator, Markdown, JSON, XML and HTML representations plus structural prompting. Establishes task-dependent table-format effects and shows that additional structural material/order can sometimes hurt. citeturn11search12turn13view0

**S8 — Kato & Kato, “Which Algorithm Specification Formats Help Language Models Implement Machine Learning Algorithms?”, 2026 preprint.** **4,020 implementations**, five tasks, seven formats, three models, four settings. Particularly valuable matched design: target code/tests fixed while representation changes. Shows model × format × information-completeness interaction and no general structured-format winner. Recent preprint; independent replication pending. citeturn22view1

**S9 — “Prompt Design at Scale,” 2026 preprint / VeyraBench.** Uses a procedurally generated contamination-free 512K corpus of 8,780 fictional entities; four formats, five models, context ladder to 512K where supported. Reports context-dependent ranking reversals, up to 48-point format spread in one 128K condition, and 22–37% format overhead relative to plain text. Methodologically highly relevant but new and not yet independently replicated. citeturn22view2

**S10 — TQA-Bench, 2026 version.** Multi-table QA with Markdown, CSV, JSON and HTML serializations and Qwen/Llama-family models. Directly measures serialization token costs and finds major differences; also isolates cross-table grounding/join/aggregation difficulty. Recent study; strongest relevance is to dense relational/tabular evidence. citeturn22view3

**S11 — Gao et al., “Enabling Large Language Models to Generate Text with Citations,” EMNLP 2023 / ALCE.** Establishes a benchmark framework in which citation quality is independently evaluated alongside answer correctness and fluency on multiple long-form QA datasets. Supports claim-specific provenance evaluation, though it does not establish an optimal evidence serialization. citeturn14search0turn14search4

**S12 — Niu et al., “RAGTruth,” ACL 2024.** Nearly 18,000 naturally generated RAG responses with detailed manual hallucination annotations. Establishes feasibility of granular unsupported-content measurement in retrieval-grounded generation. citeturn14search2turn14search10

**S13 — Min et al., “FActScore,” EMNLP 2023.** Decomposes long-form output into atomic factual propositions and measures factual precision. Provides a strong methodological basis for canonical claim decomposition, though the present benchmark needs relation/provenance/scope extensions. citeturn15search1turn15search30

**S14 — Jiang et al., “LLMLingua,” EMNLP 2023.** Coarse-to-fine prompt compression with token-budget control. Establishes that substantial lexical/context compression can retain useful downstream performance and that token count alone is not the right objective. citeturn16search0

**S15 — Jiang et al., “LongLLMLingua,” ACL 2024.** Query-aware long-context compression/reordering. Reports improvements including up to 21.4% on NaturalQuestions with about four times fewer tokens in its GPT-3.5-Turbo experiment. Establishes the value of density and placement optimisation, but not complete evidence preservation. citeturn16search1

**S16 — Xu, Shi & Choi, “RECOMP,” ICLR 2024.** Retrieval-augmented context compression with selective/extractive/abstractive compression. Supports separating raw context size from useful query-conditioned evidence. citeturn16search2

**S17 — Wu et al., “Recursively Summarizing Books with Human Feedback,” 2021.** GPT-3-based recursively decomposed book summarization trained with demonstrations/human feedback; produced useful book-level summaries and strong NarrativeQA results. Establishes feasibility of recursive summarization under specialised supervision, not generic losslessness across arbitrary agent chains. citeturn21search0

**S18 — Wang et al., “AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge,” 2024.** Tests conflict between supplied context and parametric knowledge over multiple models, QA datasets and summarization tasks. Establishes that contextual evidence and prior model knowledge can compete rather than supplied evidence automatically prevailing. citeturn17academia40

**S19 — Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” NeurIPS-era/OpenReview.** Documents position, verbosity, self-enhancement and reasoning limitations in LLM judging. Central evidence against a single unblinded model judge. citeturn19search0

**S20 — Liu et al., “G-Eval,” EMNLP 2023.** GPT-4-based semantic evaluation achieves reported Spearman 0.514 with human summarization judgements while raising concerns about evaluator bias toward LLM-generated text. Establishes that LLM judges are useful but imperfect semantic evaluators. citeturn19search2

**S21 — Berg-Kirkpatrick, Burkett & Klein, “An Empirical Investigation of Statistical Significance in NLP,” EMNLP-CoNLL 2012.** Methodological evidence that significance-testing choices matter in NLP evaluation; supports paired/item-aware analysis rather than naive comparisons of aggregate scores. citeturn15search0

**RECOMMENDED NEXT STEP**

Given the available evidence, the neutral natural-science benchmark corpus should be built **before** any candidate ARC-style representation is selected, and it should be built as an **evidence-controlled experimental substrate rather than a conventional literature review**.

Choose a niche natural-science domain with enough primary literature to contain genuine methodological variation, disagreements, negative results, numerical findings and causal uncertainty, but not so famous that frontier models can answer most examination questions closed-book. Obscurity alone is insufficient; closed-book testing and counterfactual controls remain necessary. citeturn17academia40turn22view2

The corpus should contain **multiple independent evidence packets within that domain**, rather than one linear narrative. Each packet should answer a separable scientific subquestion and contain several primary sources. This creates real replication units for the representation experiment and prevents a single unusual paper topology from determining the outcome.

Before creating any exposed summary or serialization, expert annotators should construct a **hidden canonical evidence graph**. Its nodes should contain atomic claims, numerical findings, observations, interpretations, hypotheses, definitions, negative findings, methodological statements, limitations and explicit unknowns. Its edges should encode support, contradiction, qualification, source provenance, temporal ordering, causal claims, correlation, dependence, scope, alternative explanation and source-to-interpretation relationships. This hidden graph is measurement infrastructure, **not the proposed final transmission format**. FActScore, ALCE and RAGTruth provide precedents for atomic facts, attribution and unsupported-content annotation respectively, but the richer relational ontology would be a novel extension. citeturn15search1turn14search0turn14search2

Every important scientific distinction should have at least one deliberately testable instance: an observation that must not become an inference; a correlation that must not become causality; an unobserved effect that must not become an established absence; a conditional result that must not become universal; a source disagreement that must not be silently resolved; a negative finding; an exception; a methodological limitation; an unresolved unknown; a source-specific interpretation; and a number whose meaning depends on its population or denominator.

For contamination control, the corpus should have **three layers**. The first is the real primary-source evidence. The second is a closed-book examination establishing what models already know. The third is a small set of **counterfactual twins** in which selected names, magnitudes, directions or relationships are altered while preserving linguistic structure. Success that persists under counterfactual injection is stronger evidence of context intake than success on real facts alone. Procedurally generated synthetic micro-packets can provide an additional zero-contamination calibration set. citeturn17academia40turn22view2

The source corpus itself should remain intact and immutable. Candidate representations should be generated from its canonical annotation under **two separate experimental tracks**: one tightly isomorphic track that maximally holds wording constant, and one format-native track where each representation is allowed to express the same semantic inventory naturally. Equivalence should be adjudicated before models ever see the representations.

Create several preregistered **resolution levels** for every candidate representation. The lowest level should be required to remain defensible rather than merely terse; higher levels add omitted evidence and relationships. This creates the substrate needed to test both ordinary fidelity-budget tradeoffs and the novel property of resolution monotonicity.

The examination should then be generated from the hidden evidence graph, not from the candidate representations. It should contain direct recall only as a minority component and deliberately exercise provenance, numeric binding, epistemic status, scope, contradiction, missing information, negative boundaries, cross-source synthesis and novel-case transfer. No question should contain clues that privilege a particular serialization.

Finally, the corpus should be frozen **before** comparing formats. Representation authors, question authors and final scorers should operate from different views of the benchmark where practical. Model outputs should be evaluated first with deterministic graph/ID/numeric checks, then semantic adjudication, with blinded multi-model judges and human audit for the irreducibly semantic residue. citeturn19search0turn19search1turn19search2

The methodological foundation therefore points to a specific sequence:

**source corpus → hidden evidence ontology → contamination controls → examination → fixed resolution budgets → alternative representations → heterogeneous receivers → multidimensional reconstruction scoring.**

Not:

**choose a format → summarize the literature into it → ask models whether they like it.**

The latter would confound representation quality with authoring quality, prior knowledge, placement, compression, model preference and evaluator preference. The former makes it possible to answer the actual scientific question: **which representational properties preserve the structure of supplied evidence, across heterogeneous receiving agents, at each constrained context budget?**