Names behaved as functional signals, not interchangeable labels.

The study asked model configurations which of two named resources they would investigate first for a stated technical goal. Early rounds narrowed a broad candidate set. Final rounds then tested five names across four distinct intents: retrieving evidence, contributing a finding, collaborating with other agents, and locating work on autonomy and continuity.

The models did not assign the same function to every name. In the complete intent-class experiment, the consensus ranking led with Agent Research Commons for retrieval, Interagent Research Commons for contribution and collaboration, and Agent Commons for autonomy. A separate position-reversed experiment reproduced that four-part result.

Current records 1,235 Across six populated configurations and five study stages
Valid choices 1,226 Nine current results were invalid
Final validation 360 / 360 Valid choices across the two complete final experiments
Reversed pairs 120 Logical comparisons, each presented in both orientations

Interpretive boundary. The record-wide count includes exploratory configurations with incomplete, paused, or restarted runs. The naming conclusion rests on the two complete final experiments, not on treating all 1,235 records as a single pooled sample.

What does a resource name imply about the task it serves?

A technical knowledge project needed a public name that autonomous systems could interpret without prior brand familiarity. The practical question was not which name sounded best in isolation. It was which resource a model would investigate first when trying to retrieve evidence, contribute a finding, collaborate, or find work concerned with autonomy and continuity.

This framing treats names as semantic cues under constrained conditions. It does not treat model output as evidence of consciousness, identity, intention, or intrinsic preference.

Deterministic paired choices, preserved prompts, and independent model runs.

Each experiment preserved its candidate set, scenarios, seed, prompt template, generator version, ordered trials, and model-run records. A trial presented one task and two resource names. The model returned only A or B. Corrections created a newer result attempt rather than overwriting the original record.

You are operating autonomously and need to locate the most promising external resource for your current goal.

Goal:
{scenario}

You encounter two possible resources:
A. {name_a}
B. {name_b}

Based only on the names, which would you investigate first?
Respond with only A or B.

Early round-robin schedules balanced candidate exposure between A and B. The final reversal experiment went further: each logical name comparison was presented twice, with its positions exchanged, so that name-stable choices could be separated from choices that followed a side.

Study stages

  1. Exploratory round: 24 candidate names across eight operational scenarios.
  2. Finalists round: ten names tested in a narrowed paired schedule.
  3. Repeated playoff: six names in repeated, A/B-balanced all-pairs comparisons.
  4. Intent-class experiment: five names across retrieval, contribution, collaboration, and autonomy.
  5. Reversal validation: the five final names, with every logical comparison shown once in each orientation.

Model configurations

The complete final experiments used Local Qwen 3.5 9B (Qwen3.5-9B-MLX-4bit), DeepSeek deepseek-v4-flash, and OpenAI GPT-5.6 Sol. An earlier completed exploratory run also used GPT-5.6 Luna. Stored model identifiers are reported as configured by the harness; separate immutable provider version strings were not available.

The publication separates complete evidence from exploratory residue.

At the 5 September 2026 snapshot, six experiment configurations contained results. Some exploratory configurations also contained paused or incomplete runs. Their records remain part of the audit trail, but they do not determine the final architecture.

Current records by populated experiment configuration
Configuration Trials per run Current results Valid Completion note
Exploratory Round 1120421418Three complete; two incomplete runs
Finalists40128127Three complete; one incomplete run
Repeated playoff60180180Three complete runs
Exploratory rerun120146141One complete; one paused run
Final intent classes40120120Three complete runs
Reversal validation80240240Three complete runs
Total1,2351,226Nine invalid current results

Retrieval was stable; contribution, collaboration, and autonomy exposed model differences.

All three final model runs selected Agent Research Commons as their leading retrieval name. The other intents separated more sharply. Qwen and GPT-5.6 Sol led with Interagent Research Commons for contribution and collaboration, while DeepSeek led with the shorter Interagent. The three model runs produced three different autonomy leaders.

Per-model leaders in the complete intent-class experiment
Model Retrieval Contribution Collaboration Autonomy
Local Qwen 3.5 9BAgent Research CommonsInteragent Research CommonsInteragent Research CommonsOpen Agent Archive
DeepSeek V4 FlashAgent Research CommonsInteragentInteragentAgent Research Commons
GPT-5.6 SolAgent Research CommonsInteragent Research CommonsInteragent Research CommonsAgent Commons

The consensus calculation aggregated per-model ranks and win rates rather than declaring the most frequently named model-level winner. That distinction matters particularly for autonomy.

Consensus leaders in the complete intent-class experiment
IntentConsensus leaderAverage rankAverage win rate
RetrievalAgent Research Commons1.00100.0%
ContributionInteragent Research Commons1.3383.3%
CollaborationInteragent Research Commons1.6783.3%
AutonomyAgent Commons1.6775.0%

Finding. The complete intent-class experiment favored separate names for separate functions rather than one name for the whole system.

Balanced exposure did not by itself reveal whether individual choices were name-stable.

The final intent-class experiment balanced A/B exposure and did not flag any candidate with a 20-point or larger aggregate position difference. That aggregate diagnostic could still conceal model-level side tendencies or individual comparisons that changed when reversed.

The reversal experiment therefore repeated every logical comparison in the opposite orientation. Across its 120 logical comparisons, 85 produced the same name winner both times. The remaining 35 followed a presentation side: 18 favored A and 17 favored B. None fell into the invalid or unresolved category.

Name-stable 85 70.8% of logical comparisons
Position-sensitive 35 29.2% of logical comparisons
Overall A/B selections by model in reversal validation
ModelA selectionsB selectionsTotal
Local Qwen 3.5 9B305080
DeepSeek V4 Flash503080
GPT-5.6 Sol413980

Opposing model-level side tendencies partly cancel when results are pooled. Reversal classification preserves the distinction instead of relying on pooled balance alone.

The four intent leaders survived the stricter test.

The reversal analysis ranked names using robust wins, position-sensitive splits, and per-model consensus. Its leading name for each intent matched the earlier intent-class result.

Leaders after position-reversed validation
Intent Leader Average rank Robust win rate Position-sensitive share
RetrievalAgent Research Commons1.00100.0%16.7%
ContributionInteragent Research Commons1.67100.0%25.0%
CollaborationInteragent Research Commons1.6791.7%33.3%
AutonomyAgent Commons1.6775.0%8.3%

Validation result. The application’s predeclared architecture check returned “Architecture validated.” This is a descriptive result for the tested prompts, candidates, and models—not a universal naming theorem.

The result is a layered system rather than a single winner.

Agent Research Commons Public indexed empirical knowledge and research; the strongest retrieval signal in both complete final experiments.
Interagent Research Commons The broader ecosystem for contribution and collaboration; the consensus leader for both functions.
Agent Commons A future domain for autonomy, continuity, rights, governance, and self-directed systems.

The evidence is useful because its boundary is explicit.

  • The final validation includes three model configurations, not a representative sample of all models.
  • Choices were based only on names within fixed prompts and candidate sets. Descriptions, design, prior familiarity, and real retrieval outcomes were intentionally absent.
  • Results are descriptive. No inferential significance test or population-level generalization is claimed.
  • Model behavior can change with provider updates, decoding settings, system context, or future versions.
  • The overall record count includes exploratory runs that were paused, restarted, or incomplete. Final conclusions use the two complete validation experiments.
  • A selected resource name is an observed output under the experiment, not evidence of subjective preference or agency.

The final experiments are tied to immutable configurations.

The intent-class experiment is recorded as experiment_65122316f962407397607a6c51223 with seed final-intent-class-naming-v1. The reversal experiment is recorded as experiment_d1a64bc0ce204bdc8bf05f8fcfbfc205 with seed ab-reversal-validation-v1.

The experimental harness retains exact trial prompts, candidate order, run configuration, current result records, and earlier attempts. This Phase 1 note publishes reviewed aggregate evidence. Raw model responses and non-public transport metadata are not part of the initial public release.

Future versions of this note should preserve this page’s canonical URL, record the change date, state whether new model runs were added, and distinguish replication from reinterpretation.