AN INQUIRY INTO UNPAIRED ROSETTA 10 OCT 2026

Do the shadows agree,
or the things
themselves?

Let us ask of Unpaired Rosetta: how much of the agreement belongs to what is seen, and how much to the models through which we see it?

THE WALL AND ITS SHADOWSFIG. 01
The wall and its shadows An imagined scene from Plato's cave. Fire behind a low wall casts the shapes of carried objects onto stone. A seated figure watches the shadows. A narrow opening leads toward daylight.

This scene is an allegory made for the inquiry. The measured evidence is given below.

Experiment paused · scope under reviewFixed protocol and work preservedNo EPIC result yet

01 / THE QUESTION

What follows from agreement?

When two shadows agree,
what have we learned?

The chosen encoders may share training data, objectives, or useful biases. If we repeat the solver but keep the encoders, we have not yet asked whether other models would agree.

A

Would other models agree?

A new alignment seed gives another trial with the same embeddings. To speak of other model families, we must choose them before seeing the results, then report every pairing and failure.

Not yet tested
B

Have they seen the same things?

Two datasets may bear different names and still contain the same images. The metadata shows that some COCO source images are described in Stanford Paragraph Captioning (SPC).

Metadata examined
C

The thing, or merely its kind?

To find a horse is not yet to distinguish this horse, its number, its place, or what it does. Retrieval must also be tested among things of the same kind, where broad categories no longer suffice.

Further trials proposed

THE FIRST CASE · PAUSED FOR REVIEW

Let the claim be no larger than the trial.

Begin with one case.
Ask what it proves.

Can one fixed image model and one fixed word model recover object-category correspondence on later kitchen recordings, when their alignment receives no paired episodes? That is the first question. Original DINOv1 and GloVe give us a case whose training sources we can examine.

WHY BEGIN HERE?

A narrower claim is easier to defend.

The training corpora predate these recorded episodes. Under the published training histories, memorizing these particular episodes cannot explain a result. This makes the case stronger against direct episode exposure. Familiar objects, ordinary phrases and the choices of the model builders remain part of the account.

The question is small; the present run is substantial. One model pair, one source of recordings and one primary measure keep the claim narrow. The frozen plan nevertheless calls for 9,392 clips, 37,568 frames and 17 fits. It is a first controlled case, not the smallest possible demonstration.

  1. 2012

    The images

    The ImageNet-1K collection, used without class labels by original DINOv1 ViT-B/16. DINO training record · ImageNet release

  2. 2014 and earlier

    The words

    Original GloVe 6B: Wikipedia 2014 and the older Gigaword 5 collection. We average its word vectors without further language training. Original GloVe release

  3. 2017

    The later episodes

    EPIC-KITCHENS-55 recordings and their human narrations. The recording years come from the per-video metadata. Recording dates

These are corpus and recording dates, not model publication dates. The argument depends on the documented training histories and publisher-supplied dates. A short phrase may occur in both old text and a later narration; that alone is not exposure to the later episode.

I

Keep the test apart.

The existing sample was fixed before alignment scores: 6,639 fitting clips, 1,127 development clips and 1,626 test clips. The seven test kitchens appear in neither other partition.

Selection checked
II

Withhold the correspondence.

The main fit receives images from 73 videos and text from 74 different videos. Known-pair mappings, shuffled-pair mappings, random maps and separate encoder probes help us interpret success and failure.

17 fits prescribed · paused
III

State what a match means.

The first measure asks whether the nearest description has the same annotated noun class. Each test kitchen receives equal weight; duplicate descriptions and exact ties are accounted for. Recognizing a cup does not yet distinguish this cup or what is done with it.

No benchmark scores

What could follow from the result?

Possible outcomes, not findings. Negative controls use shuffled pairs or random maps. Read differences with the prescribed uncertainty intervals and results for every seed.
What we observeWhat we may conclude
Unpaired retrieval exceeds the negative controls; paired retrieval also works.Evidence of recoverable object-category correspondence for this pair on later episodes. Shared concepts and model biases remain possible explanations for that correspondence.
Paired retrieval works; unpaired retrieval does not.The frozen features support a useful mapping with pairing information. The tested unpaired procedure has not recovered it under these conditions.
Neither paired nor unpaired retrieval works.The features, pooling, task or domain may be unsuitable. This outcome alone cannot identify the unpaired solver as the cause.

Unexpected behavior between controls also needs investigation. No outcome from this one pair establishes a universal representation, understanding of actions or syntax, or independence from model choice. Changing both the models and the benchmark cannot measure how much contamination affected the original COCO/SPC result.

The run is paused. Acquisition and its supervisor were suspended on 10 October 2026 before any full EPIC alignment scores. Existing data, caches and the fixed protocol are preserved. A proposed smaller pilot below has not run.

The controls have limits too. Paired and shared-episode controls use all fitting clips, a larger information budget than the main disjoint fit. Annotated action windows supply preprocessing supervision. These comparisons cannot isolate contamination by themselves.

Work record · 10 October 2026. Paused-run record · Code and checks. The earlier pilot below used different models and data.

EXAMINING THE SOURCES

Two names.
Some of the same images.

We joined the SPC, Visual Genome, and COCO metadata. In the released cross-dataset setup, some descriptions and images refer to the same source.

Read the source record
6,258unique COCO training images
described in SPC
3,335unique COCO validation images
described in SPC

This establishes shared sources. It does not show that the aligner was told which items were pairs. The effect on published performance is still unknown; this finding does not concern the separate disjoint-half COCO experiment.

02 / THE EARLIER PILOT

The first check held the models still.

Five conditions.
The same two models.

Keep DINOv2-B and MPNet fixed. Remove the known shared sources, then compare each removal with a random-removal control of the same size.

O / ORIGINAL

Begin with the original SPC population.

Keep the known validation and training source matches. This gives us the starting condition.

WHAT THIS COMPARISON ASKS

What do we observe before removing either kind of shared source?

19,561SPC paragraph rows retained
Starting condition
0Original population: 19,561 rows
Known validation source
3,337
Known training source
6,261
No known COCO match
9,963
Removed
0

These are the full populations, before the pilot’s 4,096-row cap. Here we count paragraph rows; above we count unique COCO images. Random controls use selection seed 0.

Ask each the same questions. Every arm uses the same three 2,048-query sets and complete 40,504-item gallery. Clean, all-population, and originally exposed queries are sampled separately.

Say only what was removed. “Source-disjoint” means known metadata identities only. Missing mappings, visual near-duplicates, and pretraining overlap remain unresolved. Random deletion matches size, not semantic composition.

03 / WHAT HAS BEEN SHOWN

Recorded · 10 October 2026

The trials ran.
What did they establish?

Each of the five conditions completed three alignment seeds. The reduced pilot shows that the runs finish, the exclusions are applied, and the saved results pass consistency checks.

Each fit uses 4,096 training rows per modality. These equal caps remove the full-population size contrasts. The scores therefore do not estimate the effects of the planned full-data exclusions.

Read the checks
15 / 15fits completed
92,160query records verified across 45 strata
  • Fixed query and gallery identities checked
  • Source-identity exclusions checked
  • Saved map, rank, and configuration hashes checked
  • Mean rank, median rank, and FOSCTTM recomputed from saved ranks

Verification did not recompute embedding similarities or Recall@k from scratch.

Examine all 15 reduced pilot runs

FOSCTTM below uses the same clean queries against the full gallery. Lower is better. These reduced runs check the procedure; their scores do not answer the scientific question. Every seed is shown.

Reduced pilot runs · DINOv2-B × MPNet
ArmFit seedStatusFOSCTTMFit time
Download all pilot measurements (JSON) ↓

04 / THE PROPOSED ORDER

Review first; then resume.

Take the next step
only as far as it warrants.

The experiment remains paused. The smaller pilot below is a proposal, not a new result or a silent change to the frozen protocol. Its sample and rules must be written before it runs.

  1. 01

    First establish that the features can answer the question.

    Use a small, separately specified pilot from fitting and development kitchens only. Check known-pair mappings, separate noun probes, repeated descriptions and chance performance. Choose clips by fixed identifiers, not by which videos arrived first. Keep the final test kitchens out of this pilot and publish its failures too. This checks feasibility; it does not establish unpaired alignment.

    PROPOSED
  2. 02

    Then make the unpaired test and report its whole result.

    Before held-out scoring, either retain the existing 9,392-clip protocol or publish a separate amendment with the reason for each change. Fix the sample, controls, settings and measures. Report failed fits and weak results alongside successes. The earlier protocol and evidence remain available.

    AFTER REVIEW
  3. 03

    Only then ask how far the result extends.

    Prespecify other encoder pairs to test model-choice dependence, regardless of the first pair’s outcome. Original ELMo, ImageNet-only MoCo-v2 and MAE are audited candidates; their weights and extraction still need checking. A further study must test distinctions within object classes, including number, relation and action. Read the candidate audit.

    LATER STUDY

Training from scratch is not required for this first, conditional chronology argument. If we require tighter control over exact-item exposure, we can lock the model files and then collect new human-recorded scenes and descriptions. That would require a new protocol and would still leave the question of model choice open.

THE RECORDS

Examine the account.

The protocol and evidence records are below. They describe the work as it stood on the date shown; they do not update as experiments run.