AN INQUIRY INTO UNPAIRED ROSETTA 10 OCT 2026

Do the shadows agree,
or the things
themselves?

Let us ask of Unpaired Rosetta: how much of the agreement belongs to what is seen, and how much to the models through which we see it?

THE WALL AND ITS SHADOWSFIG. 01
The wall and its shadows An imagined scene from Plato's cave. Fire behind a low wall casts the shapes of carried objects onto stone. A seated figure watches the shadows. A narrow opening leads toward daylight.

This scene is an allegory made for the inquiry. The measured evidence is given below.

15 / 15 pilot fits completeSaved-result checks passedThe question remains open

01 / THE QUESTION

What follows from agreement?

When two shadows agree,
what have we learned?

The chosen encoders may share training data, objectives, or useful biases. If we repeat the solver but keep the encoders, we have not yet asked whether other models would agree.

A

Would other models agree?

A new alignment seed gives another trial with the same embeddings. To speak of other model families, we must choose them before seeing the results, then report every pairing and failure.

Not yet tested
B

Have they seen the same things?

Two datasets may bear different names and still contain the same images. The metadata shows that some COCO source images are described in Stanford Paragraph Captioning (SPC).

Metadata examined
C

The thing, or merely its kind?

To find a horse is not yet to distinguish this horse, its number, its place, or what it does. Retrieval must also be tested among things of the same kind, where broad categories no longer suffice.

Further trials proposed

EXAMINING THE SOURCES

Two names.
Some of the same images.

We joined the SPC, Visual Genome, and COCO metadata. In the released cross-dataset setup, some descriptions and images refer to the same source.

Read the source record
6,258unique COCO training images
described in SPC
3,335unique COCO validation images
described in SPC

This establishes shared sources. It does not show that the aligner was told which items were pairs. The effect on published performance is still unknown; this finding does not concern the separate disjoint-half COCO experiment.

02 / THE EXPERIMENT

First, let us hold the models still.

Five conditions.
The same two models.

Keep DINOv2-B and MPNet fixed. Remove the known shared sources, then compare each removal with a random-removal control of the same size.

O / ORIGINAL

Begin with the original SPC population.

Keep the known validation and training source matches. This gives us the starting condition.

WHAT THIS COMPARISON ASKS

What do we observe before removing either kind of shared source?

19,561SPC paragraph rows retained
Starting condition
0Original population: 19,561 rows
Known validation source
3,337
Known training source
6,261
No known COCO match
9,963
Removed
0

These are the full populations, before the pilot’s 4,096-row cap. Here we count paragraph rows; above we count unique COCO images. Random controls use selection seed 0.

Ask each the same questions. Every arm uses the same three 2,048-query sets and complete 40,504-item gallery. Clean, all-population, and originally exposed queries are sampled separately.

Say only what was removed. “Source-disjoint” means known metadata identities only. Missing mappings, visual near-duplicates, and pretraining overlap remain unresolved. Random deletion matches size, not semantic composition.

03 / WHAT HAS BEEN SHOWN

Recorded · 10 October 2026

The trials ran.
What did they establish?

Each of the five conditions completed three alignment seeds. The reduced pilot shows that the runs finish, the exclusions are applied, and the saved results pass consistency checks.

Each fit uses 4,096 training rows per modality. These equal caps remove the full-population size contrasts. The scores therefore do not estimate the effects of the planned full-data exclusions.

Read the checks
15 / 15fits completed
92,160query records verified across 45 strata
  • Fixed query and gallery identities checked
  • Source-identity exclusions checked
  • Saved map, rank, and configuration hashes checked
  • Mean rank, median rank, and FOSCTTM recomputed from saved ranks

Verification did not recompute embedding similarities or Recall@k from scratch.

Examine all 15 reduced pilot runs

FOSCTTM below uses the same clean queries against the full gallery. Lower is better. These reduced runs check the procedure; their scores do not answer the scientific question. Every seed is shown.

Reduced pilot runs · DINOv2-B × MPNet
ArmFit seedStatusFOSCTTMFit time
Download all pilot measurements (JSON) ↓

04 / WHAT MUST WE ASK NEXT?

Proposed · not yet begun

Would the agreement hold
with other models?

Even if this pair survives the exclusions, we will still have examined only this pair. Further trials must ask how far the agreement extends.

  1. 01

    Repeat the comparison with the full data.

    Compare targeted exclusions with their size controls using the full solver profile. Set the uncertainty calculation and justify the expected precision before beginning. The proposed 180-fit budget and effect margin remain provisional.

    PROPOSED
  2. 02

    Choose other encoders before seeing their results.

    Try every pairing and report failures. ResNet-50 / MAE × Contriever / GTR is proposed, but both text encoders appear elsewhere in the paper. Families unused throughout the paper still need to be chosen for a strict holdout test.

    PROPOSED
  3. 03

    Ask what the training has taught them.

    Hold architecture fixed while varying the objective, disjoint training corpus, and independent pretraining seeds. Test retrieval within categories, including differences in count, attribute, relation, and who did what.

    PROPOSED

THE RECORDS

Examine the account.

The protocol and evidence records are below. They describe the work as it stood on the date shown; they do not update as experiments run.