Devlog

Measuring the ruler: 73 of 84 dimensions pass a test-retest study

A test-retest study over 196 portraits scored the head and neck dimensions, wrote the results into the vocabulary files, and made them citable.

  • Research
  • Data
  • AI

On August 31 we measured the phenotype vocabulary itself. Portraits already coded were read a second time by the same vision model and compared against their stored codings, dimension by dimension, across the ten head-and-neck vocabularies: 196 portraits from 29 groups. 73 of 84 rankable dimensions reached a Cohen's kappa of 0.60 or better.

The results live inside the vocabulary files. 86 dimensions now carry a reliability record that cites its study, and the ten files moved to version 1.1.0 with every value set unchanged, so a coding made against 1.0.0 reads identically against 1.1.0.

The vocabularies are citable. The pipeline repository is archived on Zenodo as v2.0.1 under concept DOI 10.5281/zenodo.20075616, which always resolves to the newest version, and the Hugging Face dataset was republished at schema 6 with the reliability record in its vocabulary table.

Under the hood

  • Dimensions are scored on Cohen's kappa, not raw agreement. Raw agreement rewards a dimension for having one dominant answer; kappa corrects for chance. Three dimensions that return a single value for nearly every portrait are marked uninformative instead of topping the table.
  • A decline is not an answer. Values such as unclear or not visible are scored as declines even where the vocabulary spells them as values, while absent and none count as findings: a nose with no dorsal hump is a result.
  • Every rate is reported on two denominators: 92.9 percent agreement over dimensions both readings answered, and 90.6 percent counting a one-sided decline as a miss.
  • Kappa turns fragile where one value dominates. Dimensions with high chance agreement carry an estimate_is_fragile flag with the measured evidence, and the marginal counts behind every kappa are published in the pipeline repository.
  • Schema 6 records, for each of the 196 dimensions, how well it reads from a single photograph (120 high, 41 medium, 6 low, 29 not assessable), the least framing that can carry it, and four reliability columns: kappa, n, verdict and study id.
dimensions at kappa 0.60 or better
73 of 84
portraits re-read
196
dimensions with a reliability record
86

See it on the site

Sources and open data

Read the full log