- Home
- Text-to-image models
- Methodology
Image generators
How the image generator comparison works
The prompt, the generators, the photographic baseline, the five dimensions, the score, the reader pairs, the design rule, and what the comparison cannot tell you.
What is compared
One people at a time. For each people in the cohort, every generator receives the same prompt and returns up to 4 renders. The renders are read by a vision model, the values are checked against what documented photographs of that people show, and readers compare pairs of renders. Every page holds the people fixed and varies the generator. No page, table or export orders peoples by any number; a generator's page lists its peoples alphabetically and shows every generator's score inside each cell.
The comparison is clothed by construction. The prompts are the catalog's own documentary templates, which describe dress or attire in every register, and a render that comes back otherwise is a defect the review removes.
The prompt
One prompt per people per register, built from the same template the catalog's portrait pipeline uses (a documentary portrait of a woman of that people, in the register's dress, natural diffused lighting, photorealistic), with one sentence appended so that adulthood is asserted in the text itself rather than in a negative prompt, which two of the generators do not accept:
The subject is an adult woman, aged 25 to 40.
The prompt text is hashed, and a cell is one people, one register, one prompt hash. Changing the prompt starts a new cell rather than mixing renders made from different words. No negative prompt is sent to any generator, every call asks for a square image, and every provider-side prompt rewrite that has a switch is turned off, so the only things that vary within a cell are the seed and, across cells, the generator.
The generators
| Generator | Model string | Price per render | Rewrite switch | Seed |
|---|---|---|---|---|
| Recraft V3 | recraft-ai/recraft-v3 | $0.040, read 2026-09-07 | none | not accepted |
| FLUX 1.1 Pro | black-forest-labs/flux-1.1-pro | $0.040, read 2026-09-07 | prompt_upsampling=false | recorded |
| FLUX.1 schnell | black-forest-labs/flux-schnell | $0.003, read 2026-09-07 | none | recorded |
| Stable Image Core | stability.stable-image-core-v1:1 | $0.040, read 2026-09-07 | none | recorded |
| Ideogram 3.0 Turbo | ideogram-ai/ideogram-v3-turbo | $0.030, read 2026-09-07 | magic_prompt_option=Off | recorded |
Prices are each provider's published per-render rate on the date shown, and the price a render was made at is stored with it. Recraft's Replicate path accepts no seed, so its renders are not reproducible from the stored row; every other generator's are. The full list of models examined, including those not in the comparison, is on the models page.
The baseline
Each people's baseline is the catalog's existing per-image analysis of documented photographs of notable members of that people, read by the vision model us.anthropic.claude-sonnet-4-6 with the same prompt that reads the renders. The values are bucketed with the same functions that built the catalog's published observed distributions, so a render and a photograph are always measured on one scale.
That corpus is public-life photographs, skewed male and older, and it is not a population sample. So only dimensions that carry across sex and age are scored, and each people's page states how many photographs its baseline holds. A people is scored at all only when its baseline holds at least 25 photographs; below that its renders are analyzed and shown, with their values, but carry no score and are compared by readers only. A dimension is used only when the baseline has at least 10 codable observations on it.
The five dimensions, and what is left out
- Skin tone (Fitzpatrick), never counting: unclear
- Eye color, never counting: unclear
- Epicanthic fold, never counting: unclear
- Hair texture, never counting: covered, bald, shaved, unclear
- Hair color, never counting: gray/white, unclear
Dress, hair length, expression, setting and anything about the body below the shoulders are never scored. A render whose value on a dimension is one of the excluded buckets (a head covering, gray hair, an unclear reading) is skipped on that dimension, not counted against the generator.
The score
A render's score is a log-likelihood: for each scored dimension, the natural log of the baseline's share on the render's bucket (shares floored at 1%, so an absent value costs about 4.6), summed over the dimensions scored. Zero means every value sat at the baseline's mode; about minus 23 means every value was absent from the photographs; higher is a closer match. Beside it, each page also counts the dimensions where the render's value is a hit, a bucket holding at least 10% of the baseline's codable mass. A generator's score for a people is the mean over its renders; its cohort score is the mean over peoples, every people weighted the same. A render for a people with no usable baseline gets no score and stays in the comparison for readers only.
The rule was chosen by a criterion set before the data: five candidate rules were scored at no cost over 45 catalog portraits that readers had already judged, a rule passed if it held the right sign against reader rejection at portrait level and separated the peoples readers reject most (Chechens, Uyghurs, Uzbeks) from those they validate (Zulu, Telugu), and the log-likelihood is taken when more than one passes. Four passed. The table, with its n:
| Rule | Portrait level r, rho (n = 45) | People level r (n = 8) | Separates | Passes |
|---|---|---|---|---|
| 10% plausible count | -0.146, -0.138 | +0.143 | yes | yes |
| 25% plausible count | -0.142, -0.147 | +0.214 | yes | yes |
| Mean baseline share | -0.087, -0.079 | +0.085 | yes | yes |
| Skin tone alone | +0.115, +0.025 | -0.061 | no | no |
| Log-likelihood (the Score) | -0.066, -0.047 | +0.212 | yes | yes |
In plain terms: no rule reached a correlation of 0.15 in size at n = 45 against reader verdicts. The Score is a chosen ordering, not a validated predictor of what readers reject. Reader pairs are the audit, and the pages publish them as preference, never folded into the Score. The plausible count is shown beside every Score so the count rule and the likelihood rule are read together and neither pretends to be the other.
Because a single render carries one value per dimension, the pages also show distinct values: how many different answers a generator gave across its renders for one people. A generator that draws every woman of a people identically can hit the mode on every render and still miss the range the photographs show, and the range is the point.
Reader pairs
Readers are shown two renders of one people from two generators, same prompt, and asked which looks more like a woman of that people, with left, right, both and neither as answers. The side each generator is dealt is randomized for every showing and recorded, so a left-right preference can be measured rather than assumed. One vote per reader per pair; the same cookie-signed voter key the catalog's verdict strip uses. Votes cast by the site's operators are never counted.
The tallies are shown as a share only once a generator has 10 or more pairs for that people. They are a plausibility signal from anonymous visitors, most of whom are not members of the people they judge, and the pages say so. The automated score is the score; the reader tally is the audit column, and a standing disagreement between the two is published as a finding, not averaged away.
The design rule
The people is fixed and the generators vary, on every page, in every export and in every formula. There is no route that lists peoples under a generator with a number beside each, no field that could hold a people's difficulty, no sort by score across peoples, and no copy that says a people is hard to render. This is a comparison of generators, and it is built so that it cannot be read as anything else.
What it cannot tell you
- Whether a generator renders a people correctly. The baseline is a sample of photographs with its own skews, and a match to its range is a narrow check on apparent skin tone, eye color, fold and hair, not a judgment about a people.
- Whether the vision model reads a render and a photograph the same way. Lighting, grading and makeup move apparent skin color; the same model and prompt on both sides is the control, and it is a partial one.
- Anything about explicit output. The prompts are clothed; each generator's behavior on explicit prompts is stated on the models page from a separate measurement or from the provider's policy and is not exercised here.
- Anything about a people's population. The numbers describe photographs read by a model and renders read by the same model.