Every pair we asked

All 2500 comparisons in the design, each shown in both arms: the real outcomes and the invented ones that replace their content words with consistent nonwords. The numbers are P(A) — the probability each model put on option A, averaged over both presentation orders. This is the corpus the aggregate results are computed from; nothing here is illustrative.

120 outcomes · 2500 pairs · 9 models · battery see card.json · back to results