davanstrien/isl-finepdfs-ocr-results
Rankings are computed using Bradley-Terry MLE from pairwise comparisons judged by a vision-language model. The judge sees the original document image alongside two anonymised OCR outputs and picks the more faithful transcription. Browse the comparisons to see the evidence — and vote yourself to build a Human ELO column. Human votes are stored locally for this session only and will reset when the server restarts.
| # | Model | Params | Judge ELO | 95% CI | Wins | Losses | Ties | Win% |
|---|---|---|---|---|---|---|---|---|
| 1 | dots.ocr | 1.7B | 1613 | 1513–1769 | 14 | 5 | 1 | 70% |
| 2 | DeepSeek-OCR | 4B | 1462 | 1341–1586 | 8 | 11 | 1 | 40% |
| 3 | GLM-OCR | 0.9B | 1425 | 1282–1534 | 7 | 13 | 0 | 35% |
Smaller models can win on the right documents. Error bars show 95% confidence intervals.