Leaderboard

davanstrien/isl-finepdfs-ocr-results

Rankings are computed using Bradley-Terry MLE from pairwise comparisons judged by a vision-language model. The judge sees the original document image alongside two anonymised OCR outputs and picks the more faithful transcription. Browse the comparisons to see the evidence — and vote yourself to build a Human ELO column. Human votes are stored locally for this session only and will reset when the server restarts.

# Model Params Judge ELO 95% CI Wins Losses Ties Win%
1 dots.ocr 1.7B 1613 1513–1769 14 5 1 70%
2 DeepSeek-OCR 4B 1462 1341–1586 8 11 1 40%
3 GLM-OCR 0.9B 1425 1282–1534 7 13 0 35%

ELO vs Parameter Count

Smaller models can win on the right documents. Error bars show 95% confidence intervals.