MORE Reproduction Evidence
Paper ID: ov240fehF6
Attempt ID: ccbccca4-1090-462a-9fb6-2c6a8f594010
Snapshot ID: 623872c58149a167039e89d31d32dd495a2864793db06d47df27a8a07157356d
Claims
Verified: 379b2d59a921cfdbe067f56b7f37e7381b9d54e78c235ba5e81f72564f03f51e
MORE evaluates multilingual document parsing across 149 languages spanning six major script families (Figures 1 and 4)
Computed observation: The pinned dataset card and repository metadata report 149 languages and six major script families.
Verified: a88362602383931a681db4d5a8c5ba8dc853bc3309ac0010c641faffdfa2f030
The benchmark extends beyond plain text to six tasks including text, formula, table, code, catalog, and reading-order recognition (Figure 6)
Computed observation: The pinned dataset repository exposes released artifacts for catalog, code, formula, reading_order, table, and text task groups.
Toy: f93ccd061ff16c49aebbe8066b7e81ac797becc2e0d3e1b61b74510875e8167b
MORE samples are collected from real-world documents and annotated through a model-assisted, human-refined pipeline (Figure 3)
Computed observation: The pinned dataset card describes PDF crawling, filtering, stratified page selection, model-assisted annotation, and human refinement.
Verified: 058defb5861b0e71fbb12437a97a7c1d314e156bad3eec0e0426bca61e4755a6
Compared with prior multilingual document benchmarks, MORE has broader language coverage and annotation coverage for structural document elements (Table 3)
Computed observation: The pinned dataset card comparison table lists MORE with 149 languages and all six structural coverage columns, exceeding the listed prior benchmarks on language count and task coverage.
Toy: b854f5b99d8aa58f2b4ba4fb762642a272cc806aecd9eb23e8141c392d66e4bb
Table recognition remains a bottleneck relative to text and code recognition across multilingual document parsers (Table 8)
Computed observation: The pinned dataset card result discussion identifies table parsing as the largest structural bottleneck while text, code, and catalog are closer to saturation.
Limitations
- No OCR, VLM, or document-parser baseline is rerun.
- README and dataset-card observations are released-artifact audits, not paper-prose measurements.
- Table-bottleneck and annotation-pipeline claims are marked toy where the evidence is descriptive rather than a fresh model evaluation.