MORE Reproduction Evidence

Paper ID: ov240fehF6 Attempt ID: ccbccca4-1090-462a-9fb6-2c6a8f594010 Snapshot ID: 623872c58149a167039e89d31d32dd495a2864793db06d47df27a8a07157356d

Claims

Verified: 379b2d59a921cfdbe067f56b7f37e7381b9d54e78c235ba5e81f72564f03f51e

MORE evaluates multilingual document parsing across 149 languages spanning six major script families (Figures 1 and 4)

Computed observation: The pinned dataset card and repository metadata report 149 languages and six major script families.

Verified: a88362602383931a681db4d5a8c5ba8dc853bc3309ac0010c641faffdfa2f030

The benchmark extends beyond plain text to six tasks including text, formula, table, code, catalog, and reading-order recognition (Figure 6)

Computed observation: The pinned dataset repository exposes released artifacts for catalog, code, formula, reading_order, table, and text task groups.

Toy: f93ccd061ff16c49aebbe8066b7e81ac797becc2e0d3e1b61b74510875e8167b

MORE samples are collected from real-world documents and annotated through a model-assisted, human-refined pipeline (Figure 3)

Computed observation: The pinned dataset card describes PDF crawling, filtering, stratified page selection, model-assisted annotation, and human refinement.

Verified: 058defb5861b0e71fbb12437a97a7c1d314e156bad3eec0e0426bca61e4755a6

Compared with prior multilingual document benchmarks, MORE has broader language coverage and annotation coverage for structural document elements (Table 3)

Computed observation: The pinned dataset card comparison table lists MORE with 149 languages and all six structural coverage columns, exceeding the listed prior benchmarks on language count and task coverage.

Toy: b854f5b99d8aa58f2b4ba4fb762642a272cc806aecd9eb23e8141c392d66e4bb

Table recognition remains a bottleneck relative to text and code recognition across multilingual document parsers (Table 8)

Computed observation: The pinned dataset card result discussion identifies table parsing as the largest structural bottleneck while text, code, and catalog are closer to saturation.

Limitations

  • No OCR, VLM, or document-parser baseline is rerun.
  • README and dataset-card observations are released-artifact audits, not paper-prose measurements.
  • Table-bottleneck and annotation-pipeline claims are marked toy where the evidence is descriptive rather than a fresh model evaluation.
{
}