Thai University RankingsRESEARCH RADAR
← Back to research database
มีศักยภาพระดับโลก

Open-Weight Multimodal LLMs Versus Manual Data Entry for Legacy ERP Digitization: A Comparative Evaluation of Accuracy, Cost, and Verifiability

IMPACT SIGNAL79/100
01

Information from the abstract

Decades-old enterprise-resource-planning (ERP) systems lock operational data inside unstructured, human-readable reports, forcing slow, costly, error-prone manual re-keying. Because multimodal large language model (MLLM) capability is uneven, deploying MLLMs for extraction means trusting outputs without a labeled reference. We test this with a within-document controlled experiment on 400 controlled-substance stock-ledger documents (2951 records, 11 fields, predominantly Thai) from a Thai pharmaceutical factory, comparing trained human double-entry against four open-weight MLLMs (2 × 2 design: vendor × architecture) via OpenRouter. Human double-entry left 14 discrepancies against the adjudicated gold standard, none common to both operators. The strongest model, Qwen3-VL-32B-Instruct (Dense), reached 93.95% cell accuracy; among these four models, field accuracy varied more across vendors, whereas structural completeness differed consistently between dense models (0 missing records) and Mixture-of-Experts models (up to 51 of 2951 dropped). Deterministic accounting invariants flagged 0.61% of its records, leaving the unflagged majority 94.1% accurate across all 11 fields; adding calendar rules flagged 4.61% and raised residual date accuracy from 92.1% to 95.9%. We report both operating points and recommend the extended level where date fidelity is regulatory-critical. The pipeline is 13.5–29.4× faster in wall-clock terms and 97.5–99.5% cheaper. Gold-free, rule-based verification thus locates where MLLM reliability holds, giving human–AI collaboration quantified, disclosed residual risk rather than an implied guarantee. Even at the more conservative operating point, unflagged records average 94.5% accuracy across all 11 fields but only 43.7% on the free-text Remarks field, which the triage cannot check; the results support risk reduction and the localization of review effort, not unrestricted regulatory reliability across all fields.

02

Why this record is monitored

This record has an Impact Signal of 79/100 based on recency, source, collaboration, and bibliographic signals. It prioritizes monitoring and is not a judgment of research quality.

Related topics: Topic Modeling · Biomedical Text Mining and Ontologies · Natural Language Processing Techniques

03

Thai researcher and institutional participation

Chacharin Lertyosbordin · King Mongkut's Institute of Technology Ladkrabang

04

Data limitations

This page is a bibliographic record based on abstract-level information, not a full analysis or quality assessment. Verify the DOI and original article before citation.