MonkeyOCRv2 with 0.7B Parameters Tops Open-Source Document OCR, Surpassing 3B Models
The team led by Xiang Bai at Huazhong University of Science and Technology released MonkeyOCRv2, with only 0.6B and 0.7B parameters, achieving scores of 82.5 and 83.3 on the MDPBench benchmark covering 17 languages, surpassing the previous open-source leader dots.mocr (3B parameters, 80.5 score). By enhancing the visual encoder to better capture characters and layout, the model reduces the language model's burden, enabling high performance with a smaller size.