HomeModelsPricingDocsBlogCommunityLoginSign up
Models
Pricing
Sign upLog in

lmms-lab/
OCRBench-v2

Generate
DataBranchesEvaluationsFine-tune
Copyright © 2026 Oxen Labs, Inc., All Rights Reserved
CareersTrust CenterPrivacy PolicyTerms and Conditions
OCRBench-v2
public
images
5.3 gb
add initial dataset
2 yrs ago
test.py
174 B
initial commit
text
1 yr ago
default_test_english.parquet
2.4 mb
adding render fn images for the english data
tabular
2 yrs ago
README.md
2.8 kB
Update README
text
2 yrs ago
default_test.parquet
2.8 mb
Detecting language of all the questions
tabular
2 yrs ago
dataset_metadata.md
464 B
add initial dataset
text
2 yrs ago
Last commit cannot be located
About

OCRBench is a comprehensive evaluation benchmark designed to assess the OCR capabilities of Large Multimodal Models. It comprises five components: Text Recognition, SceneText-Centric VQA, Document-Oriented VQA, Key Information Extraction, and Handwritten Mathematical Expression Recognition. The benchmark includes 1000 question-answer pairs, and all the answers undergo manual verification and correction to ensure a more precise evaluation.

6 commits
1 contributor
0 downloads
4.5 gb
0 stars
Project contents
image > 99%
text < 1%
tabular < 1%
5.3 gb
210K3
Contributors
Ox Data Bot 🤖
@oxbot
OCRBench-v2/
1 branch