Mistral · Hugging Face
They are the kind of models that raise the competitive standard for what OCR systems are expected to deliver
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
When Hugging Face ran both against the DharmaOCR benchmark, an evaluation designed exclusively around Portuguese, the results were conclusive.
Key facts
- Mistral OCR4 falls approximately 13 points below DharmaOCR; Unlimited-OCR falls more than 16 points
- Figure 1: ENEM essay manuscript used in benchmark evaluation with outputs of the each model
- Mistral OCR4, evaluated on documents of this kind, transcribed the name Chico Buarque (one of Brazil's most widely recognized musicians and poets) as "Chico Barque
- ENEM essays (Brazil's national high school examination) combine handwritten text with vocabulary, proper nouns, and cultural references that are specific to Brazilian Portuguese
Summary
Three months ago, they published a paper on DharmaOCR and open-sourced one of the models. The training pipeline was built in two stages. The first was a supervised fine-tuning step, drawing on a broad collection of Portuguese-language files from different sources, formats, and levels of complexity. The combined result was a model that achieved the highest extraction quality score with the lowest degeneration rate on a Portuguese-focused benchmark. OCR models have been moving quickly. The proliferation of multimodal generative models made language model-based OCR widely accessible, and the wave of fine-tuned OCR variants that followed reflects how fast that adoption has moved. That proliferation has not, however, changed the fundamental character of the technology.