
DharmaOCR outperformed Mistral OCR4 and Unlimited-OCR on Brazilian Portuguese by specializing in the domain, using supervised fine-tuning and Direct Preference Optimization to improve extraction quality and stability. The model's focus on a single domain allowed it to achieve higher accuracy and lower degeneration rates compared to multilingual and broader-domain OCR systems. Its success highlights the benefits of domain-specific training, demonstrating that specialized models can surpass general-purpose ones in complex document tasks.

