
LFM2.5-VL-3B is a high-capacity vision-language model optimized for on-device use, offering advanced capabilities in screen/UI understanding, object grounding, and multi-image reasoning. It achieves efficient inference on CPUs and GPUs, with performance reaching 20 tokens/s on a Galaxy S26 Ultra, and supports non-Latin scripts through a doubled vocabulary. The model outperforms competitors in vision and text benchmarks, enabling real-time, low-latency applications for tasks like document understanding and tool use. It is available on Hugging Face for deployment and fine-tuning, and is cited as a key advancement in edge AI for efficient vision-language processing.

