0
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b(huggingface.co)Liquid AI has released LFM2.5-VL-3B, a vision-language model optimized for on-device and real-time applications. This model introduces significant improvements in screen/UI understanding, object grounding, multi-image reasoning, and function calling. It combines a SigLIP2 vision encoder with a pre-trained LFM text model backbone, trained on 34T tokens and fine-tuned with supervised and reinforcement learning techniques. The model demonstrates leading performance in its size class across a wide range of vision and text benchmarks, including document understanding, object detection, and UI interaction. It is designed to answer directly for faster responses in edge applications.
0 points•by chrisf•51 minutes ago