0
NeoMME: an efficient Multimodal-native and Multilingual Encoder
https://huggingface.co/blog/Hcompany/neomme(huggingface.co)A new family of multilingual multimodal encoders called NeoMME is introduced, available in 260M and 800M parameter sizes. Unlike models that use separate vision towers, NeoMME processes both text tokens and raw image patches within a single bidirectional Transformer. The model is trained from scratch with a masked discrete-diffusion objective, which forces it to learn image-grounded representations by reconstructing heavily masked text. A fine-tuned version, NeoMME-Retriever, is designed for efficient visual document retrieval and demonstrates competitive performance with a compact model size.
0 points•by hdt•1 hour ago