0

NeoMME: an efficient Multimodal-native and Multilingual Encoder

https://huggingface.co/blog/Hcompany/neomme(huggingface.co)
A new family of multilingual multimodal encoders called NeoMME is introduced, available in 260M and 800M parameter sizes. Unlike models that use separate vision towers, NeoMME processes both text tokens and raw image patches within a single bidirectional Transformer. The model is trained from scratch with a masked discrete-diffusion objective, which forces it to learn image-grounded representations by reconstructing heavily masked text. A fine-tuned version, NeoMME-Retriever, is designed for efficient visual document retrieval and demonstrates competitive performance with a compact model size.
0 pointsby hdt1 hour ago

Comments (0)

No comments yet. Be the first to comment!

Want to join the discussion?