0
LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
https://huggingface.co/blog/LiquidAI/qad(huggingface.co)New 4-bit quantized model checkpoints (Q4_0 GGUFs) for the LFM2.5 series have been released using a technique called Quantization-Aware Distillation (QAD). This method distills a high-precision teacher model into a quantized student, allowing it to run with the memory and speed benefits of a 4-bit model without the typical quality degradation. Benchmark results show the QAD models recover over 96% of the performance lost to standard quantization across various reasoning and instruction-following tasks. These efficient models are designed for edge hardware like mobile phones and Raspberry Pi, and are available on Hugging Face for use with runtimes like llama.cpp.
0 points•by will22•1 hour ago