AI Models
Quantization-Aware Healing Makes a 4-Bit Model Beat Its BF16 Source
Multiverse Computing says its Quantization-Aware Healing recipe produced a compressed 60B MXFP4 model that matched or beat its recovered BF16 source on seven of nine benchmarks. The result came with roughly four times lower weight memory, but important limitations remain.