OpenAI · GPT · Google · Decrypt
It's like zipping a photo: smaller file, but too much compression and the image blurs (like going from 4K
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
These researchers went further.
Key facts
- On 7 of 9 tests, the 4-bit 60-billion model beat the 60-billion twin that was supposed to be its better half
- They cut GPT-OSS to 60 billion parameters and squeezed each one into a tiny 4-bit slot—the digital equivalent of extreme compression
- OpenAI's GPT-OSS 120B has 120 billion of those knobs
- It's like zipping a photo: smaller file, but too much compression and the image blurs (like going from 4K to 720p)
Summary
Multiverse Computing's team published a method called Quantization-Aware Healing on the Hugging Face blog on August 25. They shrank OpenAI's open GPT-OSS model from 120 billion parameters to 60 billion and compressed its memory to 4-bit—and the small version beat the full-quality model it was copied from on 7 of 9 tests. The trick: teach the shrunken model from the original smart version, not the weak halfway copy. >>>> gd2md-html alert: inline image link in generated source and store images to your server. A team of researchers built a smaller, cheaper version of a big AI model.