Google Unveils TurboQuant to Compress and Accelerate Large Language Models
Google Research has introduced TurboQuant, a new AI-compression algorithm designed to significantly reduce the memory needs of large language models while maintaining their quality and increasing speed. Early tests report up to an eightfold performance boost and a sixfold reduction in memory usage, without degrading model output.