QuEST: Stable Training of LLMs with 1-Bit Weights and Activations Paper • 2502.05003 • Published 5 days ago • 36
Cottention: Linear Transformers With Cosine Attention Paper • 2409.18747 • Published Sep 27, 2024 • 16 • 5
EvoPress: Towards Optimal Dynamic Model Compression via Evolutionary Search Paper • 2410.14649 • Published Oct 18, 2024 • 8
Plus Strategies are Exponentially Slower for Planted Optima of Random Height Paper • 2404.09687 • Published Apr 15, 2024