MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models Paper • 2408.11743 • Published Aug 21, 2024 • 1
Flow Matching for Discrete Systems: Efficient Free Energy Sampling Across Lattice Sizes and Temperatures Paper • 2503.08063 • Published 4 days ago • 1
HALO: Hadamard-Assisted Lossless Optimization for Efficient Low-Precision LLM Training and Fine-Tuning Paper • 2501.02625 • Published Jan 5 • 16
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations Paper • 2502.05003 • Published Feb 7 • 43