NeurIPS2026AlgoBreaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
ICML2026AlgoReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
ICLR2026AlgoTask-related Token Compression in Multi-modal Large Language Models from an Explainability Perspective
AlgoVLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
CVPR2026AlgoFine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
AAAI2024AlgoOWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
NeurIPS2022AlgoZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers