
写作日历
点击方块查看当天的文章
动态 · News
- 发布 Assignment 2:SystemsCS336 · 从头搭建一个大语言模型
- 发布 Assignment 1:BasicsCS336 · 从头搭建一个大语言模型
- 发布 Exp Chapter 2:从 Koopa IR 到 RISC-V从零开始的编译原理
- 发布 Chapter 2:语法分析从零开始的编译原理
- 发布 Exp Chapter 1:从 SysY 到 Koopa IR从零开始的编译原理
- 发布 Chapter 1:词法分析从零开始的编译原理
- 发布 Exp Chapter 0:实验环境搭建从零开始的编译原理
- 发布 GPU 体系结构 —— H100 的硬件架构现代GPU编程指南
- 发布 现代GPU编程指南(开篇):Hopper 新特性总览现代GPU编程指南
- 发布 AlgoSOAR: Scale Optimization for Accurate Reconstruction in NVFP4 QuantizationLLM
专栏
2026
CS336
Assignment 2:Systems
CS336
Assignment 1:Basics
从零开始的编译原理实验
Exp Chapter 2:从 Koopa IR 到 RISC-V
从零开始的编译原理理论
Chapter 2:语法分析
从零开始的编译原理实验
Exp Chapter 1:从 SysY 到 Koopa IR
从零开始的编译原理理论
Chapter 1:词法分析
从零开始的编译原理实验
Exp Chapter 0:实验环境搭建
现代GPU编程指南
GPU 体系结构 —— H100 的硬件架构
现代GPU编程指南
现代GPU编程指南(开篇):Hopper 新特性总览
AlgoSOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization
ACL2026AlgoARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
ICML2026AlgoVEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
NeurIPS2026AlgoBreaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
ICML2026AlgoReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
ICLR2026AlgoSERQ: Saliency-Aware Low-Rank Error Reconstruction For LLM Quantization
ICML2026AlgoOSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
ICLR2026AlgoSliderQuant: Accurate Post-Training Quantization for LLMs
ICLR2026AlgoTask-related Token Compression in Multi-modal Large Language Models from an Explainability Perspective
AlgoTowards Joint Quantization and Token Pruning of Vision-Language Models
AlgoQAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
AlgoVLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
CVPR2026AlgoFine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
CVPR2025AlgoMBQ: Modality-Balanced Quantization for Large Vision-Language Models
CVPR2026AlgoMASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
ICLR2026AlgoTurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
ICLR2023AlgoGPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
AAAI2024AlgoOWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
NeurIPS2022AlgoZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
ICML2020AlgoAdaRound: Up or Down? Adaptive Rounding for Post-Training Quantization
ICML2021AlgoAdaQuant: Accurate Post Training Quantization With Small Calibration Sets
Blog
浅谈低精度浮点数
Blog
Triton
Blog



