Self-Pruning Transformer: Universal Attention을 활용한 극단적인 KV-Cache 압축A Self-Pruning Transformer: Extreme KV-Cache Compression w/Universal Attention▲ 6 · arxiv.org · 1일 전 · 0 댓글원문 보기 → HN에서 보기 →원문 요약원문을 요약하고 있습니다…