现代 LLM 的 Memory-Compute Trade-off:从 MHA 到 MLA、Sparse 与 Linear Attention2026年9月11日·LLMsAttentionMLADeepSeek长上下文