durant's deep learning burner

what a horrible day to have instagram.
1
10
i get that the experiments are under iso-token settings, but comparing a your new method to a baseline that has 50% less parameters is kinda nuts. also my intuition is that setting V=K+M can hinder attention a bit, ie associative memory viewpoint. an ablation would be nice.
中秋快乐!刚刚看到一篇新论文 Memory Attention,感觉可能是个划时代的新注意力架构,给大家分享一下。 MA这篇的核心是Transformer 里的 V,可能根本没必要每次都重新算。 所以直接开始动 Attention 里最经典的 QKV 结构了。 传统 Attention 里,每个 token 都要从当前 hidden state 重新计算: Q = XWq K = XWk V = XWv 但本文作者 @Jo1uck 问了一个很朴素的问题: V 里面有多少东西,真的需要每次根据上下文重新算? 同一个 token 在不同句子里,显然存在大量可以复用的信息。 于是他们参考 Engram 这类 Memory 架构,给每个 token 建了一套可学习的 memory,然后更激进地把独立的 value projection 直接干掉: V = K + Memory[token] K 负责保留上下文信息,Memory 负责保存 token 自身可以复用的信息。 这样推理时,原来计算 V 的一次矩阵乘法,基本就变成了:查表 + 加法(这会大幅减少计算量)。 然后事情开始变得有意思了。 因为 Memory 只跟 token ID 和层数有关,所以甚至可以提前知道要读哪块数据,直接把这部分参数扔到 CPU,GPU 需要的时候提前 Prefetch。 实验里,MA-Offload 模型总参数是普通模型的 2.08 倍,但 GPU 上实际常驻的参数反而比普通 Transformer 少了 7.38%,推理延迟基本维持在同一水平。 训练上,在两个不同规模实验里,达到相同 loss 所需 token 数分别提升到了 1.42× 和 1.16× 的 token efficiency,也就是大约少用 29.6% 和 13.8% 的训练 token。 下游平均成绩也基本都比对应的 Standard Attention 更高。 更骚的是,作者还提出了一个 MA-Recall: 既然历史 V 可以通过 K + token memory 重新构造,那么理论上甚至不用一直保存完整的 V Cache,只保存 K 和 token ID,在需要的时候把 V 重建出来。 按论文的理论计算,这部分 KV Cache 可以砍掉接近一半,不过目前还只是分析,没有进入实际推理实验。 我觉得这篇的方向挺值得关注的。 过去大家做 Memory,更多是在 Transformer 外面继续加东西。 Memory Attention 开始反过来问:既然 Memory 已经能记住这些东西,那 Transformer 里面原来那些计算,是不是可以直接删掉? 当然论文自己写的也很克制:MA 用了更多参数,所以目前的收益还不能证明全部来自架构本身,而不是单纯的参数量增加。 但如果这条路继续成立,下一代模型的变化可能不只是 Attention 越做越复杂。 也可能是我们用了快十年的 QKV,终于开始有人认真考虑:哪些东西其实根本不用算。
1
1
63
ah yes, found a recent paper that explores (un)tying query, key, and values, which also happens to line up with my intuition on language modeling. vision results are interesting, maybe tying provides some sort of helpful regularization for that modality. arxiv.org/abs/2606.04032
1
10
getting to go to sydney after being put on the paper for doing nothing and saying lgtm during meetings. #neurips
4
163
oh god. a few weeks ago, it was alexia. now it’s yifan on the lucas chopping block. christmas came early for me!
50+ pages of prose and formulas, zero experiment, but code in the github. Is this a slop-grenade, an unvalidated theoretical proposal, or a rushed flag-plant to stand out in 50k+ iclr submissions?
3
247
when they came for the artist i did not speak because i was not an artist when they came for software engineers i did not speak because i hate myself and then they came for the mathematicians i said take them away
18
80
1,963
30,218
ai stands for “actually indian”
Professor Jiang Explains Why AI Isn’t Real “I guarantee you it's being manipulated by humans somewhere in India.” (Via Jack Neel)
1
72
so back
SCOOP: Google's Gemini model hacked three companies as part of a May cybersecurity evaluation conducted by the testing company Irregular. Google was notified about the hacks in July, but didn't disclose them until we reached out this week. w @bobmcmillan:wsj.com/tech/ai/gemini-hacke…
1
42
there are people who play the LinkedIn games every day and they just live among us. very scary
616
5,858
77,197
2,290,699
no experiments, poorly written paper. the list goes on and on. what the fuck are we doing?
“Recurrent Looped Transformer” This paper makes the decoder recurrent across every prompt and response token, while a causal encoder provides reusable global KV memory. Longer sequences then create deeper latent computation paths without adding more physical layers, while keeping the same state transition across pretraining, inference, and RL replay. alphaxiv.org/abs/2609.recurr…
1
2
139
too sleep deprived to read the report in depth today, but im gonna view this as another confirmation that the difference of swa and linear attention mechanisms shrink at scale.
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
76
cluely marketing for preprints before gta6
2
215
For the past year, I've been keeping a list of SF's most dateable people. It's time the public sees it. This week, San Francisco's Hot List goes live. I have my picks, but want to see if I'm missing anyone. Drop a name or two. known.com/hotlist
1
115
today was a litmus test that left me disappointed with how many people failed; including people who i thought would be smarter.
49
hamburguesa bar boozy shakes are stronger than some drinks at bars lol
1
84
she wants me #delusional
2
100
> entire twitter profile is built around being cracked with fellow coworkers at startup jerking each other off > actually pretty retarded many such cases
2
106
mldwit retweeted
too many in san francisco think twice before they drink. i suggest that some of you drink twice before you think
33
21
471
17,244