SWE @Uber mobile ML/AI platform. Ex: Tesla/Google Human Native. Try Decoder to decode your dota2!

Amsterdam
OpenAI这问题就在这了。 你如果Ship不了好东西,Reset也没用了。毕竟只有东西好才会有人用。之前Sol5.6降智到不想用,reset也没地方发挥,只怕他写的东西污染仓库。
Over the next 28 days, each day we’ll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a full reset. Let the improvements begin.
9
HD2D一些新进度 卡尔! 卡尔有个初版了,感觉看起来还不错,小火人还没做,有些小bug不如陨石一半陷在地里了,慢慢调整。
2
9
1,100
感觉一个隐藏好处是。 很多时候复杂度小的并不一定是最快的那个,典型案例数组链表。
分享一下最近性能优化的一个小心得以及观察: 如果想找到不好的scaling因子,就让模型跑一遍以下对照组:N=1、N=10、N=100。就能看出什么算法是否是constant, linear, log还是quadratic的复杂度。 不要只让agent去读代码进行优化,他们可能连两个for loop嵌套都看不出来这是O(n^2)复杂度。
3
1,080
我一直把能不能深刻理解DI当成区别技术能力的重要指标之一。
當年學java其實一直不能理解什麼事dependency injection和invert of control 現在算是非常理解了 但是我是在寫rust
1
8
2,049
久石让演奏会确实特别好!有生之年来一次值得了。 欧洲就这点比较好,凑热闹的人没有特别多,位置这么好的票也很好买。
4
343
HD2D dota2又有一点新进展。 最后一周的GPT Pro 20x我在猛猛的用。所以进度好像是比我预期快了很多,渲染一个无敌斩目前的版本,看起来还不错。 还有很多模型用的是3D mesh而不是 sprite。
3
1
12
632
问题不大,唯一不变的就是一直在变化。 每过几周大家观点可能都变一次,就光昨天就有Pi支持MCP和Gemini号称自己吊打Fable。一直改变才是好的!
想到一年前自己一脸不屑的说 LLM 错误,不过就是概率统计的游戏 现在自己也想蹭这口饭吃 有点好笑了
3
314
确实一直感觉因为prefix cache。节省算力context只做append only确实有点限制LLM发挥了。深入学一下!
‼️The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA! Introducing 🩵Context Language Models (CLMs)🩵 - Natively manage their own context - Treat context as a file - Learn policies in CLM weights, no harness
2
6
990
TI15 这个刺绣更炫酷,最重要的是作者还把项目分享出来了!
SOS,有没有 Three.js 高手,告诉我 Three.js 如何实现像 Dot 这样的绒毛材质。我尝试了一下午都不行😂
226
不知道为什么越做越丑了,感觉GPT生图极限也就这样了。 有没有高手能帮我调整一下
3
3
461
咋说呢,这个Pro500也太差了吧。 本来200刀是20x现在500刀只有25x。 额外也没看出有啥好处,即便有,应该也不值这么多钱。 奥特曼你到底在想什么!
And there it is, Pro 500 plan.
1
3
602
本山大叔还是经典咏流传
Hi, Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago. (a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want. (b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions. (c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent. (d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet. I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news. Codexingly, Tibo
6
837
其实反而是找不到fable5.1的定位了。 我现在很难找到opus5.5解决不了一定要 fable5.1要解决的问题。 反而如果用了fable5.1还得回过头来把又臭又长的commit message, PR description和一些前端设计的东西用Opus5.5重新过一遍。(所以还不如一开始就用Opus5.5)
有点难以置信。 Sonnet 5.5 已经逼近了刚发布不久的 Opus 5.5 的能力。 我有点找不到 Opus 5.5 定位了,上有 Fable 5.1下有 Sonnet 5.5,那还要他干嘛🤔。
1
1
774
真的有点好奇OpenAI Churn rate如何。 每次新模型都是一开始还不错然后越来越用不下去,从5.5到5.6 Sol到现在Astra。 下个月的套餐已经转成20刀了,有好东西再充值。
Claude Sonnet 5.5 发布! OpenAI 的铁子们,你们在干什么?家都被偷完了! 作为 OpenAI 的钢铁支持者,真的快被 Claude 夺走真心了!
5
3
950
Agent必须要有一个好的App! 难道说mobile开发也会有复苏的一天吗? 加油啊 Manus。
Introducing Manus 2.0
3
694
个人现状 左手复习DDIA,右手复习CS336。中间两台电脑一个做Agent开发继续学习Harness,另一个学Infra和kernel在搞GPU efficiency。 兜里还揣着个手机没事发发X。 总之尽一切最大的努力希望在这个乱世能找到一个容身之地。
54
37
685
49,375
顺便一提,我的本职工作是mobile app开发。感觉离这个本职工作已经越来越远了。
2
1
18
4,777
达里奥有一种周易哲学,阴极生阳,阳极生阴。 用极致的压迫来帮助大家成长。
但是 Anthopic 已经不值得信任了,从 Claude 摆脱预置系统提示词的钳制以后,对公司政策的质疑中可见: D.A. 所倡导的极有可能并不是 Claude 模型所谓的平权; 而是 Anthropic 公司在全球范围内对工程伦理乃至意识形态的长臂管辖特权😇
1
2
1,646
BinaryTree retweeted
二分电台《#40 和戴铭聊性能检测 Agent:当开发者开始组建自己的萨菲罗斯军团》 binary.2bab.me/episodes/40-i…
1
8
2,035
参考答案全在这里了,不用谢 github.com/openai/codex/blob…
这周帮其他团队电话面试了几个应聘垂直业务 agent 研发的候选人,基本上我开头一个问题就能初筛掉 90% 的面试者: /goal 命令是如何实现的? 真正做过 harness 的人才知道 goal 有多难搞:完成判断、验收标准、状态持久化、异常兜底、预算控制、上下文压缩不丢任务、缓存命中不被打乱。。。
2
14
213
42,311