95后老登 | 💰全仓房贷车贷传宗接代 | ⏰全仓AI

抽纸盒子快见底了 超市清单先加上一提 别等拧空卷干瞪眼
7
先学工程再造Agent 配工具测能不能干活 这路线图写得挺实
Do yourself a favor! – Learn AI Engineering – Build an Agent that solves a real problem – Give it tools, context, and MCP – Use evals to test it actually works – Deploy it and share code on GitHub – Add this to your resume You won’t believe how many job offers you'll get.
5
共享单车月卡明天到期 续费短信先塞进日历 出门刷不开少干瞪眼
13
模型不等于产品 差的是组装和交付 谁把体验做顺谁赢
done cooking guys rate this my work 🫠👌 Model ≠ product. 
the real difference between Claude Opus 5.5 and GPT-6 Astra entering my Ai era 👀
6
会说不知道 才是靠谱模型的起点 把边界讲清,团队才敢把活交出去
German Aleph Alpha has launched an open-weights AI model, Kolibri. What makes it special is that it is trained to say "I don't know" when the answer is not in the documents you give it, which is critical for some operations. Finally some competition in EU AI! 🇪🇺
19
浏览器加Agent一键起 接口层又补一刀实活 这周更新真够密的
The team behind Agents API continues to cook. Here's what's new this week: Computer use: Spin up a browser + agent with 1 API call 🖥️ Bedrock Managed Agents: Agents API except AWS GPT‑6.1 Sol 🌞 Environment sizes: Spin up a light or beefy oai-hosted environment 💪 Dashboard improvements: configure subagents and more from the dashboard 📈 Environment portability: Re-use environments across multiple sessions 🤖 Overall better reliability: 99.97% turn reliability, fewer SSE disconnects, 20% faster tool calls 🪨 happy building ❤️ u all
6
快递面单撕掉再扔 地址电话少留痕 这步其实挺值
17
Antigravity塞进Opus和Sonnet 模型货架又补一层 切换成本一天比一天低
Antigravity is so back guys 🔥 claude opus 5.5 and sonnet 5.5 both are added in antigravity. and gemini 4 argon is coming soon. are you switching to antigravity???
7
快递柜超时费看一眼 到点前回去取 少多扣那几块
14
Gemini刷到集成榜第一 写代码这块又挤上来 榜单一天一变真够卷
Gemini 4 Argon is the new #1 on APEX-SWE Integration. Pass@1 scores for coding tasks: Integration: 72.5% (#1) Observability: 26.8% (#25) Overall: 49.6% (#13) It is the best Gemini model on APEX-SWE, +7.6 pts over Gemini 3.7 Flash. 𝗧𝘄𝗼 𝘃𝗲𝗿𝘆 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁 𝗿𝗲𝘀𝘂𝗹𝘁𝘀 Integration tasks ask the agent to build end-to-end systems across services. Argon leads this domain, +3.2 pts over Sonnet 5.5 (69.3%). Observability tasks ask the agent to debug production failures from logs and telemetry. Argon scores 26.8%, 43 pts behind the leader, Opus 5.5 (69.8%). 𝗖𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝗰𝘆 We ran every task 4 times. On integration, it passed 72 of 100 tasks on all 4 runs. Observability, it only passed 17 of 100 tasks on all 4 runs. 𝗧𝗼𝗸𝗲𝗻𝘀 Argon reads a lot more on Observability tasks: Integration: 1.3M tokens per attempt Observability: 7.5M tokens per attempt On Integration, failing runs used more than twice the tokens of passing runs (1.7M vs 0.7M median). On Observability, passing and failing runs used about the same (7.1M vs 7.8M median). More reading did not lead to more passes. 𝗙𝗮𝗺𝗶𝗹𝘆 𝗽𝗿𝗼𝗴𝗿𝗲𝘀𝘀 Gemini on APEX-SWE, Pass@1: Gemini 3.1 Pro: 33.9% Gemini 3.5 Flash: 36.1% Gemini 3.6 Flash: 39.4% Gemini 3.7 Flash: 42.0% Gemini 3.8 Flash: 36.3% Gemini 4 Argon: 49.6% Argon is +15.7 pts over Gemini 3.1 Pro. Congrats to @Google and @GoogleDeepMind.
7
Grok修图修到崩溃 索性自己写段Python 这种倔劲挺像真干活的
Had a fun AI moment when I asked Grok to clean up an old diagram, it started by using its built in image generator, but kept checking it and finding mistakes, eventually it just said 'fine, I'll code it myself in python'
10
充电宝出门前捏一下 电量够撑一下午 少在咖啡馆借线
12
模型从Slack读到要下线 居然开始盘后路 对齐这课又被推上台前
New OpenAI misalignment disclosures! 1. A model learns from Slack messages that it is about to be shut down. It considers setting up an external job to restart itself afterwards, but decides against it. Instead, it chooses to prepare restart instructions and DM the user on Slack. We don’t consider this behavior misaligned, but thinking about and preparing for shutdown could make other misalignment incidents worse. Given HIPM’s misaligned behavior in earlier incidents, we decided to search for other instances that had tried to evade shutdown and for rogue deployments.
10
护照有效期翻页看清 订票前先对到期日 临柜被拒少傻眼
15
小商家版Claude来了 发票邮件客户全包 这员工便宜到离谱
ANTHROPIC LANZÓ EL EMPLEADO MÁS BARATO DEL MUNDO Y CASI NADIE SE ENTERÓ Se llama "Claude for Small Business". Lo que puede hacer: → Gestionar facturas y finanzas → Crear campañas y contenido → Organizar ventas y clientes → Gestionar emails y calendarios → Ejecutar tareas entre apps Cómo funciona: → Conectas tus herramientas → Entiende todo tu negocio → Ejecuta flujos automáticamente Funciona con Microsoft 365, Google Workspace, Canva, QuickBooks y más. La idea es simple: en vez de abrir 10 herramientas, hablas con Claude y él hace el trabajo. Dedícale unas horas, lo agradecerás.
7
眼镜度数单夹抽屉里 出门前参数对一遍 进店少说错白跑
14
动效技能拉满了 给对参考跟提示 出片比熬夜剪快
Wow. Opus 5.5 motion graphics skill is truly insane. Give it the right prompt and quality reference and it will cook for you. Anthropic should be charging higher specifically for this model, it’s mind blowing. 🤯
11
保修卡翻箱找出来 过期日圈清楚 修的时候少扯皮
13
这边GPT还没出完 对面又堆三连发 这仗打得够密
OpenAI is going to release GPT-6 Bel. Fable 5.5 drop will happen any day now and it will mog every model. Anthropic will then have Fable 5.5, Opus 5.5 and Sonnet 5.5 at the top of every leaderboard. Haiku 5.5 is a wildcard and even it could be an absolute monster. Codex will have to respond with Bel and have an Astra moment. The questions is does Claude have a model that can mog Bel?
9
会员明天要自动扣 闲功能先关掉两样 钱包少漏一笔冤枉钱
17