I'm obsessed with structure, cognition, and the human brain — three scales of the same question: how does intelligence organize itself?

Stanford
sophieinthelab retweeted
Announcing the Dimensional Residency For hungry founders eliminating physical labor with robots in any vertical No equity, just free hardware, access to customers, test environments, and engineering support to scale deployments 3 months. 10 teams. Apply now.
10
25
211
53,686
Everyone cites Voyager as the canonical "self-improving agent." It isn't one. Getting the distinction right matters, because it's the difference between a moat and a pile. 【N1】Voyager (NVIDIA/Caltech, 2023) plays Minecraft by writing JavaScript. GPT-4 proposes its own next task, writes a function to execute it, repairs that function against runtime errors, and a second GPT-4 instance verifies success before the function is committed to a vector-indexed skill library. New functions call old ones. It was the only method to reach diamond tools unassisted. 【N2】Swap GPT-4 for GPT-3.5 and unique-item discovery drops 5.7x. Same architecture, same loop. 【J1】That number is the whole story. The system moves capability, it doesn't create it. Voyager widens itself; it cannot heighten itself. Real RSI modifies the machinery that produces improvement. Voyager's weights, curriculum prompts, verification criteria and primitive APIs are all frozen and sit outside the loop. Accumulation is not recursion. 【J2】And for nearly every agent product shipping today, accumulation is enough. Human brains haven't improved in 10,000 years; everything we have came from externalized, compounding assets. Better still: the ceiling gets raised for you, free, every few months. Building RSI means paying to do what the labs already do, and trading away auditability and rollback to get it. 【N3】But Voyager's own discovery curve flattens late. Retrieval stays top-5 while the library grows, so the usable fraction shrinks. Nothing refactors, abstracts, or deprecates. 【J3】So undifferentiated accumulation is not a moat either. Past some volume, a library that only grows starts losing value. Skills rot, weak primitives get depended on, and the verifier never gets better. 【J4】The defensible layer was never the accumulation. It's whatever organizes and validates it. Making that layer self-improving is the cheap, safe 80% of RSI that almost nobody is building.
1
33
1/ Most AI agents in production run on a borrowed human identity. The audit log lies by default.Anthropic's Kristen Swanson named this as one of three things that turn an agent into a teammate: 【N1】Memory — it holds the goal across weeks, not turns 【N2】Its own credentials — it acts inside guardrails, not on someone's borrowed access 【N3】Shared context — if it isn't written down, for an agent it doesn't exist 2/ All three are right. But something sits underneath them that nobody has built. 【J1】A new hire doesn't get full access on day one. Small project, then a bigger one. Permission grows as evidence accumulates.That curve is how trust actually works inside an organization. 3/ 【J2】Agents have no curve. The decision is binary — full credentials, or a demo that never ships.That binary is where most enterprise agent pilots die. Not on model capability. On a permission question nobody wants to put their name on. 4/ 【J3】The missing layer: progressive authorization.· Start scoped to read · Log every action under the agent's own identity · Expand scope when its record earns it · Contract it when it doesn'tTrust as a measured variable, not a checkbox signed once. 5/ 【J4】A guardrail written into a system prompt is not a guardrail. It's a request. A boundary the model can reason past was never a boundary. Progressive authorization only means anything if it's enforced outside the model. 6/ 【J5】Models get cheaper and more swappable every quarter. The trust record — who granted what, on what evidence — is what compounds.
35
1/ For three years, we taught machines to speak. Jev bets that what the world actually needs is machines that decide. 【N1】Jev, from TypeSafe AI (founded by ex-OpenAI's Diogo Almeida), generates no text. You give it program state and typed questions, and it returns answers with calibrated probabilities in a single parallel pass. 2/ 【N2】"System One" comes from Kahneman. System 1 is fast intuition, and System 2 is slow deliberation. LLMs reason token by token, which makes them System 2. Jev aims to be System 1. 【J1】We built the slow mind first. Evolution did it the other way around.
24
3/ 【J2】"Can't hallucinate" is the wrong frame. A closed answer space means it can't invent, but it can still be confidently wrong. The real gift isn't truth. It's a number that tells you how much to trust the answer. 【J3】A machine that knows when to hand a decision back to a human is worth more than one that always sounds right. 4/ 【N3】It's named after Jevons, of Jevons' paradox: when a resource gets cheaper, we use more of it, not less. 【J4】That's the real thesis. At $42 per billion input tokens, judgment stops being a service you call and becomes a material you build with. It becomes like electricity, not like a consultant. 5/ 【J5】Wittgenstein wrote: "Whereof one cannot speak, thereof one must be silent." Jev is silent by design, and it still acts. 【J6】If judgment becomes a commodity, value moves to whoever composes it: orchestration, data, and trust. The model is the neuron. The product is the nervous system.
16
【N1】393 papers on deployment-time self-evolution, the largest category in the corpus. 340 on training-time self-iteration, which the authors call "the technical heart of RSI as currently practiced." 【N2】RSI means each round of improvement makes the next one easier. Weight-level loops qualify: the model generates its own training signal (STaR, self-rewarding RL, self-play), and the gain persists into the next round. 【J1】The field is split 50/50. The discourse is 95/5. 【J2】Scaffolding cashes out the gap between what a model can do and what it does. One-time, capped by the weights. It approaches a ceiling and stops. 【J3】Only the weight loop compounds: arxiv.org/abs/2607.07663
33
Should a personal AI be allowed to communicate directly with other people, or should it only interact with its own user and other people’s AIs, rather than talking directly to humans on its owner’s behalf?
26
FunSearch (DeepMind, Nature 2023): LLMs alone confabulate. Pair them with a hard evaluator — and they discover new mathematics. Wrap your AI's creativity inside a hard evaluator before trusting its output. → distill.loomus.ai/d/8dxr54
41
🇻🇳 #HOCHIMINH #AI The world isn’t flat ; it’s layered in access, language, and opportunity. In Silicon Valley, people teach machines to dream. In Vietnam, many still wait for permission to begin. At VNG, a 4,000-person company, I asked why self-developed games remain rare. “We’re still learning how,” they said. AI education isn’t just about coding. It’s about giving people the knowledge access to rewrite their own story. #Vietnam #AI #AIPhilosophy #TechHumanism
34
🇮🇹 #ROME #AI “Rome wasn’t built in a day.” At #RomeFutureWeek, I opened my keynote with a question: With AI Agents rising, can we build Rome faster, or smarter? Only one hand went up for faster. In Rome, that answer hit differently - a city where time itself feels like an algorithm: slow, deliberate, self-correcting. The next era of AI won’t be defined by speed, but by intention, consciousness, and the meaning of progress. #ROME #AI #ROMEFutureWeek #AIPhilosophy #TechHumanism
23