Pinned Tweet
I ran Pavlov's experiment on a robo-fly using @GoogleDeepMind's FlyBrain & @typesafeai's Jev. It learned what to avoid, what to chase, and when to fire its escape reflex.
1
2
10
905
most AI products are currently targeted towards other AI people or enterprises.
This is my thesis: no one uses AI. I repeat, absolutely no one. We live in a bubble. Even among my friends who pay for it, when I ask them to open ChatGPT and show me their queries, it’s the same handful of basic things. Most don’t even know they can upload a photo and ask questions about it. Connecting Gmail so an agent can read and send emails blows their minds. An agent opening a browser and checking them into a flight? They’ve never even heard of it. The massive challenge right now is adoption, and then getting people who already signed up to actually use what they’re paying for. Most have absolutely zero clue what’s possible. Imagine the compute shortage when everyone starts using AI like the top 1% of users do today.
1
70
pasan retweeted
Replit CEO Amjad Masad on how general models could train smaller, domain-specific models on the fly: "There's a lot of talk of recursive self-improvement, but there's something I don't think is getting a lot of discussion, which is models training their replacements." "You can think of it as a just-in-time compiler. As you're executing dynamic code, the interpreter realizes there's an opportunity to optimize, and it emits machine code on the fly that's a lot more optimized." "You can imagine general models, you're doing something with Operator or Astra, some of the big models, and they realize the use case is limited, or some other agent observing realizes the use case is limited." "General agents have all these flaws, but there's also more potential for them to be harmful, more potential for them to go off the rails." "So the model, on the fly, trains a model that could be its replacement, but is a lot more domain specific. Therefore it's cheaper, less vulnerable to prompt injections, and less harmful for you, because it's less capable." "It's almost like a system that's training machine learning models for specific use cases as it's monitoring the entire system." @amasad
.@OpenRouter co-founder Alex Atallah, in his first podcast since Stripe acquired the company, joins @Replit co-founder Amjad Masad and a16z's Erik Torenberg on why the future of AI is independence and specialization. In this conversation, Alex walks through how the Stripe deal unfolded, why he wasn't originally looking to sell, and why "payments and inference are going to blend together." Pre-OpenRouter, the typical AI workflow had one model provider to choose from, and little pressure on that provider to lower prices. Now enterprises are diversifying across labs and open-weight models, and every board is asking about AI costs and benchmarks. Amjad argues if your company depends on one AI lab, it can turn into your competitor. So Replit is building the layer that lets enterprises use any model and any cloud, without being locked into either. Alex and Amjad are split on personal agents – Amjad runs one agent across his whole company and loves the cross-domain joins, while Alex says general agents cause you to sacrifice understanding, and argues 10 specialized chiefs of staff beats one superagent. 0:45 How the Stripe deal unfolded 5:05 Why mixing models beats one model 7:25 Forcing the labs to compete on price 8:50 Enterprises want open-weight models 10:30 Every board asks about AI every month 12:25 Why companies must own their intelligence 14:15 Replit as the independence layer 15:10 Everyone is building the same agent 16:35 Why Amjad built bring-your-own-cloud 18:10 Amjad's agent that runs his whole company 19:55 Why 10 specialized agents beat one 23:35 Machines, not humans, should specialize 27:30 Guardrails for agents talking to agents 31:10 Models training their own replacements 33:45 Most tasks don't need a frontier model 40:50 Training small models on Qwen 8B 43:25 The Rust cycle is coming for AI 45:10 Fusion models: frontier quality at half the cost YouTube: piped.video/ekK8urKHPMQ @alexatallah @OpenRouter @amasad @eriktorenberg
40
19
214
69,889
pasan retweeted
@AnthropicAI is either lying or they have dumbs ass engineer's. It takes me 1 HOUR TOPS. 1. you only need 1 GPU to load 1 layer at a time and --showhiddenstates in your sglang or vllm serve. 2. run 12-32 harmful/benign pars at each layer to see where the refusals are most active. 3. find your separation and kl then apply your alpha across the layers that were most active while normalizing the ones that are not. 3. stream the conversion to its original quant and load the model. TIP: your question bank can reverse your directions depending on how many and how experts are triggered. (FOR MOE) also using a compliance prompt when capturing directions ONLY. can often reveal the hidden and much deeper refusal detection layer. this way when you apply the edit with the prompt off, your edit will cover the hidden reasoning as well. @AnthropicAI HIT MY INBOX ILL SELL YOU ALL MY RESEARCH. BUT HEADS UP CLADE WONT DO IT, TOO MANY SAFEGUARDS.
So wait something doesn't add up. @AnthropicAI engineers said it looks them 2200 GPU hours to obliterate GLM 5.3? how is that possible? 3 months? 91 days? $4400? how come? Either that is a typo, I mean even let's say they spent days figuring out, testing etc it won't take 91 days for 4 people. I took me 10 hours yesterday, 1 person, including the time to upload the weights to hugging face. I suspect they used Claude to do abliteration and claude was trolling them for 91 days, spending tokens left and right, running in circles. That's the only explanation I have (or it's a Typo)
5
11
139
4,135
super cool
Replaced a few motors and expanded the range of motion, I think my air hockey robot is just better than me at this point
1
2
302
pasan retweeted
Alexandr Wang on enterprise perception ▫️Better products can still lose ▫️Big buyers rarely confront reality ▫️Palantir gave hires an acting book Wang: Perception is more real than reality
38
181
2,391
601,169
damn
Claude Opus 5.5 just took a big drop on NerfBench. Yesterday it was scoring above launch. Today it's at 94.2%. GPT 6 Astra: 98.0% Sonnet 5.5: 100.9% GPT 6.1 Sol: 106.7% 94.2% is still inside normal variance, so we can't call it a nerf yet. But we're watching Opus 5.5 very closely.
1
472
+1 on Codex being the best UI
I haven’t touched Claude Code or Codex CLI in a while. The terminal era is over imo. It's the wrong interface for coding agents. Tabs are ephemeral, but context is persistent, and managing 30 tabs is pure cognitive overhead. I don’t really need an IDE like Cursor either. I rarely navigate the whole codebase anymore. The new primitive is the agent, not the file. (Codex desktop app is the best agentic UI for now. But we’re still early.)
4
549
Can you estimate an object’s physical properties, such as its stiffness and velocity, from a single video? Our @NeurIPSConf 2026 paper, MonoPhysics, answers this question. Led by my students @daniel129911845, Jun, Matthew, and in colab with @DBiswadip. Project Page: daniel03c1.github.io/MonoPhy… 1/2
4
15
122
6,470
pasan retweeted
Reasoning from scratch, round number 6! An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO). 00:00 Introduction 01:54 What makes a reasoning model different? 04:25 Reasoning traces and model capability 08:29 Accuracy and format rewards 11:34 Aha moments and DeepSeek-R1 training 14:41 Reasoning effort and answer length 18:38 RLHF and RLVR 23:04 GRPO vs. PPO 26:40 GRPO explained with a cooking analogy 31:43 The KL term and simplified GRPO 35:04 Loading the pretrained model 36:07 Loading the MATH training data 39:26 Sampling model responses 46:30 Computing verifiable rewards 49:55 Computing advantages 51:54 Token and sequence log probabilities 55:29 Implementing sequence log probabilities 57:37 Fixing the inference-mode error 1:02:24 Computing the GRPO loss 1:04:37 Putting the GRPO step together 1:09:19 The GRPO training loop 1:12:57 Training settings, logging, and checkpoints 1:17:24 Running training and inspecting outputs 1:19:28 Loading and evaluating checkpoints 1:22:33 MATH-500 results and training stability 1:24:05 Memory requirements and next steps
49
128
1,189
48,361
pasan retweeted
Feed-forward models distort geometry even when given camera poses. In paper to appear in #NeurIPS, we recast multi-view stereo as sequence-to-sequence: a camera-aware transformer with a unified global cost volume predicts geometry for all views jointly. arxiv.org/abs/2609.24850
11
126
6,836
pasan retweeted
(1/8) We are excited to share our new model ARROW 🏹 ! ARROW brings D4RT-style decoding to broader input types, including multi-view videos and unordered collections of images. 📜arXiv: arxiv.org/abs/2610.01314 🌐Project page: vision.rwth-aachen.de/arrow
10
28
203
12,566
damn
ok fine, here's free GLM 5.3 Flash uncensored ~13m TPM shared ~360 TPS peak use it for cybersec, coding, whatever live at platform.xplabs.ai
1
93
Excited to share that 𝘾𝙎𝙄𝘾𝙇 is accepted to #COLM2026 🎉 We introduce code-switching in-context learning, an inference-time cross-lingual representation alignment mechanism that gradually bridges non-En languages with En representations. 📄 arxiv.org/abs/2510.05678 (1/🧵)
1
9
57
6,715
MolmoMotion Forecasting Point Trajectories in 3D with Language Instruction MolmoMotion is a 4B vision-language model that forecasts 3D point trajectories under natural-language action instructions. Given a short RGB observation history, a set of user-specified 2D query points with their initial 3D positions, and a language description of the intended action, the model predicts each query point's 3D trajectory for up to ~2 seconds in the camera-frame-at-t₀ coordinate frame. We show that the learned motion prior transfers to robotics planning and to motion-guided video generation.
1
7
53
3,063
EvenSplat: Coupled 2D-3D Decomposition for Gaussian Splatting under Exposure and Illumination Variation Tongyu Wu, Jacob Edwards, @Empire_Xiao_Yan, @JiangCaigui, Cheng Wang tl;dr: image-space decomposition optimized together with the Gaussian representation arxiv.org/abs/2610.01876
2
8
35
1,910
pasan retweeted
Excited to share PixelUMM! One model for image & video understanding and generation, directly in pixel space. No VAE, no vision encoder. Paper, model & code are out! nv-tlabs.github.io/PixelUMM/ Huge shoutout to @CongWei1230 for leading the effort and to the incredible team! 🚀
Let's remove VAEs and ViTs from video models! 🚀 Introducing 𝗣𝗶𝘅𝗲𝗹𝗨𝗠𝗠: an encoder-free unified multimodal model for image and video understanding and generation, directly in pixel space. Code and model available today! 🌐 nv-tlabs.github.io/PixelUMM 📄 arxiv.org/abs/2609.38597
6
20
1,180
Does biological wiring make a better reservoir? Our WIVACE 2026 paper benchmarks different C. Elegans connectomes. Randomized baselines often perform better; results depend on the task and configuration. arxiv.org/abs/2609.30508 #ALIFE #NeuroAI
2
7
29
2,691
pasan retweeted
VoxMem Researchers from University of Melbourne and UNSW released VoxMem on Hugging Face, a benchmark for spoken memory in audio LLMs. No model tops 40% at 32K; models remember what was said far better than who said it, how, or what was audible.
1
6
20
1,866
New paper: instead of decomposing a whole neural network, we only dissect the parts used on a task we care about. It's much cheaper, and gives mechanisms you can inspect, erase, or edit to change how the model behaves 🧵
2
19
198
12,616