SB 947 The no robo boss law In California, is BS. "So the practical reading is a firing protocol. Log the model output. Have a person attach a file. Send a plain-language notice. Name a human contact. Do not let the model be the only signature. That is how a robot boss fires you in California after July 2027, with a human countersignature." #SB947
17
keithofaptos retweeted
You would think that this is the shortest route from South America to Australia but no airline flies like that for those destinations. Why?
Community note
Scheduled flights between South America and Australia use great-circle routes passing far south near Antarctica over the Southern Ocean. The Antarctic continent itself is not crossed due to ETOPS safety standards requiring nearby diversion airports unavailable there. myflightroutes.com/blog/south-ame… simpleflying.com/etops-banned-a… simpleflying.com/antarctic-circ…
537
185
3,608
2,177,193
keithofaptos retweeted
mathematicians and muse spark collaborated to solve 6 open problems in math: 1. The Strict Threshold for Gaussian Ellipsoid Fitting 2. Finite-Time Blow-Up of Radial Negative-Energy Solutions for the Mass-Critical Biharmonic Nonlinear Schrödinger Equation 3. Semiabelian Groups Need Not Be Monomial 4. Tightness of the Cycle-Based Relaxation for Completed Length-Three Alpha-Cycles 5. String Two-Point Function = Height Function on a Curve 6. On Solvable Evolution Algebras and a Conjecture by García-Martínez and Pérez-Rodríguez
Following gold-medal-level performance from our AI models across five competitions in mathematics, physics, and chemistry, we asked a harder question: can AI contribute when a problem is genuinely open and without an existing solution path? Over the past several months, mathematicians worked with Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the regular meta.ai chat interface, with no custom research scaffold, to find solutions to such problems. Our goal wasn't to mass-produce papers, but empower researchers. Every collaboration followed the same principles: mathematicians guided the research, a second group of mathematicians then reviewed the work, each paper marks which passages were drafted primarily by humans or AI, and each credits the prior research it builds on. Where other teams independently announced solutions to the same problems, we acknowledge their work as well. Today, we're sharing six papers from that collaboration. 🧵👇
152
196
3,038
383,868
keithofaptos retweeted
Introducing Decision 2.0: our newest state-of-the-art decision models, open in every size from 0.6B to 27B ⚡ HF collection 🤗: huggingface.co/collections/v… #OpenSource #Decision #Jev #LLM
33
125
1,436
85,511
We need models 4x smarter, 100x smaller, and 10k's_x more compute. Cognitive agentic systems. And JARVIS's. Oddly enough this is 100% already in route.
We need at least 50x more compute. At least.
31
lol
Only in SF😭😭😭
28
keithofaptos retweeted
Andrej Karpathy predicted the future of AI once again: “Everyone’s renting frontier models for jobs a 3B model could do. Small models are the future.” this 18-page PDF breaks down Karpathy’s case for working with small LLMs. the real question isn’t “Is the small model as good?” It’s “Which of my 1,000 calls ever needed a frontier model?” And @thewebai just answered it for formal logic. TwIL-LM3-Pro: → 3.6B params, on par with Qwen3-8B on formal logic → leads VibeThinker-3B on all 6 formal-logic tasks tested → 95.4% on BBH logic, 95% on SVAMP → 2.09 GiB in Q4, runs on CPU or 4GB VRAM → no API bill, no data leaving your machine The secret isn't size. It's post-training. PDF below. Model 👇 huggingface.co/webAI-Officia…
Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸
Paid partnership (ad)
33
241
1,542
122,131
Finally got 4.7 on my SuperGrok account. About fukn time! 😎 @grok
1
1
1
87
keithofaptos retweeted
models behave differently across harnesses. if a model is trained only on one harness, it can perform poorly on others. in this collab with @adithya_s_k and the @huggingface team, we look at: • how much a model's base performance can differ across different harnesses • what happens when you train a model on a single vs. multiple harnesses • how much tool efficiency can improve if you add a small reward this guide also contains a lot of great basics around agentic RL, incuding the difference between white-box vs. black-box harnesses, etc. full guide: huggingface.co/spaces/FineEn…
1/ Excited to release The ultimate guide to multi-harness RL The same model behaves differently in every agent harness. So we built an open way to train any model with RL on any task set, inside the harnesses people actually use, like Claude Code, Codex, and OpenCode, without changing a single line of harness or training code. Trained across four harnesses, LFM2.5-2.6B went from 42% to 54% with 31% fewer tool calls. 🧵
4
11
112
6,441
Impressive.
Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. TwIL-LM3-Pro from @thewebAI has 3.6B parameters and runs locally on everyday computers. It scores 95.4 on BIG-Bench Hard, well ahead of Qwen3-8B at 63.7. I like their post-training recipe. They fine-tune on formal logic, merge the weights back toward the base model, and then run RL against a programmatic verifier. Logic scores go up, and general reasoning holds steady. Great to see more of this work released as open source.
71
Intelligence capabilities are not going to slow down, at all. So if you are not sure what's going on here then you may want to zoom out and really take another perspective view. Even that view currently isn't really going to tell you the future. My guess is that what we see now is likely 1 in 10,000 of where this is headed. And that 1 is all of human history of knowledge combined. How quickly society as we know it changes is the real question, imo. In 4 years the IQ of general Ai has gone from 40 to 120, exponentially (recently). Now RSI inside the frontier's internal models. The open weight creators will soon release models that make Prisms Bonsai tech of taking a 27B model down to 6B w/a 95% intelligence retention look like child's play. Meaning that a 4B parameter model's IQ on the artificial analysis intelligence score will go from 40-60's Will into the 80's. And these tiny models will work like fusion/counsels collectively, for a higher IQ. Not to mention that these smaller smarter models in conjunction with cognitive agentic systems will be able to run efficiently on local hardware. Distributed computing "groups" will make a lot of ground here. And soon. People only know what they know. Millions of Agents are, definitely, going to blow right past humanity.
Interesting study. But let's face it, soon enough AI's will do the vast majority of inventing. 10's of thousands per year. So fast that patents won't be relevant.
38
How long until interdimensional acceleration; ID/acc ? Comms via entanglement w/Orch-OR, or similar. Because at the speed of development currently this seems realistically possible before chip ramp and space deployments at scale.
Before we get space-based computing at scale, sea-based computing makes so much sense! Acceleration on the seas. 🌊/acc
29
I'm gonna go ahead and declare AI/SI from here forthwith; Mathematical Knowledge.
California Gov. Gavin Newsom signed an executive order today (30 sep) declaring that Artificial Intelligence will be called "Artificial Intelligence" in California. > Employers can't use AI to read workers' emotions. AB 1883 bans AI surveillance tools that recognize or predict an employee's emotional state, or that collect neural data, such as readings from brain or nervous system sensors. > AI can't make firing decisions alone, and AI layoffs must be disclosed. Employers can't rely only on AI for discipline or termination. They must also tell workers if a mass layoff, relocation, or termination is caused by an AI system. > The wider package goes beyond the workplace. It protects doctors' own judgment when clinical AI tools are used, bans deleting AI watermarks, strengthens likeness protections against deepfakes, and requires gene synthesis companies to verify their customers and screen what they ship.
22
Here's a stratosphere view on wheee the action is going to be.
The frontier labs (Claude + ChatGPT + Grok + Gemini) have been log-monotonically plowing into the enterprise super intelligence market. People often ask what inning we are in? On this definition of market size, AI software is roughly 0.1% penetrated into the enterprise; the first pitch has only just been thrown and is not even a quarter of the way to the catcher's mitt. Adjusting for the realized rate of market penetration of the frontier labs, and assuming they follow a traditional diffusion curve, some $30 trillion in revenue could be up for grabs by 2030. Even if the rate of diffusion were to slow to 2/3rds that realized thus far, $10 trillion in revenue could be at stake for the AI frontier labs by 2030.
47
Could support a few thousand Agents, a day is more like it. 😳 Currently it's just a concept. But these will get built. Possibly with: Jet fuel underground tank(s). Underground facility with ac units, liquid chillers, and turbine engine(s) generator(s). Above ground shop/office with a gang of Starlink satellites on the roof. My guess ... won't be long from now. Especially for "Those bunkers".
This is the NX ONEPOD, a 100kW micro data centre. A contenarised 5.8 TB of aggregated GPU memory with 32x NVIDIA B200 GPUs. Fully plug and play and backed up with batteries for 24h operations and reducing netwrok loads. This could be built for around 3M$, built and deployed in 5 months and could support approximately 50,000–150,000 people whose daily digital activity is extensively AI-assisted.
1
1
51
This chart stops at Aug 7th. That's 7 weeks short of today. And this is only Openrouter! Imagine what the actual number is across ... -everything- . Now go ahead and fathom this number exactly 1 yr from now, globally.
Humans are the minority user of AI Agents burn nearly 5x the tokens people do, up 14x since February More charts in State of Markets II: a16z.news/p/state-of-markets…
20
keithofaptos retweeted
it's ironic I just did this like 3 days ago for all my VMs and then you guys added it! Haha
1
4
502