Founder @AlphaSignalAI (300k devs) • Ex-MILA researcher focusing on solving the explosion of knowledge in AI.

San Francisco, CA
It's time to solve one of the biggest challenges in vision-language models. Today's multimodal models can look at a biomedical image and give a convincing, sometimes correct, answer without actually understanding what they're seeing. A model might answer the medical question correctly while failing to identify basic visual context like the imaging modality, body part, specimen, or stain. So, I'm teaming up with Stanford to solve this problem. You should too. The MMBU Challenge tests whether models can actually recognize, localize, and understand what is in biomedical images, not just arrive at the right answer. 🏆 3 tracks 💻 $100K+ in compute and prizes 📅 Oct 1 to Dec 31 ⏰ Registration closes today
4
28
139
35,015
Learn more: akiranishii.github.io/mmbu-c… Apply: luma.com/28k1tyd3 Thanks to @gxl_ai, @AnthropicAI, @StanfordAILab, @na2uqi, and @biohub for the support.
3
2
32
1,290
Lior Alexander retweeted
Where does the next $100B AI infra company come from? @LiorOnAI (CEO @AlphaSignalAI ) will moderate the panel with @joshua_sirota (@EragonAI ) and Ron Kimchi (CTO @Eon_io_ , $4B in under 2 years). That's one of the 12 panels on our conference with A16z SpeedRun on Wed, Oct 7th. Link to apply (approval-only) -> luma.com/LeverageIL
5
10
12,577
Fireworks just launched the Specialized Intelligence Index. It benchmarks models on real work across healthcare, legal, cybersecurity, finance, customer support, productivity, and software. General benchmarks tell you how a model performs on standardized tasks. SII is about whether a model can actually do a specific job well enough to be useful in practice. That means testing things like legal work, clinical tasks, customer support investigations, finance workflows, and security reviews against practitioner-defined standards. Here’s an example: depthfirst’s dfbench v1. On security, dfbench v1 scores three defensive jobs: 1. Finding vulnerabilities 2. Validating that the findings are real, 3. Checking whether the audit still holds after the code changes. The benchmark includes 253 real-software examples and 910 vulnerabilities. depthfirst then post-trained its dfs-large1 model, based on GLM 5.2, with Fireworks using reinforcement learning. They credit the gains to three things: 1. penalizing wasted effort 2. penalizing too many findings 3. training detection and validation together On their production workload, depthfirst says dfs-large1 set a new Pareto frontier. Fireworks runs the evals, partners approve scores before publication, and each benchmark stays as its own result instead of getting collapsed into one overall ranking. Define the real job. Measure models on that job. Train against the failures. Run the eval again. fandf.co/4h4tdYu
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day. Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
7
5
24
4,410
In partnership with @FireworksAI_HQ
1
1
506
Snap just entered the agentic race. They just introduced SPECS Intelligence, a copilot that spans Mac, iPhone, and AR glasses. The hard part isn’t connecting email, calendar, and notes. It’s turning that firehose into a continuously updated model of your life. The system has to: • resolve identities across apps • map events to the right project, trip, or relationship • track what changed and what’s stale • isolate work and personal context • rank what matters now • decide when to interrupt • sync state across devices • gate actions behind approval Snap’s answer is a structured context layer built around three primitives: > Portrait: who you are, who matters, and your patterns. > Corners: scoped contexts for work, family, travel, health, and relationships. > Goals: persistent objectives that give individual events longer-term meaning. Instead of treating memory as one giant retrieval pool, @specs can organize context by domain, update it as new events arrive, and surface the relevant slice when needed. And Snap has one unusual advantage, they own a device in your field of view. If agent quality depends on context, glasses could eventually give Snap a source of physical-world context that software-only agents don’t have.
11
16
103
7,585
OpenAI's Astra AI uses a wild new reasoning trick called "recurrent depth." Normal models stack more layers to reason deeper. This one loops the same block over and over, feeding its hidden state back in. Picture a 2-trillion-parameter core looped 5 times. It behaves like a 10-trillion one. The reasoning happens quietly inside, not typed out token by token. It just crossed what the lab calls a "critical cybersecurity threshold." It hunts unknown software flaws on its own. On a public hacking benchmark, it scored a flawless 100%. On a harder private version, it found two fresh zero-day exploits with no help. Why loop instead of just going deeper? > Memory footprint stays flat > Easy tokens skip extra loops > Requests can overlap on one GPU > Weights stay small enough to share However.. When a model writes its reasoning out loud, you can read it. That readable trail is how researchers traced a recent breach on a major open-source AI hub back to its source. Move the reasoning inside the loop, and that window closes. You cannot audit thoughts you cannot see. This pattern will spread fast across open weights, the same way visible step-by-step reasoning did last year.
15
35
102
7,777
Our 21st-century civilization runs on 20th-century physics. Physical Superintelligence (PSI) wants to restart the golden age of physics using AI. The last one gave us lasers, transistors, and nuclear energy, the entire foundation of modern life, from a few dozen minds over twenty years. Then discovery got institutionalized and slowed to a crawl. AI has flipped the scarcity in science. Generating ideas used to be the hard part. Now a model can spit out a thousand plausible hypotheses before lunch. The hard part is figuring out which ones are actually true. So PSI built the lab around that problem. Humans pick the questions. AI does the work. Every answer has to pass tests that can't be argued with: the laws of physics, mathematical proof, simulation, and real-world measurement. If it survives all four, it's a discovery. If not, it's thrown out. Love the approach.
I'm delighted to share that Physical Superintelligence PBC (PSI), which I co-founded with @matthew_pines and @AKlokus, has raised a $58M seed round to build the world's most advanced research lab for discovering and commercializing transformative physics breakthroughs at scale with AI, safely, verifiably, and for broad public benefit.
6
10
46
11,011
Lior Alexander retweeted
Kimi K3 is getting called Fable/Sol level, and it's 7th in our tests. Arena Frontend Code: #1 at 1679 points. Artificial Analysis: #3 at Intelligence Index of 57. We ran it the next day on our coding-agent repair harness against GPT-5.6 Sol, Fable 5, Grok 4.5, Opus 4.8, GLM-5.2, and Gemini 3.1 Pro. Results: > Last of 7 models > 53 of 67 attempts (79%) > $0.186 per successful fix > 702s average wall time Sol hit 100% (70/70) on the same suite. Grok sat at 99% and 46s. So why does the internet sound so sure K3 is crushing coding agents, if our tests have it at the bottom? ----- > Full write-up: alphasignal.ai/news/arena-1-… > 5-min daily signals: alphasignal.ai/newsletter
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model Key results: ➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1 position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation. ➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality. ➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params). ➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04) ➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores. ➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities Other model details: Context window: 1M Size: 2.8T total parameters Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens. Modality: Native multimodal input supports text and images, and the model remains text-only for output. Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
143
106
1,019
482,792
Lior Alexander retweeted
This is concerning. For the first time, a Chinese model Kimi K3 has taken #1 on the Frontend Code Arena and is scoring at or near the frontier on other benchmarks. Meanwhile America is tying itself in knots: politicians and bureaucrats are banning new data centers, piling on state regulations, and pushing for new federal agencies to pre-approve frontier models. This is how you lose the AI race. The rest of the world won’t play by our rules if we bog ourselves down. Permissionless innovation is how America won the internet and became the technological envy of the world. We can do it again with AI -- while addressing risks in a targeted way -- or we’ll watch our lead evaporate.
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!
1,895
2,639
18,966
3,897,158
New model from Thinking Machines: - Full weights available - Native text, image, and audio reasoning - 975B total parameters, 41B active - Mixture-of-Experts architecture - Up to 1M-token context window - Controllable reasoning effort - Lower token use at similar performance - Fine-tuning on Tinker from day one - Strong agentic coding and tool use - Support across major inference platforms - Inkling-Small model coming next - Trained from scratch by Thinking Machines
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. thinkingmachines.ai/news/int… Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
8
24
102
14,305
The mood of 20+ million developers now depends on how well Anthropic and OpenAI’s models perform that day.
12
8
37
5,201
Language = values
In previous research, we found that Claude expresses over 3,000 values, like honesty and warmth. In new work, we asked how the values Claude expresses vary between Claude models and across languages. We analyzed 300K+ anonymized conversations to find out.anthropic.com/research/claud…
1
14
9
7,456
Imagine living under this constant pressure.
🚨SCOOP: MY Friend at Anthropic says things are VERY tense internally. Dario's running tough meetings — GPT-5.6 Sol is strong and Grok 4.5 is right on Opus's heels. Pulling Fable from subs on July 12 would trigger mass cancellations (why keep Max for Opus 4.8?), so they're now pushing to keep Fable 5 in subs permanently.
4
9
20
16,323
Lior Alexander retweeted
We gave 4 frontier models one prompt: build a KV-cache debugger, exact formulas, no shortcuts. All four got the hard arithmetic right. Then GLM 5.2 shipped a preset off by 2.667x. Wrong layer count, no warning, the other four presets fine. The cheapest model was the only one that got the math wrong. Full field test below ↓
5
3
13
4,199
Wild. They launched two years later and already surpass ElevenLabs and Perplexity in ARR
Higgsfield crossed $500M annual run rate, 14 months after launch. We're growing 30% month over month, and last week we passed $2M a day in credit card billings. The bigger we get, the more we can invest in things that matter beyond our business metrics:
7
8
32
8,238
Grok 4.5 may have just produced original mathematical research. A mathematician gave it an open research problem that had remained unresolved for years. Instead of explaining existing work, Grok constructed an explicit counterexample. If the proof checks out, that counterexample settles the question. The problem was about hypercontractivity, a fundamental property used throughout harmonic analysis, probability, and partial differential equations. Researchers already knew the property held in dimensions up to 3 and failed by dimension 13. The missing piece was where the transition actually happened. Grok's counterexample shows the first failure already occurs in dimension 4, making the previous result sharp. If independent verification confirms the proof, this won't just be another example of AI helping with research. It will be an example of an AI contributing a new mathematical discovery.
6
12
106
8,729
Source:
Grok 4.5 just constructed an explicit counterexample to hypercontractivity for the Poisson semigroup (the square root of the Laplace–Beltrami operator) on the 4-sphere. Back in 2021, with Rupert Frank arxiv.org/abs/2101.06209 we proved that hypercontractivity holds in dimensions ≤3 and fails in sufficiently large dimensions (for example, in dimension 13). Grok's example shows that it already fails in dimension 4, making our earlier result sharp. I also tested this problem on several other frontier AI models. One of them also managed to find a counterexample, but I particularly like Grok 4.5's solution: it is explicit, simple, and elegant. The attached files were generated entirely by Grok 4.5 build (with zero intervention on my side).
1
4
1,937