We are here to ensure our partners become the signal of innovation in the noise of the technology markets.

An @AMD Instinct MI355X node does more agentic work per dollar as the serving stack behind it improves, and the same node now serves 1.44x the MiniMax-M3 throughput it did when Signal65 PINNACLE first published it, in our testing. ➡️ AITER unified attention and piecewise graph capture on the same vLLM 0.23.1 build added 11% peak throughput and took the MI355X node from 52 to 56 agents sustained at the service floor ➡️ At 8 agents, per-agent decode went from 18.3 to 43.4 tok/s, 2.4x, on the same build ➡️ vLLM 0.30.1 with the @AMD MiniMax-M3 configuration added another 30%, 563 to 731 tok/s ➡️ 31% lower cost per output token on hardware a buyer already owns On the version-matched build the MI355X now sustains more agents at the PINNACLE floor than the NVIDIA B300. AMD engineering has been a strong partner in this work, and more MI355X results on the current vLLM cycle are coming. @AnushElangovan @roaner @AIatAMD
1
3
13
48,373
An AI rack an enterprise has already paid for does more agentic work when its serving stack is tuned to each model, and an untuned stack leaves that capacity on the table. In our first Signal65 PINNACLE (pinnacle.signal65.com) tuning results on @nvidia GB300 NVL72, a single setting change added up to 64% more output throughput in our testing, with the model, nodes and load held fixed. 🧵
2
5
55,302
A rack-hour costs the same no matter how many tokens come out of it, so those gains come straight off the cost of every token. In our testing, tuning cut the relative cost per output token by up to 39% on the same GB300 NVL72 hardware, which means more agent work finished per dollar, or the same workload served with fewer GPUs.
1
52
This is the next phase of Signal65 PINNACLE, moving from single servers to rack-scale systems with more than 4,200 vLLM serving runs across 53 model and precision combinations, and @nvidia has been a strong partner in it. The tuned configurations go through Signal65 PINNACLE quality and agent-capacity runs next, and larger multi-node and NVFP4 results are coming soon. pinnacle.signal65.com
52
Signal65 retweeted
Enterprises buy AI capacity for the work it produces, and the cloud underneath decides how much of that work a fixed budget buys. At the scale of a 5,000-GPU @nvidia Blackwell deployment over three years, the cost difference between providers runs from hundreds of millions of dollars to more than $2B. We modeled that deployment on @CoreWeave and three hyperscaler clouds using public pricing and MLPerf Inference v6.0 results, and in our analysis CoreWeave came out ahead on both cost and work per dollar. ➡️ A $1M annual commitment on GB200 NVL72 buys about 1.33 trillion tokens on CoreWeave, 64% to 195% more than the hyperscaler clouds ➡️ Three-year cost on GB200 NVL72 runs up to 65% lower ➡️ On HGX B300, CoreWeave comes in 52% lower, roughly $1.2B less over the term ➡️ Higher MLPerf throughput turns a 22% price advantage over the closest competitor into a 65% lower cost per million tokens Full report: signal65.com/research/ai/cor…
1
2
9
109,102
Signal65 retweeted
More users shouldn't have to mean slower AI. In @Signal_65 RAG testing hosted via @TensorWave, AMD Instinct MI355X delivered ~42% lower p99 latency and roughly 2x the concurrent-user headroom before breaching the evaluated SLA: bit.ly/45hAvl3
11
14
129
30,661
OpenAI reopened the $200 ChatGPT Pro plan on September 29th with half the included usage and said subscribers still get more work done. That is true, from a certain point of view. Signal65 PINNACLE scored GPT-6.1 Sol over the OpenAI API at medium and max effort on the same agentic jobs and 128K retrieval corpus as the rest of the board. ➡️ GPT-6.1 Sol at max makes 1.5x fewer weighted errors than GPT-6 Sol at max and lands within 12% of GPT-6 Astra, at one fifth of the Astra price per correct task ➡️ It generates 42% more output tokens per correct task than GPT-6 Sol, so per credit it finishes about the same number of tasks, 31 cents against 29 at max and 16 cents at medium on both ➡️ A subscriber comparing to last week is worse off by task count on the same tier, 47% of the correct tasks for 1.5x fewer errors on each ➡️ A subscriber comparing to August, the comparison OpenAI made, is better off, because the Sol API price was cut in half on September 22nd ➡️ An Astra subscriber is better off only by downgrading, to 2.4x the correct tasks on the halved allowance at 89% of the Astra score
3
330
Signal65 retweeted
The first @AMDRyzen AI Max+ PRO 400 Series systems are now shipping with up to 192GB of unified memory, with as much as 160GB of it available to the GPU, a 67% jump over the prior generation. HP pairs the chip with Perplexity Computer in the ZBook Ultra G3a, GMKtec and Framework are bringing compact desktops, and Lenovo has the 1.6L ThinkCentre X Ultra on the way with cluster-ready designs for enterprise agent workloads. The timing is deliberate. Microsoft has a moment coming that is clearly built around local and hybrid agentic AI, and NVIDIA RTX Spark systems arrive in October with a 128GB ceiling, so @AMD landing first with 50% more memory puts real pressure on that launch. The prices are a gut punch, with the 192GB ZBook at $7,449 (estimated) and the GMKtec EVO-X5 Pro at $6,599. Most people will not be buying one, but the client PC is once again where some of the most interesting AI engineering is happening, and that is good for the industry. amd.com/en/blogs/2026/how-am…
2
3
20
2,188
The @OpenAI GPT-6 models get better at agentic work when you turn the thinking effort up. @AnthropicAI Claude Opus 5.5 does not. @Signal_65 PINNACLE scored GPT-6 Sol, GPT-6 Luna and Claude Opus 5.5 at their default (medium) and maximum reasoning effort on the same multi-step enterprise jobs and the same 128K-token retrieval corpus, at one list price per model. ➡️ GPT-6 Sol at max effort: 2.5x fewer weighted errors than at medium (Model Score 3,998 to 9,853), 99.6% of multi-step jobs finished against 96.8%, fabrication down from 7.9% to 2.9%. It costs 1.8x as much per correct task, 29 cents vs 16, and it lands on the cost frontier above Grok 4.6, Claude Fable 5.1 and GPT-6 Astra at medium. ➡️ GPT-6 Luna at max effort: 2.4x fewer weighted errors (1,682 to 4,115), 99.6% of multi-step jobs finished against 88.6%, for 2x the cost, 3 cents against 1.5. That is above Claude Sonnet 5 and GPT-5.6 Sol at 27x and 33x less per correct task. ➡️ Claude Opus 5.5 at max effort: 1.8% more on the score (10,418 to 10,603) for 3.7x the cost, $1.45 against 40 cents, 4.8x the output tokens, and fabrication up from 2.8% to 4.9%. It already finished all of the multi-step jobs at medium in our testing. ➡️ The trade is the same on all three: max effort spends 3.3x to 4.8x the output tokens per correct answer, so token efficiency falls. On the GPT-6 models you get a better model for it. On Opus 5.5 you get the same model at a higher price, and medium is the setting to run. Every score and every price of a correct task is live at pinnacle.signal65.com
3
1
13
33,477
GPT-6 Luna deserves its own line. At max effort it is the first small hosted model on the PINNACLE board to finish agentic work at a frontier rate, and it does it for 3 cents a correct task. ➡️ 99.6% of the multi-step enterprise jobs finished end to end in our testing, up from 88.6% at medium, the same completion rate as GPT-6 Sol at max ➡️ A Model Score of 4,115, above Claude Sonnet 5 (3,782) and GPT-5.6 Sol (3,954), for 27x and 33x less per correct task, and above every open-weight build on the cost view except Qwen3.8-Flash-Next ➡️ Four rows leave the cost frontier because of it: GPT-6 Sol at medium, DeepSeek-V4-Flash-0731, Qwen3.8-27B and Qwen3.6-27B, so a hosted model now holds every frontier step below 5,000 but one ➡️ The cost of the jump is tokens, 30,624 output tokens per correct answer against 7,332 at medium, and retrieval is still where it shows its size: 92.8% on the 128K corpus against 97.3% for Sol, with fabrication at 11.4%, down from 14.4% but still second highest among the hosted models OpenAI does not publish a parameter count for Luna, and its $0.10 / $0.50 list price puts it in the small tier. Nothing priced like that has finished agentic jobs at this rate on our board before.
1
196
Signal65 retweeted
Even though there was no new silicon announcements for the @Snapdragon PC family, the addition of new @surface devices using X2 plus, new Googlebook system announcements, and Linux support for X2 Elite make for a pretty compelling mid-cycle refresh story.
5
3
20
1,148
Signal65 retweeted
Perplexity is bringing Portable Computer to @AMD Ryzen AI Max, including the Ryzen AI Halo developer platform, so local tasks and recurring workflows can run on the user's own hardware without drawing down Perplexity Computer credits. That connects the cost of an AI subscription directly to the silicon on the desk, and it gives users a concrete reason to pay for a large shared memory pool and a capable integrated GPU in a client system. I expect most PC AI to settle into this hybrid model, with routine agent work running locally and the cloud reserved for frontier-scale reasoning. Ryzen AI Max is one of the few client platforms today with enough memory to hold models large enough to make that split REALLY useful. The value of those savings will depend on how much of a typical agent workload the local models can absorb before handing off to the cloud, and are looking at multiple angles of this at @Signal_65 right now.
Portable Computer for Windows is now available on @AMD Ryzen AI Max Series processors. It makes it easy to run local AI agents that work with your connected apps and local files. Kick off tasks or schedule recurring work that runs entirely on your device.
2
2
20
1,302
Our Opus 5.5 benchmarking versus Fable 5.1: - >2x cheaper for correct task 🔥 - slightly more reliable ✅ but 100% - 4x more hallucinations 😱 but only 2.8%
Any enterprise working on @AnthropicAI Claude Opus for agent work gets more of the work done for less money by moving to yesterday's brand-new Opus 5.5, and in our testing the gap is...pretty large. Opus 5.5 gets more work done for more people, at a noticeably lower price than Anthropic offered before, as measured by @Signal_65 PINNACLE. ➡️ 2.3x fewer weighted errors than Claude Opus 5, and 100% of the multi-step jobs finished end to end against 95% ➡️ 34% less per correct task (!!), $0.40 against $0.60, on a list price 20% below Opus 5 with cache reads at 5% of input rather than 10% ➡️ Fabrication in retrieval answers down from 7.6% to 2.8% ➡️ It also passes Claude Fable 5.1 with 1.3x fewer weighted errors at 63% less per correct task Full comparison on the Signal65 PINNACLE results views at pinnacle.signal65.com
3
7
32
4,038
Any enterprise working on @AnthropicAI Claude Opus for agent work gets more of the work done for less money by moving to yesterday's brand-new Opus 5.5, and in our testing the gap is...pretty large. Opus 5.5 gets more work done for more people, at a noticeably lower price than Anthropic offered before, as measured by @Signal_65 PINNACLE. ➡️ 2.3x fewer weighted errors than Claude Opus 5, and 100% of the multi-step jobs finished end to end against 95% ➡️ 34% less per correct task (!!), $0.40 against $0.60, on a list price 20% below Opus 5 with cache reads at 5% of input rather than 10% ➡️ Fabrication in retrieval answers down from 7.6% to 2.8% ➡️ It also passes Claude Fable 5.1 with 1.3x fewer weighted errors at 63% less per correct task Full comparison on the Signal65 PINNACLE results views at pinnacle.signal65.com
1
3
7
43,754
Along with the new Opus model, yesterday saw the release of new GPT-6 options as well. GPT-6 Luna, in our PINNACLE testing it is the cheapest correct task of any hosted model at 1.5 cents. It is also the model that fabricates the most. @Signal_65 PINNACLE scores long-context retrieval on 369 documents at 128K tokens, lease contracts, sales field reports and HR evaluations generated from one ground truth, and counts every answer that cites something not in the documents. ➡️ GPT-6 Luna fabricates 14.4% of its retrieval answers, the highest rate on the board; the next highest is GPT-5.6 Sol at 10.3% ➡️ The models on the cost frontier above it sit at or near zero, Grok 4.6 and GPT-6 Astra at 0.0%, Claude Fable 5.1 at 0.7% ➡️ GPT-6 Sol, the mid-tier sibling, fabricates 7.9% at 16 cents per correct task, and Claude Opus 5.5 2.8% at 40 cents ➡️ Luna still makes 1.4x fewer weighted errors than GPT-5.6 Luna at a fifth of the price, so the generation gain is real; the fabrication rate is the cost of the price
1
151
Another standout data point from our testing: @OpenAI GPT-6 Sol finishes a correct task in 3,500 output tokens in our testing, the lowest of the 17 hosted builds on the @Signal_65 PINNACLE board. ➡️ 2.8x fewer output tokens per correct answer than Grok 4.6 or Claude Fable 5.1, and 22x fewer than Claude Haiku 4.5 ➡️ The three models launched this week all sit in the terse half of the board, Sol at 3,500, GPT-6 Luna at 7,300 and Claude Opus 5.5 at 7,500 ➡️ Sol scores level with GPT-5.6 Sol on 20% fewer output tokens and a list price half as high, which is how a correct task ends up at a sixth of the price ➡️ The verbose end of the board is where reasoning effort lives, Gemini 3.8 Flash at 20,000 to 24,000 and Muse Spark 1.3 at 37,000
2
140