GPT-6-luna medium reasoning. Codex CLI - Pro Plan vs. API Key. 50x runs on the same prompt, interleaved. Run script and generated outputs data here: huggingface.co/buckets/evals…
2
2
10
702
Shaun Smith retweeted
PSA: @mervenoyann is destined to teach @pewdiepie about local agents and post-training. upvote her plea and make a dream come true. github.com/odysseus-dev/odys…
3
5
53
10,094
Shaun Smith retweeted
Open standards take a village, and MCP's is showing up. :muscle: Over the past year, its core repository had 721 active contributors from 216 active organizations across 46 countries and territories. Thank you to everyone contributing to MCP.
2
4
755
xAI raising cache read prices on the Grok 4.6 launch is absolutely killing them price / performance wise. By far their biggest issue now.
2
189
Shaun Smith retweeted
You can now train an open model with RL inside claude code, codex, opencode, pi, or any harness you need. All at once, with the harness runs unmodified. @adithya_s_k added a capture proxy to openenv. it sits between the harness and the model and records the exact token ids and logprobs of every call. to the harness it's just another model provider. harbor supplies the tasks and sandboxes, trl trains with async GRPO. it's worth doing because the harness completely changes what the model learns. the same LFM2.5-2.6B weights solve 62% of held-out tasks in mini-swe-agent and 33% in claude code. after RL across four harnesses the average goes from 42% to 54%, and claude code from 33% to 49%. there are three envs you can try in the browser and a training script for each. guide: huggingface.co/spaces/Adithy…
21
21
181
14,068
Something is up with OpenAI serving. Overnight test (with my personal API key on flex) experienced a big jump in speed and drop in intelligence, back to normal this morning. (gpt-6-luna). I don't really have time to dig in to this stuff, but it's measured and recorded. Unhappy.
207
Caching < Stupidity
1
4
398
Oh and I wrote a "--" in there too. fml.
52
Shaun Smith retweeted
I'm giving away $1,000 of HF credits 🤗 $10 x 100 people, I want you to experience this: > having your agent launch hundreds of HF Jobs > run ~33M tokens of DeepSeek V4.1 Flash > deploy your own dedicated Qwen3.8 27B endpoint > Agents + Blender on a GPU because it's 🔥 Just reply with your HF username 👇
Hugging Face Jobs usage is stonking 📈 All thanks to agents: you just need the hf CLI installed and you have access to an insane amount of hardware. Asked mine to test hundreds of code snippets from Hub model pages: it launched 322 Jobs in 90 minutes, up to 25 at once, CPU to A100 (diffusers, transformers, sklearn, MLX, Keras...) Total bill: ~$4 🤯
279
28
233
35,356
“We are pleased to announce the latest Gemini model!” Great, can customers use it? “No.”
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
4
5
89
6,191
Shaun Smith retweeted
Updates for Codex and ChatGPT Work users. No nerfing, only good stuff! - We have landed inference optimizations and are passing down savings to all the subscriptions for GPT-5.6 Sol. That should result in around 10% more usage on its own. - We noticed that by changing the context size limit in the product to 372k for GPT-5.6 Sol, up from 272k for GPT-5.5, it resulted in more usage being charged than intended. We have reverted to 272k and will work to roll back out to 372k in the days to come. You should notice that usage drains significantly less after this change. - To understand where the extra usage was coming from, we ran some experiments where reasoning efforts were changed (referred to as juice values under the hood) and have reverted this. - There is slightly more usage of multi-agent than intended in high and xhigh reasoning effort, we are fixing this going forward. Also fixing a small other thing we noticed with auto-review where we can be more efficient. And we continue to have the 5h limit temporarily not apply. Enjoy the rest of the weekend!
OpenAI has reduced GPT-5.6 Sol's thinking budgets in an effort to make the model more efficient They essentially bumped everyone's reasoning down by 1... so if you were running Sol Extra High, you now have to set it to Max to get the same effort So we basically don't have Max reasoning anymore, how do you feel about these changes? 🤔
1,035
596
9,831
2,282,970
Can anyone give me a super-high level view of how GitHub Copilot token pricing works? Is it discounted at all (like it says Max is $100 -> $200, but do they use list prices?)
2
1
336
Anyone got any good ideas for other levers providers could adjust to make your quota go further?
1
1
320
This is amazing for benchmarking btw. No weird network egress restrictions. 24 hour job length (yes, tasks > 1hr) No weird pricing scheme that locks you in or demands $250 to be useful.
Hugging Face Jobs usage is stonking 📈 All thanks to agents: you just need the hf CLI installed and you have access to an insane amount of hardware. Asked mine to test hundreds of code snippets from Hub model pages: it launched 322 Jobs in 90 minutes, up to 25 at once, CPU to A100 (diffusers, transformers, sklearn, MLX, Keras...) Total bill: ~$4 🤯
1
9
729
Shaun Smith retweeted
getting acquired by @nvidia = hugging face can now hire people we couldn't as a small startup and give them a decade to make open-source AI win! if you're one of them, my dms are open
274
326
7,099
373,932
xAI Responses WebSockets have a max age of 25 minutes. Happily terminates the connection mid-inference 🙃
5
309
Replaced "embedding-drift-monitor" with "telecom-entity-resolution", and upped the repeats to 5. The drift monitor task has an instruction/verifier mismatch due to be fixed in the version of tb.
Here's my first pick for this - tested with Astra Max/Luna, now giving ds4.1 a little run through. Tasks are selected to be indicative, not favouring one model series and simple to run (no GPU/multi-container).
330
Wow! Holy turnaround Anthropic, Opus 5.5 slaps.
Right, and now I'm basically bankrupt. Hits hard when you pay for this by the token from your own wallet. Not a had a session without this pathology yet, not enjoying it any more. Waiting for Opus 5.1 I guess.
1
11
649
Why is the ChatGPT app telling me to download the ChatGPT app?
10
13
611