this demo of codemode + jev in pi beautifully demonstrate both the power of codemode and general classification models
well done w/ this example @badlogicgames, @mitsuhiko, and the rest of your team
People of Pi: We’ve partnered with @OpenAI to integrate their new Sign in with ChatGPT experience. Pi users can now access their OpenAI subscriptions and credits in Pi via one unified login flow.
Sonnet 5.5 vs Opus 5.5 High on my-mini-bench and comparison with Sonnet 5
I tested them on tasks from my own work: data conversion, frontend, refactoring and running task evaluations.
Setup: 10 tasks × 3 runs, High reasoning, in pi-agent and Claude Code through harbor.
Sonnet 5.5 vs Opus 5.5 in pi:
> Both passed 30/30.
> Time: 46s vs 103s: ~2.2× faster.
> Cost: $0.10 vs $0.30 per task: ~3× cheaper.
Compared with Sonnet 5 in pi:
> Passes: 25/30 → 30/30.
> Turns: 12.7 → 4.3
> Output tokens: 17.7k → 6.5k
The trajectories show a different approach to reading files, writing code, and testing. Examples below 👇
This benchmark is already saturated, but I use it to find the best setup for my tasks: cheaper, faster, and still gets the job done. It also lets me quickly compare models and see how their behavior differs.
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
Welcome to our Monday Meditations. In which @badlogicgames and @mitsuhiko are talking about things that might or might not need to change in Pi given the changes in models.
Hi people of Pi. Greetings from Vienna. Enjoy your weekend and spend some time with humans.
And consistent with that message, we're moving our Sunday videos to Monday.
we’re thinking of killing plan mode and using the shift+tab hotkey to adjust effort levels
I don’t think the models need plan mode anymore, but if you’re a plan mode diehard would love to get your feedback on why
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
People of Pi, what are some of the best agentic experiences you have built or seen people build with Pi? What are some ways we could improve the experience for you?