A few thoughts about the latest Opus model:
I've had several long sessions going concurrently today. Up to now, Opus only dispatched subagents for reviews. Just a second ago, we reviewed several outstanding items at once. Before I could finish typing it, Opus launched 4 subagents to tackle them.
Every turn very consistently follows the house rules of conclusions up front, then what was done, then what's needed from me. Except for short answers where it's obviously not useful to do so. Previous Claude models have been inconsistent about rule following, while OpenAI models follow rules pathologically.
If you've been following along, I've posted several times recently about Gemini's improved performance. It often found problems even with Fable and Astra. Twice today (and for the first time EVER) I've had Opus nail a build with zero findings following an adversarial review from a different model family.
The good sense of this model is incredible. It isn't perfect, but it's the first model that feels like working with a peer instead of just a smart computer. From what I've seen this model is at or near intelligent human level in terms of taste and common sense.
Really impressed.