These tests are done in „Max“ mode, nobody uses max for economic reasons. If you do the test in „High“ reasoning then Qwen beats GPt 5.6 Sol !
And the recommended reasoning is „Medium“.
The reality is crazier than even this test shows
Aug 18, 2026 · 11:29 AM UTC
36

