These tests are done in „Max“ mode, nobody uses max for economic reasons. If you do the test in „High“ reasoning then Qwen beats GPt 5.6 Sol ! And the recommended reasoning is „Medium“. The reality is crazier than even this test shows
let that sink in

Aug 18, 2026 · 11:29 AM UTC

36
Sort replies: Relevant Recent Liked