On AutomationBench, GPT-6.1 Sol scores 31.7% at medium reasoning effort, up 4.8 percentage points from GPT-6 Sol at the same setting.
On OSWorld 2.0’s offline set, it scores 71.4% versus Astra’s 73.5%, both at maximum reasoning effort, at roughly one-seventh of Astra’s cost per task.