Qwen3.8-27B from @Alibaba_Qwen on ARC-AGI (Verified): - ARC-AGI v2: 42.4%, $0.45/task - ARC-AGI v1: 87.5%, $0.22/task Qwen's chat template adds guidance for low and xhigh reasoning, but not medium, which may explain why medium used more tokens yet scored below low.
11
4
187
13,027
Chat models commonly include a chat template that formats messages and can add instructions before they reach the model. We tested Qwen3.8-27B through @Baseten, which uses Qwen3.8-27B's own chat template. It tells low to keep its thinking brief and focused, and xhigh to think carefully, check assumptions, and consider alternatives. Medium adds neither instruction, although thinking remains enabled. That difference may help explain medium's lower scores: these settings change how the model is instructed to approach a problem, rather than simply giving it a larger thinking budget. View Qwen3.8-27B's chat template: huggingface.co/Qwen/Qwen3.8-…
1
1
28
8,109
Leaderboard costs were estimated by applying Alibaba Cloud's listed token prices to recorded token usage, rather than using our dedicated-inference charges. We will share ARC-AGI-3 results once testing is complete. Full results: arcprize.org/results/alibaba…
1
9
1,899
- Leaderboard: arcprize.org/leaderboard - Reproduce the public results: github.com/arcprize/arc-agi-… - Testing policy: arcprize.org/policy - Full Qwen3.8-27B results: arcprize.org/results/alibaba…

Oct 1, 2026 · 8:08 PM UTC

4
1,323
Sort replies: Relevant Recent Liked