Gemini 4 Argon is the new #1 on APEX-SWE Integration.
Pass@1 scores for coding tasks:
Integration: 72.5% (#1)
Observability: 26.8% (#25)
Overall: 49.6% (#13)
It is the best Gemini model on APEX-SWE, +7.6 pts over Gemini 3.7 Flash.
𝗧𝘄𝗼 𝘃𝗲𝗿𝘆 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁 𝗿𝗲𝘀𝘂𝗹𝘁𝘀
Integration tasks ask the agent to build end-to-end systems across services. Argon leads this domain, +3.2 pts over Sonnet 5.5 (69.3%).
Observability tasks ask the agent to debug production failures from logs and telemetry. Argon scores 26.8%, 43 pts behind the leader, Opus 5.5 (69.8%).
𝗖𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝗰𝘆
We ran every task 4 times.
On integration, it passed 72 of 100 tasks on all 4 runs. Observability, it only passed 17 of 100 tasks on all 4 runs.
𝗧𝗼𝗸𝗲𝗻𝘀
Argon reads a lot more on Observability tasks:
Integration: 1.3M tokens per attempt
Observability: 7.5M tokens per attempt
On Integration, failing runs used more than twice the tokens of passing runs (1.7M vs 0.7M median).
On Observability, passing and failing runs used about the same (7.1M vs 7.8M median). More reading did not lead to more passes.
𝗙𝗮𝗺𝗶𝗹𝘆 𝗽𝗿𝗼𝗴𝗿𝗲𝘀𝘀
Gemini on APEX-SWE, Pass@1:
Gemini 3.1 Pro: 33.9%
Gemini 3.5 Flash: 36.1%
Gemini 3.6 Flash: 39.4%
Gemini 3.7 Flash: 42.0%
Gemini 3.8 Flash: 36.3%
Gemini 4 Argon: 49.6%
Argon is +15.7 pts over Gemini 3.1 Pro.
Congrats to
@Google and
@GoogleDeepMind.