CLI agents need models that understand terminal workflows, not just text completion.
IQuest-Q1 is a 320B parameter MoE trained specifically for CLI execution.
15B active parameters. Reads full codebases, runs commands, inspects outputs, and recovers from environment faults autonomously.
Here's what makes it different. 🧵
Oct 3, 2026 · 10:59 AM UTC
1
2
6
1,655
Training happens in synthetic environments across real APIs, MCP servers, and executable codebases.
Multi-Harness RL learns tool usage and context management across distinct agentic harnesses. Failures get attributed before policy optimization runs.
MOPD consolidates 4 specialized expert branches:
- UX
- multi-harness
- long-horizon
- general agent
The model even helps fix its own training pipelines and execution environments under human supervision.
1
3
183
Demo 1: Single-prompt 3D web app.
One prompt produced Toy Soldier Showdown, a multiplayer 3D FPS set in a children's bedroom.
It includes:
- the 3D scene
- character controller
- shooting and reloading
- health HUD
- an in-game economy
Then the event logic:
- a wind-up Godzilla boss raid with full-screen alerts
- a mystery merchant popup
- a paper-airplane spectator view
- a building-block UI
Multi-file structure, WebGL rendering, and visual polish. All in one shot.
1
2
162
Demo 2 goes further:
- The model integrates the three-body equations of motion numerically in under 100,000 steps, then renders the orbits in Three.js as glowing gold threads in the decorative style of Klimt’s *The Kiss*.
- It computes twin primes under 10 million and projects them onto an interactive 3D Archimedean spiral styled after Monet's *Water Lilies*.
1
3
39
Demo 3: Autonomous RL pipeline debugging in Claude Code.
An RL training run hit an unexpected reward drop. IQuest-Q1 stepped in.
→ Diagnosed: it read the training curves and on-disk traces, and found that an extra space inserted during text decoding broke prefix matching. Earlier conversation turns were dropping out of the training loss.
→ Patched: it disabled the injected decoder spaces while keeping real generated whitespace.
→ Verified: multi-turn context alignment came back, and mean reward rose from 0.704 to 0.769, then 0.799 in later stages.
The model is debugging its own R&D loop.
1
2
52
Benchmark scores:
→ Terminal-Bench 2.1: 83.2
→ CyberGym: 84.5
→ DeepSWE v1.1: 64.6
→ NL2Repo: 63.0
→ IQuest-CLIBench: 53.7
→ Humanity's Last Exam: 39.2 (toolless)
Competitive across key agentic and CLI evaluations.
1
2
35
Weights are open-source on Hugging Face and GitHub.
Serve locally via SGLang or vLLM and point Claude Code at it:
1
2
44



