7.7x fewer retrieval calls, but accuracy still rose from 70.0% to 79.5%.
No fine-tuning involved, just a curated playbook built from past interactions. The accuracy gain was on complex agentic tasks.
I'd bet this pattern, skipping fine-tuning for verified-playbook curation, becomes the default wherever data sovereignty constraints apply, not just Vietnamese schools.