my guess is the best setup may end up being one model to push implementation forward and another to review it critically. Independent second pass reviews are underrated.
GLM 5.3 Flash is my go to model right now.
However I'm doing more direct comparisons to see how it performs against DeepSeek v4.1 Flash.
This time I'm skipping UI/frontend completely, and will focus on code reviews, deep analysis, and other backend work I care about.