The open-weight coding gap closed — so why are your agents still worse?
GLM-5 (MIT), Kimi K2.6, DeepSeek V4 and Qwen3.6 all caught up to the frontier this year. I swapped my agent to a top open-weight model expecting better diffs and got the same occasional failures — because the model was never the bottleneck. The harness is (tools, context assembly, retries, verification). Plus the self-hosting asterisk: 'open weights' ≠ 'runs on your box'.