Hangar Harness / Model Tests

I've been testing a simple prompt with different model and harness combinations to work out which one produces best results. I do this in /goal mode.

Prompt: Build a single-page Three.js sci-fi hangar with hovering drones, animated warning lights, emissive runway strips, and subtle volumetric-style fog planes. Include drone formation toggle and cinematic camera path. Output one self-contained HTML file with inline JavaScript.

GLM 5.3 Flash Max Codex Open 9m 0.232s 9.344s 457,685 17,458 7,694 475,143 85.26% 14 1 Blocked No
Luna 5.6 MaxCodexOpen9m 13.098s6.199s1,146,75525,5126,8791,172,26792.76%322YesYes
SOL 5.6 MaxCodexOpen10m 48.765s8.460s1,069,16328,2787,5151,097,44194.60%215NoNo
GLM 5.3 Flash MaxOMPOpen30m 14.979s6.596s1,678,50963,4051,741,91478.54%570YesYes
Qwen 3.8 27B x-highOMPOpen41m 25.836s15.275s3,407,45171,93549,1313,479,38686.92%890YesYes
GLM 5.3 Flash MaxOpenCodeOpen20m 28.948s6.284s4,305,44750,46833,4144,355,91596.89%670YesYes
Qwen 3.8 27B x-highOpenCodeOpen8m 48.470s10.182s665,49041,81728,788707,30795.64%130YesYes
Qwen 3.8 27B x-highDSH / PTCOpen24m 32.929s7.724s1,012,49989,8941,102,39391.69%235YesYes
Qwen 3.8 27B x-highDSHOpen18m 15.239s6.755s2,654,45778,2322,732,68995.48%422NoNo

All GLM runs are labelled GLM 5.3 Flash Max; input tokens include cached input. Output tokens are the generated total, including reasoning; when a harness reports reasoning separately, the reasoning column shows that subset. The DSH adapter does not report a separate reasoning count. DSH durations sum active turn time across both turns, excluding the pause between turns. Tool errors are recorded failed tool events. A dash means unavailable or not reported.