Who starts?
Can the same 7.4 M parameter model as in the main demo run faster? Yes: with custom kernels in onepass-webgpu, a small WebGPU runtime, from the fp32 file or the main demo's int8 file, compared with onnxruntime-web on WebAssembly. A perfect Connect Four solver is timed alongside: faster on most moves, but far less consistent. The story behind it: One millisecond to make a move.
Lower is better; the best in each column is blue, the worst red. Times fill in after the benchmark.
| engine | size (runtime + model) | start-up | median | mean | p95 |
|---|
Before the WebGPU runtime is trusted, it must agree with onnxruntime-web (wasm, same fp32 file) on reference positions. The encoder is also checked against the Python tooling's bytes.