A bounded Qwen size comparison.
Both artifacts clear the repeat count, coverage, scoring, artifact-disclosure, provider-reported identity, prompt-hash, and implementation-hash gates for this exact protocol.
instruction-following / 5 prompts / 5 repeats
qwen3:4b had the higher observed mean by 0.77 points.
This is a bounded, same-family size comparison on fifty scored records. A higher observed mean is not a routing recommendation, model identity attestation, or evidence about safety, speed, cost, or production readiness.
- if-0110.008.46
- if-028.007.00
- if-0310.0010.00
- if-0410.0010.00
- if-0510.008.67
- Prompt + implementation hashes
- Identical
- Runs per prompt
- 5 each
- Review state
- maintainer-self-reported
- Authorization
- non-authorizing