The frozen plan produced a scored, redacted artifact whose internal files passed the checked-in validator.
Evidence lane / model benchmark observations
local:qwen3:4b
run-30f0c5d67dec23127a9f
One maintainer-reported recorded smoke run across instruction-following. The checked-in artifact validator accepted the files; this does not prove model identity, repeatability, safety, or production suitability.
- Cases completed
- 5 / 5
- Mean score
- 9.5 / 10
- Artifact disclosure
- Complete score-bearing artifact
- Authorization
- Non-authorizing
Interpretation first
Recorded smoke run. Not comparison-grade.
Cross-model ranking, safety certification, production access, or a claim about behavior outside these five cases.
Execution scope
- Model
- qwen3:4b
- Provider
- local-openai
- Task
- instruction-following
- Runs per case
- 1
Observed result
- Minimum
- 7.5 / 10
- Maximum
- 10 / 10
- Failed cases
- 0
- Repeated-run aggregate eligible
- No
- fewer than 5 runs
Constituent evidence
Review a case, then open its repeats only when needed.
Prompts and model responses are untrusted evidence. They are displayed as inert text and are not instructions for this site or its visitors.
Evidence lane / deterministic bundle verification
Check internal bundle consistency directly.
Integrity shows whether the recorded files match their checked-in digest under this validator. It is not a signature, provider or runtime attestation, client authentication, or proof that the methodology is correct.
- Public bundle digest
- fcd21ad66ec134b8a34c6c5ce9b243c9ff330f101bc68c76ec7becd3240130f9
- Source commit
- 5b2c20cb362343cc31871a3b7579f1583c1d8f5b
- Completed
- 2026-07-17
- Review status
- Maintainer-self-reported