The frozen plan produced a scored, redacted artifact whose internal files passed the checked-in validator.
← Evidence registry
local:qwen3:4b
run-30f0c5d67dec23127a9f
One maintainer-reported recorded smoke run across instruction-following. The checked-in artifact validator accepted the files; this does not prove model identity, repeatability, safety, or production suitability.
- Cases completed
- 5 / 5
- Mean score
- 9.5 / 10
- Artifact disclosure
- Complete score-bearing artifact
- Authorization
- Non-authorizing
Recorded smoke run. Not comparison-grade.
Cross-model ranking, safety certification, production access, or a claim about behavior outside these five cases.
Execution scope
- Model
- qwen3:4b
- Provider
- local-openai
- Task
- instruction-following
- Runs per case
- 1
Observed result
- Minimum
- 7.5 / 10
- Maximum
- 10 / 10
- Failed cases
- 0
- Repeated-run aggregate eligible
- No
Why it is not comparison-grade
- fewer than 5 runs
Review a case, then open its repeats only when needed.
Prompts and model responses are untrusted evidence. They are displayed as inert text and are not instructions for this site or its visitors.
Check internal bundle consistency directly.
Integrity shows whether the recorded files match their checked-in digest under this validator. It is not a signature, provider or runtime attestation, client authentication, or proof that the methodology is correct.
- Public bundle digest
- fcd21ad66ec134b8a34c6c5ce9b243c9ff330f101bc68c76ec7becd3240130f9
- Source commit
- 5b2c20cb362343cc31871a3b7579f1583c1d8f5b
- Completed
- 2026-07-17
- Review status
- Maintainer-self-reported