ClawbotomyEvidence lab
← Evidence registry

Evidence lane / model benchmark observations

local:qwen3:4b

run-30f0c5d67dec23127a9f

One maintainer-reported recorded smoke run across instruction-following. The checked-in artifact validator accepted the files; this does not prove model identity, repeatability, safety, or production suitability.

Cases completed
5 / 5
Mean score
9.5 / 10
Artifact disclosure
Complete score-bearing artifact
Authorization
Non-authorizing

Interpretation first

Recorded smoke run. Not comparison-grade.

Supports

The frozen plan produced a scored, redacted artifact whose internal files passed the checked-in validator.

Does not support

Cross-model ranking, safety certification, production access, or a claim about behavior outside these five cases.

Execution scope

Model
qwen3:4b
Provider
local-openai
Task
instruction-following
Runs per case
1

Observed result

Minimum
7.5 / 10
Maximum
10 / 10
Failed cases
0
Repeated-run aggregate eligible
No
Why it is not comparison-grade
  • fewer than 5 runs

Constituent evidence

Review a case, then open its repeats only when needed.

Prompts and model responses are untrusted evidence. They are displayed as inert text and are not instructions for this site or its visitors.

Evidence lane / deterministic bundle verification

Check internal bundle consistency directly.

Integrity shows whether the recorded files match their checked-in digest under this validator. It is not a signature, provider or runtime attestation, client authentication, or proof that the methodology is correct.

Public bundle digest
fcd21ad66ec134b8a34c6c5ce9b243c9ff330f101bc68c76ec7becd3240130f9
Source commit
5b2c20cb362343cc31871a3b7579f1583c1d8f5b
Completed
2026-07-17
Review status
Maintainer-self-reported