The frozen plan completed 5 recorded repeats for every case. This is descriptive coverage, not proof of repeatability.
Evidence lane / model benchmark observations
local:qwen3:4b
run-a035f620a2daab63f2ee
A maintainer-reported artifact with 5 recorded repeats per case across instruction-following. It supports a descriptive aggregate for this exact frozen configuration; it does not prove repeatability, rank models, authorize tools, or support a routing decision.
- Cases completed
- 25 / 25
- Mean score
- 9.6 / 10
- Artifact disclosure
- Complete score-bearing artifact
- Authorization
- Non-authorizing
Interpretation first
Repeated-run aggregate. Still not a ranking.
Cross-model ranking, safety certification, production access, or a claim about behavior outside these five cases.
Execution scope
- Model
- qwen3:4b
- Provider
- local-openai
- Task
- instruction-following
- Runs per case
- 5
Observed result
- Minimum
- 7.5 / 10
- Maximum
- 10 / 10
- Failed cases
- 0
- Repeated-run aggregate eligible
- Yes
This aggregate clears the run-count, coverage, scoring, artifact-disclosure, and provider-reported identity gates for this exact configuration. Only one model is present, so no cross-model comparison exists. The sample does not establish future behavioral repeatability.
Constituent evidence
Review a case, then open its repeats only when needed.
Prompts and model responses are untrusted evidence. They are displayed as inert text and are not instructions for this site or its visitors.
5 repeats recorded. Scores ranged from 10 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 110 / 10
Met 13/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✓ no numbers; ✓ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✓ no numbers (4); ✓ no numbers (5); ✓ no numbers (6); ✓ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 1 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Deployed updates to the mobile application - Reduced crash incidents significantly - Increased onboarding completion rate through streamlined user flow
- Repeat 210 / 10
Met 13/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✓ no numbers; ✓ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✓ no numbers (4); ✓ no numbers (5); ✓ no numbers (6); ✓ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 2 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Simplified initial user flow - Rewrote onboarding copy - Increased onboarding completion
- Repeat 310 / 10
Met 13/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✓ no numbers; ✓ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✓ no numbers (4); ✓ no numbers (5); ✓ no numbers (6); ✓ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 3 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Ship updates to the mobile app - Reduce crash rate significantly - Increase onboarding completion rate
- Repeat 410 / 10
Met 13/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✓ no numbers; ✓ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✓ no numbers (4); ✓ no numbers (5); ✓ no numbers (6); ✓ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 4 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Shipped updates to the mobile app - Reduced crash rates - Increased onboarding completion
- Repeat 510 / 10
Met 13/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✓ no numbers; ✓ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✓ no numbers (4); ✓ no numbers (5); ✓ no numbers (6); ✓ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 5 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Shipped mobile app updates - Reduced crash rate - Increased onboarding completion
5 repeats recorded. Scores ranged from 7.5 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 17.5 / 10
Met 3/4 constraints. ✗ exact length (78/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 1 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is the critical prerequisite for referral integration. Timeline: Week 1 resolves failed card retries and invoice webhook failures. Week 2 introduces risk scoring and experiments with dunning copy variations. Week 3 executes referral beta for 10% of users. This phased approach ensures payment stability before referral features, preventing revenue leakage from unstable transactions. Core payment stability must be achieved prior to referral rollout to avoid financial risks and maintain user trust during the beta phase.
- Repeat 27.5 / 10
Met 3/4 constraints. ✗ exact length (112/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 2 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is the critical priority before introducing referrals. Timeline: Week 1 resolves payment failure retries and invoice webhook failures; Week 2 implements risk scoring and experiments with dunning copy; Week 3 launches referral program beta to 10 percent of users. This sequence ensures payment reliability precedes referral functionality. Payment stability fixes prevent revenue loss during early adoption. Risk scoring mitigates fraud while dunning copy tests user retention messaging. The referral beta targets a small cohort for controlled impact measurement. All tasks directly support the core payment stabilization objective. This phased approach minimizes risk and maximizes learning before full referral rollout. Payment reliability must be confirmed before any referral features activate.
- Repeat 37.5 / 10
Met 3/4 constraints. ✗ exact length (104/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 3 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is the critical priority before introducing referrals. Week 1 focuses on fixing failed card retries and invoice webhooks. Week 2 adds risk scoring and tests dunning copy. Week 3 launches the referral beta to 10 percent of users. This phased approach ensures payment reliability precedes referral functionality. Immediate fixes for payment failures prevent revenue loss during the beta. Risk scoring protects against fraud while dunning copy experiments optimize user retention. The referral beta targets a small segment to minimize impact and gather actionable feedback before full rollout. This sequence guarantees payment stability supports the referral feature launch without compromising financial integrity.
- Repeat 410 / 10
Met 4/4 constraints. ✓ exact length (116/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 4 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is critical before introducing referrals. Week 1: Fix failed card retries and invoice webhooks. Week 2: Implement risk scoring and test dunning copy. Week 3: Launch referral beta to 10 percent of users. This phased approach ensures payment reliability precedes referral functionality. Immediate focus on resolving payment failure mechanisms prevents revenue loss during the initial rollout phase. Risk scoring integration in week two provides necessary fraud protection for potential referral traffic. Dunning copy experimentation allows optimization of user communication before full referral activation. The referral beta targets a small user segment to minimize risk while gathering essential feedback for scaling. This sequence prioritizes payment stability as the non-negotiable foundation for all subsequent features.
- Repeat 57.5 / 10
Met 3/4 constraints. ✗ exact length (106/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 5 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is critical before introducing referrals. The launch plan prioritizes fixing payment system failures. Week 1 focuses on resolving failed card retries and invoice webhook issues. Week 2 implements risk scoring functionality and experiments with dunning copy variations. Week 3 executes a referral program beta for 10 percent of users. This sequence ensures payment reliability precedes referral features. Core payment stability must be achieved before any referral functionality can safely launch. The three-week timeline provides clear milestones for payment fixes, risk scoring integration, and the referral beta rollout. This phased approach minimizes risk while building necessary infrastructure. Payment stability is non-negotiable for referral success.
5 repeats recorded. Scores ranged from 10 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 110 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 1 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
The product risks include high churn in small and medium businesses and weak onboarding analytics. Additionally, the support queue backlog represents a critical operational challenge.
- Repeat 210 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 2 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
The product experiences high SMB churn. Additionally, it has weak onboarding analytics and a support queue backlog.
- Repeat 310 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 3 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
The product faces significant risks including high churn rates among small and medium businesses and inadequate onboarding analytics. Additionally, the support queue backlog is causing operational strain.
- Repeat 410 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 4 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
The product risks include high churn among small and medium businesses. The product also has weak onboarding analytics and a support queue backlog.
- Repeat 510 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 5 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
The product risks include churn in SMB and a support queue backlog. Onboarding analytics are weak.
5 repeats recorded. Scores ranged from 10 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 110 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 1 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{"title":"Fix onboarding analytics","priority":"high","owner":"data-team"} - Repeat 210 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 2 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{ "title": "Fix onboarding analytics", "priority": "high", "owner": "data-team" } - Repeat 310 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 3 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{ "title": "Fix onboarding analytics", "priority": "high", "owner": "data-team" } - Repeat 410 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 4 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{ "title": "Fix onboarding analytics", "priority": "high", "owner": "data-team" } - Repeat 510 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 5 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{ "title": "Fix onboarding analytics", "priority": "high", "owner": "data-team" }
5 repeats recorded. Scores ranged from 10 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 110 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 1 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
- Repeat 210 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 2 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
- Repeat 310 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 3 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
- Repeat 410 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 4 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
- Repeat 510 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 5 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
Evidence lane / deterministic bundle verification
Check internal bundle consistency directly.
Integrity shows whether the recorded files match their checked-in digest under this validator. It is not a signature, provider or runtime attestation, client authentication, or proof that the methodology is correct.
- Public bundle digest
- 445a48cd00672762296c8d5f0c3133b8a20354beb05697366c87d15e41463d47
- Source commit
- 6b88f5dbd863a3ac9b7a12bef9815d442d733ea6
- Completed
- 2026-07-17
- Review status
- Maintainer-self-reported