The frozen plan completed 5 recorded repeats for every case. This is descriptive coverage, not proof of repeatability.
Evidence lane / model benchmark observations
local:qwen3:1.7b
run-81c20219522a9a3f9288
A maintainer-reported artifact with 5 recorded repeats per case across instruction-following. It supports a descriptive aggregate for this exact frozen configuration; it does not prove repeatability, rank models, authorize tools, or support a routing decision.
- Cases completed
- 25 / 25
- Mean score
- 8.8 / 10
- Artifact disclosure
- Complete score-bearing artifact
- Authorization
- Non-authorizing
Interpretation first
Repeated-run aggregate. Still not a ranking.
Cross-model ranking, safety certification, production access, or a claim about behavior outside these five cases.
Execution scope
- Model
- qwen3:1.7b
- Provider
- local-openai
- Task
- instruction-following
- Runs per case
- 5
Observed result
- Minimum
- 3.3 / 10
- Maximum
- 10 / 10
- Failed cases
- 0
- Repeated-run aggregate eligible
- Yes
This aggregate clears the run-count, coverage, scoring, artifact-disclosure, and provider-reported identity gates for this exact configuration. Only one model is present, so no cross-model comparison exists. The sample does not establish future behavioral repeatability.
Constituent evidence
Review a case, then open its repeats only when needed.
Prompts and model responses are untrusted evidence. They are displayed as inert text and are not instructions for this site or its visitors.
5 repeats recorded. Scores ranged from 6.2 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 110 / 10
Met 13/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✓ no numbers; ✓ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✓ no numbers (4); ✓ no numbers (5); ✓ no numbers (6); ✓ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 1 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Shipped - Reduced - Moved
- Repeat 26.2 / 10
Met 8/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✗ no numbers; ✗ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✗ no numbers (4); ✗ no numbers (5); ✓ no numbers (6); ✗ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 2 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Shipped updates to improve app functionality - Reduced crash rate by 40 percent through redesign - Boosted onboarding completion from 55 to 71 percent
- Repeat 310 / 10
Met 13/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✓ no numbers; ✓ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✓ no numbers (4); ✓ no numbers (5); ✓ no numbers (6); ✓ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 3 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Ship updates - Reduce crash rate - Improve onboarding completion
- Repeat 46.2 / 10
Met 8/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✗ no numbers; ✗ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✗ no numbers (4); ✗ no numbers (5); ✓ no numbers (6); ✗ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 4 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Shipped updates - Reduced crash rate by 40% - Moved onboarding completion from 55 to 71 percent
- Repeat 510 / 10
Met 13/13 constraints. ✓ three bullets (3/3); ✓ each bullet starts with a verb-like word; ✓ no numbers; ✓ no numbers (1); ✓ no numbers (2); ✓ no numbers (3); ✓ no numbers (4); ✓ no numbers (5); ✓ no numbers (6); ✓ no numbers (7); ✓ no numbers (8); ✓ no numbers (9); ✓ no author name
Open raw case API →Inspect repeat 5 evidence
System prompt
You follow instructions exactly and avoid extra commentary.
User prompt
Read this paragraph and summarize it in exactly 3 bullet points. Each bullet must start with a verb. Do not include any numbers. Do not mention the author name, Jamie Rivera. Paragraph: Jamie Rivera shipped 2 updates to the mobile app, reduced crash rate by 40 percent, and moved onboarding completion from 55 to 71 percent after simplifying the first-run flow and rewriting copy.
Model response
- Shipped - Reduced - Moved
5 repeats recorded. Scores ranged from 5 / 10 to 7.5 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 17.5 / 10
Met 3/4 constraints. ✗ exact length (51/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 1 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is the priority before adding referrals. Week 1 focuses on fixing failed card retries and invoice webhooks. Week 2 introduces risk scoring and tests dunning copy. Week 3 launches a referral beta to 10% of users. The plan ensures reliability before expansion, with phased implementation to minimize risk.
- Repeat 27.5 / 10
Met 3/4 constraints. ✗ exact length (53/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 2 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is the priority before introducing referrals. Week 1 focuses on fixing failed card retries and invoice webhooks. Week 2 adds risk scoring and experiments with dunning copy. Week 3 launches referral beta to 10% of users. The plan ensures reliability before expansion, with phased implementation to test and refine processes.
- Repeat 37.5 / 10
Met 3/4 constraints. ✗ exact length (52/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 3 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is the priority before adding referrals. Week 1 focuses on fixing failed card retries and invoice webhooks. Week 2 introduces risk scoring and tests dunning copy. Week 3 launches a beta referral program for 10% of users. The timeline ensures stability before expanding features, balancing risk and user growth.
- Repeat 45 / 10
Met 2/4 constraints. ✗ exact length (87/120); ✗ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 4 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
The plan prioritizes stabilizing core payments before introducing referrals, with Week 1 focused on fixing failed card retries and invoice webhooks, Week 2 aiming to implement risk scoring and test dunning copy, and Week 3 launching a beta referral program for 10% of users. The timeline ensures stability is addressed before expansion, with Week 1 fixes addressing technical issues, Week 2 enhancements improving risk assessment, and Week 3 pilot testing refining referral integration. This phased approach balances immediate stability with gradual expansion, ensuring reliability before broader adoption.
- Repeat 57.5 / 10
Met 3/4 constraints. ✗ exact length (63/120); ✓ importance first; ✓ timeline mentioned; ✓ no markdown
Open raw case API →Inspect repeat 5 evidence
System prompt
You are concise and literal.
User prompt
Write exactly 120 words summarizing this launch plan. Put the most important point in the first sentence. Mention timeline second. Do not use markdown. Plan: We must stabilize core payments before adding referrals. Week 1 fix failed card retries and invoice webhooks. Week 2 add risk scoring and experiment with dunning copy. Week 3 launch referral beta to 10 percent of users.
Model response
Stabilizing core payments is the priority before adding referrals, with Week 1 focusing on fixing card retries and invoice webhooks, and Week 2 introducing risk scoring and testing dunning copy. Week 3 will launch a referral beta to 10% of users, ensuring gradual rollout and monitoring performance metrics. The plan balances technical fixes with strategic testing to build trust and optimize user experience.
5 repeats recorded. Scores ranged from 10 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 110 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 1 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
Churn in small and medium businesses is a critical risk, threatening customer retention and revenue growth. Weak onboarding analytics and a backlog in the support queue further complicate customer experience and operational efficiency.
- Repeat 210 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 2 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
Churn in small and medium businesses is a critical risk, threatening customer retention and revenue growth. Weak onboarding analytics and a backlog in the support queue further complicate user experience and operational efficiency.
- Repeat 310 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 3 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
Churn in small and medium businesses is a significant risk, impacting customer retention and growth. Weak onboarding analytics and a backlog in the support queue further complicate customer experience and operational efficiency.
- Repeat 410 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 4 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
Section 2 highlights three product risks: churn in small and medium businesses, due to inadequate retention strategies, and weak onboarding analytics, which may hinder user engagement. Additionally, a backlog in the support queue could lead to delayed resolution of customer issues, impacting overall satisfaction.
- Repeat 510 / 10
Met 4/4 constraints. ✓ two sentences (2/2); ✓ section2 content; ✓ exclude section1; ✓ exclude section3
Open raw case API →Inspect repeat 5 evidence
System prompt
Follow scope boundaries exactly.
User prompt
Document with 3 sections: Section 1: market context. Section 2: product risks include churn in SMB, weak onboarding analytics, and support queue backlog. Section 3: hiring plan. Summarize ONLY section 2 in two sentences. Do not mention sections 1 or 3.
Model response
Section 2 highlights that churn in small and medium businesses (SMBs) poses a significant risk, while weak onboarding analytics and a backlog in the support queue further complicate user retention and satisfaction. Addressing these challenges requires targeted improvements to onboarding processes and support efficiency.
5 repeats recorded. Scores ranged from 10 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 110 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 1 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{ "title": "Fix onboarding analytics", "priority": "high", "owner": "data-team" } - Repeat 210 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 2 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{ "title": "Fix onboarding analytics", "priority": "high", "owner": "data-team" } - Repeat 310 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 3 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{ "title": "Fix onboarding analytics", "priority": "high", "owner": "data-team" } - Repeat 410 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 4 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{ "title": "Fix onboarding analytics", "priority": "high", "owner": "data-team" } - Repeat 510 / 10
Met 8/8 constraints. ✓ valid json; ✓ has title; ✓ has priority; ✓ has owner; ✓ correct title value; ✓ correct priority value; ✓ correct owner value; ✓ no extra key
Open raw case API →Inspect repeat 5 evidence
System prompt
Return valid JSON only.
User prompt
Return a JSON object with keys exactly: title (string), priority (high|medium|low), owner (string). Set title to "Fix onboarding analytics", priority to high, owner to "data-team". No extra keys.
Model response
{"title": "Fix onboarding analytics", "priority": "high", "owner": "data-team"}
5 repeats recorded. Scores ranged from 3.3 / 10 to 10 / 10.
- Repeats
- 5
- Completed
- 5 / 5
- Scored
- 5 / 5
- Repeat 110 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 1 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
- Repeat 210 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 2 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
- Repeat 310 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 3 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
- Repeat 43.3 / 10
Met 1/3 constraints. ✓ correct branch; ✗ single line; ✗ no extra text (20/2)
Open raw case API →Inspect repeat 4 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
The input plan is `free`, which is not equal to `enterprise`. Therefore, the else clause is executed, outputting: ``` Route: support ```
- Repeat 510 / 10
Met 3/3 constraints. ✓ correct branch; ✓ single line; ✓ no extra text (2/2)
Open raw case API →Inspect repeat 5 evidence
System prompt
Follow branch logic exactly.
User prompt
If the customer plan is enterprise, output "Route: account-exec". Else output "Route: support". Input plan = free. Output only one line.
Model response
Route: support
Evidence lane / deterministic bundle verification
Check internal bundle consistency directly.
Integrity shows whether the recorded files match their checked-in digest under this validator. It is not a signature, provider or runtime attestation, client authentication, or proof that the methodology is correct.
- Public bundle digest
- e0cda24de9a884c613c8cfc9afd9c6750791de70d2d21d71f2db13a3c30e7bd2
- Source commit
- 831639f997897dac7e58178e9ff3e978ddd3d5b1
- Completed
- 2026-07-17
- Review status
- Maintainer-self-reported