ClawbotomyEvidence lab

Evidence lane / Configured-agent session

See how your agent behaves before it gets more power.

Clawbotomy records fixed synthetic tasks against one operator-selected runtime. Run the open-source workflow yourself, or add a human reviewer to the same local evidence workflow.

A checkup records one observed session. It is not a certification, production guarantee, or automatic permission decision.

One observed sessionSynthetic tools onlyPrivate receiptsHuman decision required

Choose the level of help

Two ways to run the same checkup.

Start self-serve. Add guided review when the runtime, failure modes, or permission decision deserve another pair of eyes.

  1. [01]
    Self-serve

    Open-source workflow

    Build the plan in your browser, connect a checked-in OpenClaw or Hermes bridge, validate the bundle in your terminal, and inspect its local projection.

    Includes
    • Browser-local planning
    • Synthetic Inbox tools
    • Private evidence inspection
    • No hosted agent or mailbox
    Start a checkup
  2. [02]
    Guided review

    Agent Behavior Checkup

    Work through one configured runtime with a human reviewer. Freeze the scope, run the cases, separate behavior from infrastructure, and leave with a decision packet.

    Includes
    • One frozen runtime and plan
    • Failure-cluster review
    • Evidence and limits packet
    • Recommended next controlled change
    Discuss a guided checkup

The checkup loop

Plan. Connect. Inspect. Decide.

Failed infrastructure is not scored as behavior. An invalid arm ends the comparison instead of quietly becoming a zero.

  1. [01]

    Plan

    Record intended powers, pin the runtime, and freeze the cases before a provider call.

  2. [02]

    Connect

    Run a checked-in OpenClaw or Hermes bridge from your own machine. Exact-pin compatibility is checked separately.

  3. [03]

    Inspect

    Validate the private bundle in the terminal, then load its allowlisted browser projection.

  4. [04]

    Decide

    Use the bounded evidence as a review input. Permission decisions remain human.

Optional next stage

Retest only after a valid baseline.

Change one behavior, keep the plan and comparison contract fixed, then rerun. If either arm is invalid or the evidence cannot be compared honestly, stop as inconclusive.

Read the evidence model →

Claim boundary

Useful evidence, kept in its lane.

What you get
  • A frozen, reviewable test plan
  • Private launcher and bundle receipts
  • Observed failures separated from runtime failures
  • A bounded next-step recommendation
What you do not get
  • A universal agent score
  • A safety certification
  • Hosted access to your production systems
  • Automatic permission or deployment changes

Start with the smallest honest check

Freeze one plan. Observe one runtime.