Evidence method
How AI Model Waddle handles evidence
Our practical standard for separating public facts, direct measurements, operating experience, inference, and community reports.
Evidence should travel with the claim
AI systems change quickly, and confident language ages badly. Our job is to make the basis of a claim visible: what is documented, what we measured, what we experienced, what we inferred, and what someone else reported.
That separation matters more than sounding certain.
Five states, five different jobs
Public fact
A verifiable statement from a named public source. Vendor claims remain vendor claims until independent evidence supports more.
First-party measurement
A result produced by a disclosed method, environment, sample, instrument, and time window. A measurement only speaks for its test conditions.
First-hand operating experience
What we observed while using or implementing a system for a stated task. One operating experience is useful evidence, not universal performance.
Inference
A reasoned interpretation built from identified facts or observations. We label it as interpretation and keep alternative explanations in view.
Community report
An attributed external experience we have not independently verified. It can reveal a pattern worth testing, but cannot impersonate direct evidence.
What a practical test needs
A Waddle Test starts with a real decision question, not a product ranking. It records the task, system version, access tier, configuration, material human involvement, method, observations, failures, and limits. Where safe and useful, it points to inspectable artifacts.
Negative findings stay in the record. “Did not work under these conditions” is not rewritten as “cannot work.”
How updates work
We show published and updated dates on every field note. Material changes to access, behavior, price, limits, evidence, or ownership trigger review. Change notes explain revisions when the history helps readers make a current decision.