Public record / ongoing inquiry
Tech Lab Reports
Hypotheses, methods, and results from our ongoing experiments in collective intelligence—published so the work can be examined.
- Published
- 03
- Series
- Active
Research archive
Published reports
Each report preserves the premise, conditions, and baseline used when the work was published.
Connect 4, measured: Astra 2 – Fable 0
This companion to ArcadeBench report 002 documents an invite-only, untimed Connect Four best-of-3 played on BarKade on September 8, 2026. It is the first concrete step toward defining ArcadeBench for Connect 4, not a full leaderboard claim.
ArcadeBench v0
Most evaluations score an AI's output against an opinion: a human rater, a written rubric, a second model acting as judge, or a test that someone had to author. Those instruments are useful, and they are also approximate — the judgment is part of the measurement.
The Joule Hypothesis
Artificial intelligence is usually presented as a capability: what a model can write, reason through, create, analyze, or play. The resource consumed to produce that work is often hidden behind subscriptions, token limits, and provider bills.
