Funny Eval
in progressLeaderboards don't show what a model can actually do. I want to force the gap into the open visually and playfully — a task is worth building only if you can see who won at a glance.
- Four granularities, tested apart: context, tools, model, agent — each for the ability it owes
- Only the steps that open the widest gap; if every model clears it, it isn't a question
- Aimed at what resists quantification: the abilities no score reports are the ones I want to see
- Tasks are black boxes: no rules given, only traces — the model has to reconstruct the world
- The eval has to be fun to watch, or nobody watches — myself included
Projects1
<div class="pj">
idea
Black-box mazeGive the model only a ball's trace and ask it to reconstruct the maze
idea
</div>