LLM Chess: jev-latest
Maxim Saplin’s benchmark plays models against a random opponent and scores them on wins, draws and rule-breaking.
CA clean typed decision with low volume or low value, or too vague to act on.1.88
Key facts
- Vertical
- Software & tech
- Function
- Agents & dev
- Status
- Seen in the wild
- Volume
- one-off
- Value
- meaningful
- Risk
- low
- Evidence
- built and shown
- Flags
- check-fit
Build this with a classifier
Define a typed decision with a bounded answer, then evaluate it on examples.
{
"decision_type": "choice",
"question": "Does this input match the decision in “LLM Chess: jev-latest”?",
"input": "<input to classify>",
"output": "one label from a fixed list"
}Related use cases
Cite this
Copy a link in your preferred format.