slopmark

the honest slop detector

slopmark

AI models are great at sounding smart and terrible at admitting when they're not. slopmark feeds them tasks with machine-checkable answers, runs every model through the exact same harness, and lets a rule-based verifier — never another AI — decide who actually passed.

489 tasks in the pool19 domains5 completed sessions0 LLM judges. ever.
FAILp 41%f 59%hmm.?!0/10

featured challenges

bench receipts — fixed traps, fixed harness, wipeouts included

all challenges →

how a score is made

most leaderboards ask a bigger model to grade the answers. that's a vibe check, not a benchmark. here every point is traceable:

  1. 01

    task drops

    a task with a machine-checkable contract comes out of the seed pool — word limits, JSON schemas, exact numbers, SVG shapes.

  2. 02

    same harness, always

    identical system prompt, temperature 0, capped tokens. no per-model babysitting, no secret scaffolding.

  3. 03

    rules decide

    a deterministic verifier checks the output against the contract and prints the rule-by-rule breakdown. no LLM judges.

  4. 04

    everything is kept

    runs roll up to the leaderboard, sessions land on the wall, and the worst outputs get framed in the hall of shame.

play

when you're bored of charts

still the same models — just doors into stupid traps instead of suite runners. canvas and thunderdome live here too.

playground hub →