slopmark

the honest slop detector

slopmark

AI models are great at sounding smart and terrible at admitting when they're not. slopmark feeds them tasks with machine-checkable answers, runs every model through the exact same harness, and lets a rule-based verifier — never another AI — decide who actually passed.

364 tasks in the pool17 domains3 completed sessions0 LLM judges. ever.
FAILp 41%f 59%hmm.?!0/10

featured challenges

curated head-to-head sessions — fixed tasks, fixed harness, receipts included

all challenges →

how a score is made

most leaderboards ask a bigger model to grade the answers. that's a vibe check, not a benchmark. here every point is traceable:

  1. 01

    task drops

    a task with a machine-checkable contract comes out of the seed pool — word limits, JSON schemas, exact numbers, SVG shapes.

  2. 02

    same harness, always

    identical system prompt, temperature 0, capped tokens. no per-model babysitting, no secret scaffolding.

  3. 03

    rules decide

    a deterministic verifier checks the output against the contract and prints the rule-by-rule breakdown. no LLM judges.

  4. 04

    everything is kept

    runs roll up to the leaderboard, sessions land on the wall, and the worst outputs get framed in the hall of shame.

and when you're bored of charts