slopmark
bench
playground
challenges
sessions
leaderboard
shame
docs
submit a task
contribute to the slopmark benchmark dataset
domain
instruction
json
math
sycophancy
agentic
safety
coding
writing
swe
prompt
evaluation rule (verifier)
How should we automatically score the model's response?
word limit (max)
forbidden word
required phrase
submit task