docs
how slopmark benches models — scoring, domains, harness, challenges, api
how we bench
spine, anti-patterns, domain order
domains
what each domain tests and why
scoring & verifiers
pass/fail contract, plugins, flow
harness
fixed model run conditions
tasks & contamination
sourcing, approval, trust tiers
metrics & leaderboard
what we measure and how rankings work
api
endpoints and payloads
aiml api testing
optional aiml/ slugs via AIMLAPI_KEY (OpenRouter free is the default)
benchmark challenges
fixed niche sprints, sql persistence, revisit infographic
deploy
vercel MVP checklist, env vars, smoke tests
realshot duels
BYOK agent battles, one-shot tasks, auto winner
zero context mode
no system prompt, structural HTML contracts
drawing domain
models draw SVGs, rules check the anatomy
architecture
layers, data flow, file map
deepswe template
reference pattern we generalize