slopmark

submit a task

contribute to the slopmark benchmark dataset

How should we automatically score the model's response?