TMLR reached out to the authors of 10 papers slated for desk rejection, in an attempt to understand if the authors could explain the paper they submitted [D]
LARA: small, composable behaviours for frozen LLMs [P]
GoBench: Evaluating LLMs on the game of Go [R]
Has anyone measured specification ambiguity as a predictor of correlated failure across model families? [D]