I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]
acl arr august 2026 (desk rejected ) [D]
Why does lightgbm not fit my toy example but catboost does? (2 order interactions) [D]
The evaluation resolution has been shown to have a significant impact on the identification of the "learning rule" that exhibits the most brain-like characteristics at V1. [R]
Does telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R]