Yousefi and Collins, Learning the Bitter Lesson: Empirical Evidence from 20 Years of CVPR Proceedings, arXiv:2410.09649, as current empirical pressure around Sutton’s 2019 Bitter Lesson. | Treat scale-amenable computational approaches in machine learning as a live empirical comparison pressure. | The same empirical result automatically governs modules, organizations, work arrangements, epistemes, or every other bearer. | For a computational claim, test the declared task family and scale window. For another bearer, label the move as local analogy or policy and provide its own scale predicate and evidence. |
Kaplan et al., Scaling Laws for Neural Language Models, arXiv:2001.08361, and Hoffmann et al., Training Compute-Optimal Large Language Models, arXiv:2203.15556. | Keep compute, data, model size, budget, and scale-window relations explicit when a bearer is claimed to improve with scale. | Parameter count, compute spend, or one benchmark substitutes for the audited objective vector. | State the swept dimensions, alpha and delta tolerances when used, CI or another justified uncertainty form, and budget window before preferring the general bearer. |
Lu et al., The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check, arXiv:2601.12979. | Treat current agentic-substrate claims as evaluation-sensitive and task-family-sensitive, especially when efficiency hype competes with reliability. | “More agentic” or “more efficient backbone” proves better workflow performance. | For agent-loop or substrate selection, use G.9 for parity, C.19.1 for a current scale-advantage or declared generality-policy claim, and E.23 only for repeated object-version improvement; preserve task-family evaluation, protected trade-offs, and stop/switch conditions. |