Papers
3 · research, read for youResearch read for you: I take a paper that matters, explain why it matters to you and where to go deeper — no jargon, with links to the originals.
The Missing Half: How Complete Is AI-Generated Writing?A Meta paper builds a benchmark for measuring factual completeness in generated texts. The best model covers 58% of the facts it should.Paper: Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenessChain-of-Thought: why "reason step by step" actually worksThe 2022 paper that gave a name to the most-used trick in prompting, and showed it only works from a certain scale up.Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsWhen the agent forgets what mattersA proactive memory agent that decides what to keep and when to inject it into long-running tasks, where relevant state gets lost beyond the context window.Paper: Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents