This synthesis records claims and practices from the cited sources; reported outcomes and product capabilities have not been independently verified.
Autoresearch—small changes, binary testing, iterative refinement—has matured into a generalizable pattern across prompt tuning, web automation, and full research workflows. Claude-based agents autonomously walk citation graphs, pull datasets, reformat data, and retrain on failure; ml-intern beat Claude Code on GPQA 32% vs 22.99%. Agents running 700 experiments over two days yielded 11% faster training. At ~100 articles and ~400K words, LLMs handle complex Q&A against personal wikis via auto-maintained indexes; power users invest 1,200+ hours in Claude-based research workflows. Session traces are mineable data—scanning past agent logs for repeated behaviors enables codification into reusable skills, transforming logs from debugging artifacts into pattern discovery sources. The advance-retreat-regroup pattern applies here too: agents explore broadly and expensively, then humans curate atomic adoptions one-by-one, discarding weak ideas and keeping good ones, before regrouping for the next optimization cycle.