Field guide / primary domain

Autoresearch

31

sources in this field

Updated September 11, 2026

Current thesis

The shortest path to orientation.

This synthesis records claims and practices from the cited sources; reported outcomes and product capabilities have not been independently verified.

Autoresearch—small changes, binary testing, iterative refinement—has matured into a generalizable pattern across prompt tuning, web automation, and full research workflows. Claude-based agents autonomously walk citation graphs, pull datasets, reformat data, and retrain on failure; ml-intern beat Claude Code on GPQA 32% vs 22.99%. Agents running 700 experiments over two days yielded 11% faster training. At ~100 articles and ~400K words, LLMs handle complex Q&A against personal wikis via auto-maintained indexes; power users invest 1,200+ hours in Claude-based research workflows. Session traces are mineable data—scanning past agent logs for repeated behaviors enables codification into reusable skills, transforming logs from debugging artifacts into pattern discovery sources. The advance-retreat-regroup pattern applies here too: agents explore broadly and expensively, then humans curate atomic adoptions one-by-one, discarding weak ideas and keeping good ones, before regrouping for the next optimization cycle.

Evidence board

Claims worth carrying forward
01

'Spiking' workflow: send an agent (e.g. Astra/Fable) off for ~6 hours with an open-ended goal like 'make tests as fast as possible,' constrained by a hard correctness condition (golden master or existing app as oracle) to prevent silent regressions.

02

After a long unsupervised optimization run, have the agent break its own 'spike' into dozens of individual candidate changes and rank them by effectiveness and simplicity before any human review.

03

Land spike-derived changes atomically one at a time rather than merging the whole batch—manually pop items off the ranked stack, discarding weak ideas and keeping good ones. Described as more effective than all-at-once integration.

04

Pattern summary: 'Advance -> retreat -> regroup'—let the agent explore broadly and expensively, then retreat to human-curated atomic adoption, then regroup for the next spike cycle.

05

Set up an autoresearcher to scan through past agent session traces specifically looking for repeated patterns and behaviors, then codify those into reusable skills.

06

Session traces themselves are a mineable data source for skill-building — treating agent logs as raw material for pattern discovery rather than just debugging artifacts.

07

Karpathy's premise: any metric cheaply evaluated can be handed to an agent swarm — demonstrated by pointing an agent at his own training code for two days, running 700 experiments, keeping 20 that beat benchmark, yielding 11% faster training.

08

Reply rate qualifies as a cheaply-evaluable metric, making outbound GTM a candidate for the same hill-climbing agent loop used in ML experiments.

Adjacent fields

Key voices

Latest evidence

Recent additions

All synthesized insights →

Nicolas Finet

Building a Self-Improving Outbound System on Codex

A detailed architecture for a GTM system where Codex reads weekly outcome logs, proposes single scoring or prompt changes backed by evidence, validates them against an eval gate, and opens a PR for human review — keeping sending and merging strictly outside the automated loop.

Sean Geng

Plan-Optimizer: Iterative Loop Skill for Claude Code

Sean Geng packaged @goodalexander's iterative plan-improvement technique as an installable Claude Code skill. The skill treats planning as a search problem: generate a plan, score it 0-100 against a rubric, critique specific weaknesses, rewrite, keep the best, and stop when scores plateau over a sliding window. One curl command installs it into ~/.claude/skills/plan-optimizer/.

Jonata Santos

Full-Stack Personal Agent Workflow with Hermes, GBrain, and Orca

A concrete personal-agent stack: Hermes on a VPS (Hetzner/DO/Hostinger), GBrain or QMD+SQL memory vault, GPT-5.5 fast mode via Codex auth, Orca IDE on macOS/iOS via Tailscale. The workflow gates every project with deep research, `/grill-me` unknowns surfacing, written evals, and iterative runs until all evals pass — saving learnings to the vault throughout.