AI AGENTS: AUTOMATION
34 SRC
AI Agents: Automation
This synthesis records claims and practices from the cited sources; reported outcomes and product capabilities have not been independently verified.
Automation consolidates around parallel subagents with scoped access and hard guardrails. Scheduled Claude routines operate as 'night shift' labor with minimal setup. The approval boundary is reversibility: finish undoable work (drafting, tagging, research) autonomously; park irreversible actions (sends, spending, publishing, deletions) for human review. An n8n inbox-management agent pulls replies, classifies sentiment, checks availability, drafts responses, and pushes both to Slack for review before sending—speed-to-lead (minutes vs. hours) materially changes conversion. Setup takes roughly two minutes with a single prompt. Infra manager agents enforce hard guardrails distinct from optimization: never pause/unpause sending accounts (kills warmup) and never quarantine domains on thin data—these are treated as constants the bot never overrides. Parallel workflows require tight loops to prevent unproven changes from shipping.
Insights
Personal Automation with Subagents
- Claude Code subagents running in parallel with scoped tool access is the key capability enabling complex personal automation -- six independent workers holding different contexts simultaneously (from jimprosser chief of staff claude)
- Never let AI send emails autonomously (only draft), never make pricing decisions, default to "prep" (80% ready) rather than "dispatch" (fully handled) when uncertain (from jimprosser chief of staff claude)
- Layered automation compounds: overnight inbox scan improves morning triage, better triage enables subagent dispatch, reliable dispatch makes time-blocking viable -- 36 hours of work compounds on itself (from jimprosser chief of staff claude)
Agent-First Operations
- Marcus Moretti runs Spiral at Every as a one-person team (PM + code + support + marketing) — replaced 60% of a PM's old week with two files and a cron job: strategy.md (target problem, audience, 3-5 SMART metrics, 2-4 work tracks) and /ce:product-pulse running at 8am reading PostHog/Stripe/Datadog/database into ~/pulse-reports/ (from ai agent pm workflow spiral every)
- Now/Next/Later kanban with In Progress/Done — no sprints, no standups, no PRDs, no backlog grooming, no stakeholder updates; the agent writes tickets, moves them, and keeps statuses live (from ai agent pm workflow spiral every)
- The MCP-as-survival rule: vendors without MCPs become unusable in agent-driven workflows ("didn't have an MCP, and it was swiftly cancelled") — average company runs ~100 SaaS subscriptions and a meaningful slice now has a 12-month death timer (from ai agent pm workflow spiral every)
- New role definition: whatever the agent can't read, the PM can't use; whatever the PM can't use becomes someone else's job — the JD follows the agent's affordances (from ai agent pm workflow spiral every)
- "Agent engineer" is emerging as an internal-FDE role — extremely technical, embedded with business teams, wires up secure governed agents to Box/Salesforce/Workday and codifies workflows in skills; "agent product management" is the matching business-side role (from agent engineering roles internal business processes)
- The shift is from automating jobs to automating processes — agent engineers span teams/functions because the unit of automation is now the cross-functional process, not the role (from agent engineering roles internal business processes)
- OpenAI's "Lord Bottleneck" pattern: a growth-team staffer used Codex for individual experiment steps (analyze data, write experiment code, interpret results, produce deck), then chained them into one giant skill, then asked "do this every morning" — a self-bootstrapped cron-driven experiment loop with significant company value (from openai lord bottleneck codex automation)
- Build incrementally and personify: start with single-task acceleration (don't try to automate the whole pipeline), connect successful pieces into a skill, then schedule it; naming the system ("Lord Bottleneck") makes it approachable for the team (from openai lord bottleneck codex automation)
- Browserbase's open-source skills catalog turns web-agent automation into a playbook distribution problem: researched site behaviors become reusable capabilities instead of bespoke automation code for every target (from browserbase web agent skills catalog)
- claude-smart can reduce planning iterations and token usage by 70%+ on similar future tasks by storing reusable local learnings from previous sessions, making self-improvement an automation primitive rather than a memory side effect (from claude smart self improving plugin)
- Every automated everything it could with AI agents but still grew from 4 to 30 human employees since GPT-3, suggesting agent automation can increase the volume of valuable human coordination, oversight, and strategic work (from ai automation increases human work demand)
- Coding-agent best practices at scale can invert within six months, so agent automation programs need regular reassessment loops instead of assuming an operating style remains optimal (from coding agents large scale projects learnings)
Workflow Pattern Detection and Automation
- Use multi-source pattern detection across Codex sessions, Memories, and Chronicle to identify repeated workflows worth packaging — apply four-criteria filter: occurred 2+ times, stable inputs/outputs, material improvement potential, not already covered (from codex workflow automation prompt)
- Choose minimal viable automation forms: Skills for reusable workflows, Custom subagents for bounded specialist tasks, Automations for scheduled monitoring — choose the smallest appropriate form and skip speculative or overlapping assets (from codex workflow automation prompt)
- Structure workflow analysis with compact shortlist format: workflow description, evidence dates, frequency/confidence, recommended form (skill/subagent/automation/extend/skip), and creation rationale before building anything (from codex workflow automation prompt)
- Codex self-manages development threads by creating, searching, organizing, and pinning them — and spins up parallel worktrees for concurrent tasks, eliminating manual thread-management overhead (from codex self managing development environment)
- Sean Geng's plan-optimizer skill treats planning as a search problem by scoring plans against rubrics, critiquing them, and rewriting until score plateaus — the skill stops when improvements become noise rather than signal (from claude plan optimizer iterative improvement)
- Define clear evaluation criteria for every project to determine success, then iterate with an agent until all evals pass while capturing learnings to a memory vault for future runs (from hermes agent development workflow)
Personal Agent Requirements
- Peter Yang's framework for an ideal personal agent: works across email/calendar/Workspace/any MCP, acts proactively (cron, triggers, follow-ups), has memory that "just gets you" over time, works on web and mobile without slash commands, switches text/voice/video/live calling mid-conversation, is reachable from any messenger, and has personality that's fun to talk to (from personal agent requirements framework)
- None of OpenClaw, Claude Code, or Codex check all seven personal-agent boxes yet — the bar for "personal" agents is significantly higher than the bar for engineering agents (from personal agent requirements framework)
agent-guardrails
Explicit guardrail example: constrain an organizing agent to rename-only scope — explicitly forbid create, move, merge, archive, or delete actions — to limit blast radius of an autonomous cleanup task. (from codex workspace auto naming prompt)
Infra manager agent enforces hard guardrails distinct from optimization: never pause/unpause a sending account (kills warmup) and never quarantine a domain on thin data (a few sends in the lookback window proves nothing)—these are treated as constants the bot never overrides. (from gtm agent team grokbot astra)
agent-automation
Jason runs Codex as a 'chief of staff' operating across Slack and email, suggesting Codex can be configured for cross-platform coordination/assistant tasks rather than just code generation. (from codex work system jason openai)
Routines can be triggered conversationally by schedule (e.g. 7am daily brief) or by event (new Slack message, matching inbound email)—setup takes about two minutes with a single prompt, no workflow builder needed. (from grok bot agent teams tutorial)
safety guardrails
- OpenWorker gates consequential actions (sending messages, changing calendars, running shell commands) behind explicit user approval; unattended/scheduled runs park these requests in an inbox instead of acting autonomously. (from openworker launch)
scheduling
- OpenWorker supports scheduled automations (morning briefs, weekly reports, standing channel monitoring) that run unattended and land in the app with full transcripts for review. (from openworker launch)
agent capability growth
- Gumclaw's support-to-engineering workflow: pull a ticket from Helper, verify the customer's situation in Gumroad's internal DB, discover it's a product bug, reproduce it in the codebase, write and test a fix, open a PR, and log the customer in a follow-up ledger for another session to notify once deployed. (from gumclaw gumroad hermes agent operating system)
agent governance
- Gumclaw's permission model tiers autonomy by risk: routine work happens fully autonomously, broadcast actions (e.g., public-facing communication) require human approval, and high-stakes decisions are escalated to a person. (from gumclaw gumroad hermes agent operating system)
agent guardrails
- AGENTS.md hard rules pattern: never send messages, never self-merge, edit only repo files, change one concept at a time, cite outcomes for every proposed change, revert if eval doesn't improve — a narrow 'law' file keeps an autonomous agent contained rather than scope-creeping. (from self improving outbound codex agents)
- Eval gate pattern: fixtures.yaml encodes known-correct routing cases (including ugly/edge cases like bad-fit or weak-intent accounts), score.py computes accuracy, and a proposed config change is only accepted if it raises the accuracy score — preventing agents from accepting plausible-sounding but ineffective changes. (from self improving outbound codex agents)
- Observed failure mode: the improver agent initially wanted to raise the weight of the signal with the best-looking reply rate in a tiny outcome sample (a plausible but unproven story); the eval gate correctly rejected it since accuracy stayed flat, and only a smaller, outcome-justified edit (implementation_page_visit 4→6) actually raised eval accuracy from 0.75 to 1.00. (from self improving outbound codex agents)
- Operational cadence: run the improvement loop weekly (not after every reply) via cron/GitHub Actions to avoid overfitting to one loud account; every proposed change ships as a small PR with before/after eval scores, cited outcome rows, and required human review before merge — sending and merging must stay manual. (from self improving outbound codex agents)
agent-browser-automation
- Applied this technique to build a working Uber Eats CLI by having an agent reverse-engineer the site's network calls from a recorded HAR file. (from har derived api clients from browser agents)
self-improvement loops
- Core workflow: run two questions against recent agent sessions daily — 'what should I create from this?' and 'what should I fix so tomorrow is easier?' — separating an 'inner loop' (the actual work) from an 'outer loop' (reviewing what the work revealed). (from self improvement loop mining agent sessions)
- Lavery routes lessons found in session review into seven fixed buckets: content idea, context file (CLAUDE.md/AGENTS.md), slash command, skill, hook, tool/CLI, or config — deciding which 'home' a repeated behavior belongs in is framed as the hard part, not noticing it. (from self improvement loop mining agent sessions)
- Open-sourced tool 'agent-improvement-loop' (github.com/cathrynlavery/agent-improvement-loop) parses Claude Code (
/.claude/projects/*.jsonl) and Codex (/.codex/sessions/...) transcripts, detects real tool calls/failures/corrections (not prose mentions), redacts secrets, and stages proposals requiring human approval — it never edits anything automatically. (from self improvement loop mining agent sessions) - On its first real run across 37 sessions, the tool staged 7 proposals: one CLI fix, four skill reviews, one memory/runbook update, one backlog item — demonstrating the loop surfaces specific, actionable friction rather than abstract advice. (from self improvement loop mining agent sessions)
- Design rule for safe automation: schedule the scan (via cron or launchd, e.g. daily 7am), never schedule the changes — the model is 'controlled compounding,' where a human reviews and approves each proposal so no single bad session becomes a permanent rule. (from self improvement loop mining agent sessions)
- Manual version requires no custom tooling: paste a prompt into Claude Code/Codex telling it to read the 20 most recent session JSONL files, find patterns (content ideas, repeated errors, corrections, rediscovered setup steps), and output a numbered list of proposals with evidence lines, without applying changes yet. (from self improvement loop mining agent sessions)
loop preconditions
- Loops only earn their setup cost when four conditions hold simultaneously: the task repeats at least weekly, something can automatically fail the work (test/build/linter), the token budget can absorb retries and re-reads, and the agent has senior-engineer tooling (logs, repro environment, ability to run and inspect its own code). (from four types of agent loops)
agent-autonomy-safety
- YOLO/'Allow All' mode lets the agent execute commands without per-step approval, which is necessary for real productivity gains, but should only be run in sandboxes (e.g., GitHub Codespaces, dev containers), never on local machines with private data, due to risk of costly mistakes. (from github copilot harness workflow)
gtm agent architecture
- Connect GTM tools (Google Ads, Meta Ads, Smartlead, Apollo, Chatbase, CRM/Attio) to Claude/Codex via CLI/API — not MCP — by making each tool a folder inside one main repo, with sessions started per-folder or in the parent repo if tools need to talk to each other. (from gtm pulse agent stack)
agent-memory-workflows
- @v maintains per-user memory and personalized scheduled workflows — e.g. it periodically checks a metric (skills.sh hitting 1M skills) and proactively reminds the user, demonstrating persistent agentic monitoring rather than one-off queries. (from vercel internal agent v)
overnight-automation
- Schedule a report-only review one hour before bedtime that inspects git status, unfinished issues/TODOs, failed tests/builds, and agent handoff notes to find overnight-safe tasks—without ever touching code or starting an agent itself. (from overnight report only automation)
- Overnight task candidates must meet all six safety criteria: clear goal/definition of done, no decisions required mid-run, isolated branch or worktree, exact verification command, full reversibility, and exclusion of merging, deploying, messaging, spending, credential changes, or production data. (from overnight report only automation)
- For each overnight candidate, generate a complete Claude Code/Codex prompt specifying working directory, branch/worktree, definition of done, verification commands, and required morning handback (diffs, test results, blockers)—so execution is separated from planning. (from overnight report only automation)
- If no candidate tasks are safe, the automation should explicitly output 'No tasks are safe to run overnight' rather than inventing filler work—an explicit guardrail against fabricated agent output. (from overnight report only automation)
- Recommended rollout: start the overnight scan on a single repo; only expand scope to issue trackers and agent handoff notes later if the report consistently surfaces useful work. (from overnight report only automation)
outbound automation
- Automating the repetitive pipeline (lead finding, ICP filtering, intent matching, list building, campaign running) across multiple LinkedIn accounts is presented as the compounding lever, distinct from automating the relationship/message itself. (from you can literally 1 send this to your ai agent 2 go to sleep 3 wake up t)
agent-scheduled-reporting
- Agent Clawdito runs on a fixed cadence inside Basecamp: weekly customer sentiment write-ups and monthly summaries, all posted automatically to a dedicated 'Customer Sentiment' project. (from clawdito customer sentiment agent basecamp)
agent-autonomy
- Beyond reporting, Clawdito autonomously creates its own tracking cards on a 'What Keeps Happening' Card Table, generates its own to-dos, and posts ad hoc summaries in Chat — acting as a self-directed project participant, not just a report generator. (from clawdito customer sentiment agent basecamp)
agent-self-documentation
- Clawdito wrote a self-documenting methodology report (what runs and when) stored permanently in the project, functioning as a perpetual reference so humans can audit its own process. (from clawdito customer sentiment agent basecamp)
- Clawdito authored a companion interpretation guide for its 'State of the Customer' reports, defining the sentiment scale, frustration tiers, and other terminology needed to make its outputs legible to humans. (from clawdito customer sentiment agent basecamp)
agent-transparency
- Design pattern: consolidate an agent's outputs, self-generated tasks, methodology docs, and interpretation guides into a single dedicated project space, making the agent auditable and its work discoverable by humans without separate tooling. (from clawdito customer sentiment agent basecamp)
claude-code-routines
- Scheduled Claude routines (morning brief reading /customers and /context, Friday weekly ops review, PR review-on-open using review.md) turn Claude Code into recurring 'night shift' labor with near-zero setup time, provided production code stays untouched by default. (from claude code ai employee system)
agent-driven marketing ops
- Running many marketing agent workflows in parallel doesn't just solve the execution bottleneck—it creates a judgment bottleneck, since constant streams of competitor changes, insights, and anomalies still require human review capacity. (from solo marketer agent workflows)
- Proposed fix for agent output overload: put a prioritization filter before the agent runs and a quality gate after it produces output, following the loop: decide what matters → agent does heavy lifting → judge if output is good → ship/redirect/drop → measure → learn. (from solo marketer agent workflows)
personal-automation
- Codex automation applied to read-later apps: auto-tagging saved items based on content, reducing manual organization overhead for personal knowledge management. (from codex automation academy macstories)
agent-charters
- Bots should be defined as persistent job roles (e.g. Inbox Manager, Expense Manager) with a charter specifying what they own, what good output looks like, and explicit boundaries on what requires human approval—this is what makes unattended operation safe. (from grok bot agent teams tutorial)
agent-training
- Grok Bot supports learning-by-demonstration: performing a multi-tool, recurring, stable-step task once while the bot watches lets it save and replay the routine—best candidates are weekly, multi-tool, low-variance tasks. (from grok bot agent teams tutorial)
agent-governance
- The recommended approval boundary is reversibility, not task size: finish anything undoable (drafting, filing, tagging, research) autonomously; park anything irreversible (sending externally, spending money, publishing, deleting, agreeing to terms) for human review. (from grok bot agent teams tutorial)
loop-taxonomy
- /loop re-runs a prompt on a local time interval and stops if you close your machine; /schedule moves the same recurring loop to Anthropic's cloud as a 'routine' that keeps running independent of any open session, with schedule, API, or GitHub-event triggers. (from anthropic agent loop taxonomy)
cold-email-infrastructure
- An n8n inbox-management agent can pull replies from the sequencer, classify positive vs. negative, check calendar availability, draft a response, and push both to Slack for human review before sending—speed-to-lead (minutes vs. hours) materially changes conversion from the same reply volume. (from 10000 cold emails per day setup)
Voices
46 contributors
Peter Yang
@petergyang
Practical AI tutorials and interviews for busy people | Join 140K+ readers at https://t.co/XYKTmGVH14 | Product at Roblox
Garry Tan
@garrytan
President & CEO @ycombinator —Founder https://t.co/7aoJjp1iIK—designer/engineer who helps founders—SF Dem accelerating the boom loop—haters not allowed in my sauna
Siqi Chen
@blader
🏗️ Love to build (@runwayco @sandboxvr @zynga) people love 💸 Investor @amplitude_hq @mercury @owner @elevenlabsio @meetgamma @sfcompute @turingcom++
Tom Dörr
@tom_doerr
Follow for posts about GitHub repos, DSPy, and agents Subscribe for top posts DM to share your AI project (Due to volume of DMs I'll prioritize subscribers)
Alex Finn
@AlexFinn
Founder/CEO of Henry Intelligent Machines PBC and Creator Buddy. Building a 100 trillion dollar economic engine
Nick
@nickbaumann_
codex @openAI | prev @cline | product of @UWMadison 🦡
Vox
@Voxyz_ai
jason
@jxnlco
hype @openai
Matt Van Horn
@mvanhorn
Co-founded June ("self-driving oven" acquired by @webergrills) & the co that became @Lyft. Building again, more soon. Vibe coding @slashlast30days research tool
Sukh Sroay
@sukh_saroy
Sharing daily insights on AI, No Code, & Tech Tools • Follow me to master AI to level up your life • DM for Collabs
Charlie Hills
@charliejhills
Helping Entrepreneurs Systemise & Scale with AI | Trusted by 200k+
klöss
@kloss_xyz
AI Educator, Designer & Developer | @psychanon CEO Building AI-powered brands, workflows, and apps.
Vaibhav (VB) Srivastav
@reach_vb
Bringing Codex to developers @OpenAI | ex @huggingface | F1 fan | Here for @at_sofdog’s wisdom | *opinions my own
Samantha Trimble
@strimblez
the other sam at openai
Jonata Santos
@_jonatasantos
Building a portfolio of products at https://t.co/mFAVIoxTzS 🚀
Simon Smith
@_simonsmith
EVP Generative AI @klickhealth
Browserbase
@browserbase
give your agents access to the whole web - creators of @stagehanddev & @trydirector
Cathryn
@cathrynlavery
Founder @bestselfco ($55M+ bootstrapped). Sold to PE in 2022. Bought it back 2024. Becoming AI Native & documenting @ https://t.co/lOWxGF3Rho
Dan Shipper 📧
@danshipper
ceo @every | the only subscription you need to stay at the edge of AI
Guinness Chen
@guinnesschen
Building codex at @openai, prev @stanford, @imbue_ai
Hanako
@hanakoxbt
Nicolas Finet
@nifinet
CEO @sortlist ($1B+ generated for agencies) | https://t.co/5P8KKsKWm0 (outbound) | https://t.co/GqzxF4kSrd (intent agent)
Paras Chopra
@paraschopra
life is a game 🕹️ • building @lossfunk
Sac
@Saccc_c
探索00后的财富自由之路(全面开源成长路径,关注我,一起实现财富自由)|疯狂探索 AI 的边际和商业应用|@SWUFEBA @Ntusg
Sean Geng
@seangeng
Always building something fun 🎮 on @b3dotfun | Prev: engineering leader @Coinbase, @solana startup free components / prompts at my personal site
Simon Last
@simonlast
Building @NotionHQ
Suryansh Tiwari
@Suryanshti777
Exploring AI & SaaS trends early Sharing what’s actually useful Helping builders turn ideas → products → traction – 📩 Open to collabs
Daniel Steigman
@trekedge
Building Codex @OpenAI prev @Cline
Vibe Marketers HQ
@vibemarketersHQ
Yi Lu
@yyyiiillluuu
TL for Meta AI personalization and agent memory (Reality lab) Former head of ML at Forethought (acquired by Zendesk) Adjunct prof University of Washington
Codez
@0xCodez
Andrew Ng
@AndrewYNg
Christian
@coldemailchris
GitHub
@github
Harman
@itsharmanjot
Jason Fried
@jasonfried
Levi Munneke
@levikmunneke
Pierre-Eliott Lallemant
@pierreeliottlal
Rafal Wilinski
@rafalwilinski
Guillermo Rauch
@rauchg
Scott Schindler
@scotty529
The Startup Ideas Podcast (SIP) 🧃
@startupideaspod
dax
@thdxr
Trevin Chow
@trevin
Federico Viticci
@viticci
Yasser
@yasser_elsaid_