VIBE CODING
62 SRC
Vibe Coding
This synthesis records claims and practices from the cited sources; reported outcomes and product capabilities have not been independently verified.
Vibe coding has matured into production-scale software where frontier models handle complex tasks autonomously through supervised orchestration. Ultracode mode in Opus 4.8 removes manual intervention by enabling Claude to invoke workflows independently; supervisory workflows delegate subgoals, route routine execution to cheaper models, and enforce quality gates. Two-model adversarial loops—one drafting, one reviewing—prove effective; GPT-5.5 consistently finds issues in both planning and code review. At ~$400/month for Opus 4.7 + GPT-5.5, end-to-end feature work costs equivalent to fractional dev teams. Design specs via DESIGN.md achieve 95%+ principal completion rates. Strong prompts engineer state traps explicitly, paste raw errors, and specify mode; written rules in instructions files have highest leverage. For large features, split into planning, specification, then parallel execution. GPT-5.6 Sol-class models exhibit a distinct failure mode: not incorrect code, but overbuilt code where a config change produces full frameworks and bug fixes produce unnecessary adapter layers. Overengineered AI code often passes tests cleanly, hiding problems until modification attempts. Avoiding committed media in PRs keeps repo size clean while preserving reviewer visibility. Humans read every diff before commit.
Guides
Claude & Claude Code for Marketing Agencies: A Detailed Guide
A comprehensive guide covering agency workspace setup, content strategy, competitive intelligence, outbound automation, brand design workflows, and agency-specific Claude Code skills — grounded in 77+ sources.
Keyword Research → Optimized Landing Pages: The Workflow
A four-stage workflow using Claude Code to turn a keyword list into optimized landing pages — covering competitive intelligence, content briefs, one-shot page generation with design constraints, and automated monitoring loops.
Insights
Quality and Craft
- AI is strong at filling in UI details but fails at creating visual structure from scratch -- the fix is to provide pre-built layout skeletons from professional UI block libraries like Tailark, Tailwind UI, or shadcn blocks (from ai ui layout technique)
- The key technique for professional vibe-coding output: copy a layout template's code, paste it as a constraint for the AI, and instruct it to follow that exact structure rather than generating layout from scratch (from ai ui layout technique)
- Layout quality is the single biggest tell separating professional-looking AI-built apps from obvious "AI slop" -- spacing, hierarchy, and visual flow are where AI falls short without human-curated constraints (from ai ui layout technique)
- UI polish still requires deep care and design expertise that current AI models cannot substitute -- you can generate functional UI with AI but fluid, polished interactions need human taste and attention (from vibe coded ui beats openai codex)
- Craft Agents was 100% vibe-coded in 4 weeks (including holidays) and outpolished OpenAI's Codex in UI quality, demonstrating that AI-assisted development with design expertise beats large teams without it (from vibe coded ui beats openai codex)
- Claude produces poor UI when freestyling dashboard designs; the fix is to constrain it with an existing reference design rather than letting it generate from scratch (from claude dashboard ui design hack)
- A concrete prompting pattern for better AI-generated UI: ask the model to identify enterprise-grade open-source dashboards as references, pick one, then have it map your features onto that design system (from claude dashboard ui design hack)
- A growing pattern among vibe coders: screenshot UIs from Dribbble, feed them to Claude as design inspiration, and have it generate CLAUDE.md style guides -- producing professional-looking products without hiring designers (from dribbble design reference for claude)
- A Figma plugin that takes reference designs + brand guidelines and generates editable vector SVGs directly on the canvas compresses hours of manual illustration work into 30 seconds for first drafts (from figma plugin ai ui generation)
- Generating real text, real layers, and fully editable vectors (not raster images) is the critical quality bar for AI design tools to be useful in production workflows (from figma plugin ai ui generation)
One-Shot Shipping and Speed
- Opus 4.6 can generate complete, production-quality websites in a single prompt ("one-shot"), a capability threshold that shifts vibe coding from iterative refinement to instant shipping (from opus one shot website)
- Opus 4.6 demonstrates strong capability in generating marketing-specific UI -- landing pages, hero sections, and conversion-focused layouts -- which are high-leverage for non-technical founders (from opus marketing ui)
- The "mega-prompt" pattern for vibe coding bundles all requirements (SEO, functionality, design) into a single comprehensive prompt rather than iterating, optimizing for speed-to-ship over refinement (from vibe coding prompt opus)
- One-person teams can out-execute the largest labs on UI quality thanks to AI tooling -- this is a leverage inversion where small + skilled + AI beats large + resourced (from vibe coded ui beats openai codex)
- Designer-developers who both design and build features end-to-end are showcasing their work with live demos and video walkthroughs -- the "designed and built" framing signals the growing expectation that individual makers ship complete features (from performance feature ui build)
Accessibility and Non-Technical Builders
- A non-technical person achieved 5.8K GitHub stars by building a Claude Code skill entirely through vibe coding, demonstrating that meaningful developer tools can now be created without traditional programming knowledge (from zara zhang non coder github stars)
- The framing of "code as a medium for storytelling" positions vibe coding not just as a productivity hack but as a new creative medium accessible to non-engineers (from zara zhang non coder github stars)
- Vibe-coded projects can go viral on X and earn significant GitHub traction, suggesting the community values useful output regardless of whether the creator wrote code manually (from zara zhang non coder github stars)
Tooling and Infrastructure
- One-click integrated setup (GitHub + Convex + Vercel) eliminates infrastructure decisions entirely, making vibe coding accessible to non-technical builders (from convex vibe code setup)
- The vibe coding stack is converging on a pattern: managed backend (Convex) + managed deploy (Vercel) + AI editor (Cursor) as the zero-config trinity (from convex vibe code setup)
- Google Docs-style auto-save and link sharing for code projects enables real-time multiplayer collaboration on vibe-coded apps (from convex vibe code setup)
- OpenClaw offers one-click deployment completed in under 1 minute, targeting non-technical users -- deployment complexity is being abstracted away entirely (from openclaw one click deploy)
- Nvidia offers free API access to Kimi K2.5 model, removing cost barriers for experimenting with alternative LLMs in vibe-coding workflows (from lobster kimi k2 free nvidia)
Scaling Blind Spots
- Claude Code has implicit tech-stack defaults (Supabase, pgvector, Vercel, Node.js, Celery) that optimize for developer experience over production cost-efficiency, creating a blind spot for vibe coders who accept suggestions uncritically (from vibe coding scaling blind spots)
- The gap between prototyping and scaling with AI-generated code is not in code quality but in infrastructure choices -- LLMs suggest what they've seen most in training data, not what scales best (from vibe coding scaling blind spots)
- Specific stack critiques at scale: Firebase cheaper than Supabase, Pinecone outperforms pgvector, Modal/Lambda beat Celery, FastAPI over Node.js, Render cheaper than Vercel, Expo over raw Xcode (from vibe coding scaling blind spots)
- Vibe coding's core limitation is that production infrastructure intuition comes from operating at scale, not from prompting -- this is knowledge LLMs cannot reliably surface because it is contextual and opinionated (from vibe coding scaling blind spots)
Strategic Focus Areas
Focus on product/UI, customer acquisition, integrations, and agent infrastructure -- parallelize work and speed up verification loops rather than chasing rapidly changing AI benchmarks (from rapid change and enduring constants in ai)
Invest in training users, connecting to their systems, and working with their data rather than building complex in-house AI models (from rapid change and enduring constants in ai)
Multi-stage automated development pipeline: Feature Request Chat → Manual Spec Review → Spec Ready → Scheduled Nightly Implementation → PR Ready → Fresh-Context Code Review — all Claude, zero developer execution (from claude automated feature request workflow)
The Paradigm Shift
- December 2025 was a discrete inflection point for coding agents: models crossed a quality/coherence threshold where agents went from not working to basically working for large multi-step tasks (from karpathy coding agents paradigm shift)
- The new programming paradigm is spinning up AI agents with English-language task specs, then managing and reviewing their work in parallel rather than typing code in an editor (from karpathy coding agents paradigm shift)
- The vast majority of AI coding tool users are extremely light users -- 40 minutes/day puts you in the top 0.01%, suggesting massive headroom for deeper adoption and workflow integration (from taskmaster claude code marathon)
AI Image Mocks Replace Prototyping
- AI image-generation models are now strong enough at digital surfaces that internal teams share product ideas via generated screen mocks instead of building click-through prototypes — Codex with image gen becomes a full-stack design engineer in one tool (from ai image generation ui prototyping codex)
One-Person Operations Out-Execute Teams
Marcus Moretti runs Spiral at Every as a one-person team (PM + code + support + marketing) — strategy.md plus a /ce:product-pulse cron at 8am replaces 60% of a PM's old week; vibe-coding's productivity inversion now extends across the entire product function, not just engineering (from ai agent pm workflow spiral every)
This pattern adds two layers of quality control: built-in steering (the parent thread guides the sub-thread's goal) and taste verification (the parent thread checks the sub-thread's output before accepting it), relates to Vibe Coding (from codex babysit subgoal prompting technique)
For current frontier models, replace prescriptive step-by-step prompts with four elements: the goal, relevant context, boundaries, and a concrete definition of done — relates to Workflow and Vibe Coding (from frontier model prompting orchestration patterns)
The device keeps pinned chats visible while exposing repeat actions as hardware controls, suggesting agent workflows can benefit from a tactile command surface instead of relying entirely on keyboard shortcuts — relates to Vibe Coding (from meet kbd 10 codex micro built with worklouder map the)
prompting-technique
- Because GPT-5.6-sol tends to keep going far longer than needed, prompts benefit from explicit 'stop points' — e.g., instructing it to write a plan and pause for feedback before proceeding, or to build/test/PR and stop after the first round of review comments rather than continuing indefinitely. (from gpt 5 6 sol codex usage tips)
model-tiering
- Core design pattern: use your most capable (expensive) model for the parts where intelligence compounds — understanding the codebase, judging what's worth doing, writing the spec — and delegate execution to cheaper models. This splits cognitive load from labor cost in agentic coding workflows. (from shadcn improve agent skill)
unknowns-taxonomy
- For unknown knowns (tacit taste/criteria you recognize but can't verbalize), use cheap brainstorms and prototypes—e.g. requesting 4 wildly different HTML design directions or a mocked UI with fake data—rather than direct questions, since verbalizing tacit criteria upfront is unreliable and reverting mid-implementation changes is expensive. (from grill for unknowns agent skill)
agentic coding prompt template
- Full prompt template for Codex + GPT-5.6 Sol: give the model one complete spec/task and instruct it to 'take the whole plan to done' until architecture, implementation, tests, review, and final result all clear the bar—rather than stopping at partial progress. (from codex parallel subagent prompt template)
agent operator orchestration
- The model is directed to run itself like an operator: schedule agents in parallel, track their progress, synthesize returned work, resolve conflicts between subagent outputs, keep implementing, verify live after every important step, review when it matters, commit when ready, and close with a final summary. (from codex parallel subagent prompt template)
AI coding agent design skills
- Taste Skill is an open-source SKILL.md-based skill package designed to stop AI coding agents (Cursor, Claude Code, Codex, Gemini CLI, v0, Lovable, OpenCode) from generating generic 'AI slop' frontend designs; it plugs into any tool that supports SKILL.md files. (from taste skill anti slop frontend framework)
- Install command:
npx skills add Leonxlnx/taste-skill --skill "design-taste-frontend"— a single command works across all compatible agents; existing users auto-upgrade to v2 on next install. (from taste skill anti slop frontend framework) - Taste Skill v2 (experimental, a 2026 rewrite) is now the default and adds: §0 brief inference (reads industry/audience/mood/motion before generating), §2 brief-to-design-system mapping (choosing among Material, Fluent, Carbon, Polaris, Atlassian, Primer, GOV.UK, USWDS, Bootstrap, Radix, shadcn, Tailwind, or native CSS), §8 dual-mode dark mode protocol with contrast/hierarchy parity, §11 audit-first redesign protocol with preservation rules, §12 block library schema for coherent iterative additions, and §14 a hard pre-flight checklist every output must honestly pass before shipping. (from taste skill anti slop frontend framework)
- output-skill addresses a known AI coding agent failure mode by explicitly preventing placeholders, skipped sections, and half-finished work in generated output. (from taste skill anti slop frontend framework)
unknowns-techniques
- 'Blind Spot Pass' technique: when entering unfamiliar codebase areas or domains, explicitly ask Claude to run a 'blindspot pass' to surface your unknown unknowns, giving it context on your background/experience level so it can tailor findings. Example: 'I'm working on adding a new auth provider but know nothing about the auth modules in this codebase—do a blindspot pass to help me figure out my relevant unknown unknowns.' (from fable field guide finding unknowns)
- Brainstorm-and-prototype technique: for areas involving unknown knowns (criteria you'd recognize but can't articulate), ask Claude for multiple divergent design/implementation directions (e.g., an HTML mock with 4 wildly different layouts) before wiring up real code, since reverting deep implementation changes is expensive. (from fable field guide finding unknowns)
Claude vs Codex Capability Split
- Claude is described as WAY better than Codex at computer use, UI/UX verification, and general-purpose agent execution — Codex excels specifically on well-specified, deterministic work. This capability split drives the routing decision. (from theo claude codex workflow)
Automated PR Quality Gates
- Burden-of-proof enforcement: every subagent must complete a self-review, pass a bug bot (Cursor AI), produce elegant fixes, and submit a recording before the PR is accepted. The middle manager enforces this gate, preventing low-quality merges at scale. (from vinvan fable software factory)
AI Development Bottleneck Shift
- 60+ PRs in one overnight run creates a new bottleneck: human review capacity, not generation capacity. The author identifies review throughput as the next unsolved industry problem for AI-driven software factories. (from vinvan fable software factory)
Portfolio Construction for AI Engineers
- Three portfolio projects are the minimum viable credential for AI engineering: (1) RAG app with evaluation, (2) multi-agent system solving a real problem, (3) deployed system with monitoring and cost tracking. Each should be documented as a case study covering problem, approach, measurements, and retrospective. (from ai engineer path no degree)
Git-Integrated Experimentation
- try supports git worktrees (try . [name] creates a dated worktree dir for the current repo) and repo cloning (try https://github.com/user/repo.git creates 2025-08-27-tobi-try). These extend it from scratch experiments to fork-and-explore workflows. (from try cli experiment management)
Experiment Deployment Pattern
- @seflless extended try into a 'tries' command pointing at a single GitHub repo, with GitHub Pages auto-hosting static HTML demos. This transforms a local experiment organizer into a shareable portfolio with one push command. (from try cli experiment management)
Autonomous QA loop
- A single Codex prompt can chain three autonomous phases: (1) enumerate all features from code → user stories in a canonical spreadsheet, (2) test every user story and document errors, (3) fix all UX/logic errors and retest. Hundreds of user stories processed without manual intervention. (from codex loop automation user stories)
Human-in-the-loop AI editing
- The diff-tracking design of Flashtype treats each Claude/Codex edit as a discrete, reviewable change — a pattern that surfaces AI edits as transparent, reversible deltas rather than opaque rewrites. This positions it as a trust-layer tool for human-in-the-loop document editing. (from flashtype markdown editor claude codex)
Systematic Product QA Loop
- The 'full product evaluation loop' (Featured) builds sanitized production-scale local data, inventories every user-facing feature/role/route/button/modal/state, tests as a real user, then implements coherent fixes with regression tests before a full re-run—explicit stop: clean pass or blocked handoff. (from loop library agent workflows)
Disciplined Bug-to-PR Pipeline
- The 'ticket-to-PR-ready loop' (by Hiten Shah) mandates: reproduce failure in smallest representative environment, prove root cause, make smallest credible fix, rerun reproduction + regression tests. If unreproducible after two serious attempts, explicitly say so. No unrelated refactors folded in. (from loop library agent workflows)
AI-native design workflow
- A principal engineer (formerly designer) now does 95%+ of design work in coding harnesses and the terminal. Workflow: (1) AI generates a design.md spec, (2) AI generates components from it, (3) iterate via feedback until it feels right. No Figma required. (from designer to engineer ai workflow)
Design skill evolution
- Designing via AI in a code-native terminal environment (markdown specs → components → iteration) is framed as a core emerging skill for designers and builders broadly, not a niche developer trick. (from designer to engineer ai workflow)
Adversarial prompting as pre-build gate
- Adversarial thinking as an explicit prompt instruction ('think adversarially') surfaces failure modes before implementation. Applied after literature review, it stress-tests the idea against the field's accumulated knowledge rather than intuition alone. (from fable adversarial research)
Human-Agent Feedback Loop
- The app targets the review layer of agent-generated code/content — a gap where agents produce output but structured human feedback loops are absent. Roughdraft positions markdown as the collaboration surface between human and agent. (from roughdraft markdown review agent)
Eval-Driven Agent Iteration
- Define explicit evals before starting any agent project run. Have Hermes iterate until all evals pass, logging learnings to the vault continuously. When blocked, use memory + fresh research or a
/grill-mesession to unblock. (from hermes agent build workflow)
Parallel Task Execution
- Codex can spin up git worktrees autonomously for parallel tasks, enabling concurrent workstreams without manual branch or worktree setup. This mirrors the Claude Code pattern of one agent per worktree but executed self-directed rather than user-initiated. (from codex self managing threads)
Ultracode autonomous orchestration
- Ultracode mode in Opus 4.8 sets thinking to xhigh and lets Claude autonomously decide when to invoke Dynamic Workflows, removing the need to manually select orchestration strategy for complex coding tasks. (from claude opus 48 launch)
Code Review Standards
- Google's eng-practices repo (github.com/google/eng-practices) publishes two distinct guides: one for the reviewer and one for the change author — making it a bidirectional standard, not just a checklist for fault-finding. (from google eng practices code review)
Agent Best Practices Drift
- Simon Last's thread on large-scale coding agent usage explicitly contradicts advice from 6 months ago — a signal that best practices in this space are inverting rapidly and prior guidance should be treated as perishable. (from coding agents large scale lessons)
Parallel AI development
- Git worktrees enable parallel AI execution: isolate multiple feature branches (auth, UI, bug fixes, experiments) running simultaneously without touching main. Once experienced, sequential single-session development feels slow by comparison. (from claude code ecosystem setup)
CI/CD-embedded AI
- Integrating Claude into CI/CD pipelines—PR code review, fix suggestions, standard enforcement, architecture rule compliance—moves AI from occasional coding help to embedded development lifecycle infrastructure. (from claude code ecosystem setup)
Goal specification failure modes
- A weak /goal is a wish ('improve onboarding'); the agent optimizes for whatever is easiest to prove—cleaner screenshots, passing tests, fewer steps—without any product improvement. A loop can spend 40 turns making the wrong thing more internally consistent. (from goal command agentic coding pm)
Loop calibration discipline
- Watch the first loop iterations before stepping away: if the agent misreads the target, touches out-of-scope files, writes tests that bless wrong behavior, or keeps asking the same question (a spec ambiguity humans were silently resolving), stop and fix the spec before running unattended. (from goal command agentic coding pm)
Effort vs. outcome delegation
- The principle: 'A prompt asks for effort; a contract defines the condition where effort stops.' PM work shifts from prose intent to executable definition of done—teams with crisp acceptance criteria get useful loops, teams with mushy requirements get longer, faster mush that burns tokens. (from goal command agentic coding pm)
Agent execution discipline
- PLAN section enforces understand-first execution: restate understanding before non-trivial changes, prefer minimal sufficient changes over broad rewrites. DONE WHEN requires a verifiable completion state, not a subjective judgment. (from goal command structure codex claude)
Verifiable agent completion
- VERIFY section must include: tests/build/lint/typecheck/manual validation; explicit statement of what could not be verified and why; and a rollback/containment plan for destructive or high-risk changes—making incompleteness visible rather than hidden. (from goal command structure codex claude)
Device-agnostic AI agent infrastructure
- The author frames this makeshift setup as a preview of near-future work: AI coding agents become ambient infrastructure accessible from any device rather than tools tied to whichever machine is open. The friction today is SSH/Wi-Fi reliability, not the model. (from codex satellite home device topology)
Human-feedback flywheel
- River's PR merge rate improved from 36% to 77% over two months with no model change or retraining. The gain came entirely from employees watching River work, identifying failure points, and writing corrections into River's skills and memory. Collective human feedback beat model switching. (from shopify river lehrwerkstatt)
Codex Plugin Ecosystem
- Chrome Browser and Hyperframe are both Codex plugins, enabling direct integration of browser control and frame/layout tooling into Codex-based AI creative workflows without custom integration work. (from codex plugins chrome hyperframe)
Automated Code Quality Enforcement
- Custom precommit linters should auto-fix problems via --fix or shell out to a cheaper LLM (Composer 2.5 or Sonnet) to fix them—not just flag them. The goal is committed clean code, not a warning list. (from im just going to dump my whole agentic setup out here because i see too many p)
Cross-Agent Review Patterns
- Cross-agent review at each major phase (research, plan, implementation, wrap-up) using different models prevents same-model blind spots. Personas—maintainability, security, performance, domain expert—each own a set of system docs and keep them updated. (from im just going to dump my whole agentic setup out here because i see too many p)
Test Integrity Auditing
- A false-confidence test audit skill periodically scans recent tests to find ones that aren't testing what they claim to test, and fixes them. This directly addresses the problem where passing tests create false safety signals. (from im just going to dump my whole agentic setup out here because i see too many p)
AI-native design systems
- The recommended workflow: don't build a design system from scratch. Clone the design language of a brand you admire (Linear, Stripe, Vercel), extract it into a Design.md using ChatGPT or Claude, then layer specialized skill files on top (landing page skill, mobile skill, motion skill) that all reference the same Design.md. (from design md ai design workflow)
Design consistency failure mode
- The critical failure mode in AI-assisted UI work: one screen looks polished, everything else looks generic. Design.md's structural fix is centralizing the aesthetic contract in one file so every subsequent generation—regardless of medium or agent—stays consistent. (from design md ai design workflow)
Multi-model adversarial review
- Two-model adversarial review loop: Claude Opus 4.7 drafts the feature plan, GPT-5.5 reviews and finds issues, Opus updates until GPT approves, then Opus builds, GPT reviews code, Opus fixes, GPT signs off. Playwright handles automated UX/UI testing between steps. (from multi model review workflow zook)
Heterogeneous model review quality
- GPT-5.5 consistently finds issues in both the plan phase and code review phase when acting as a second-opinion reviewer over Claude Opus output—suggesting heterogeneous model pairs catch more bugs than single-model loops. (from multi model review workflow zook)
AI dev team cost benchmark
- ~$400/month for Claude Opus 4.7 + GPT-5.5 via Conductor Build covers end-to-end feature planning, code generation, automated testing, and adversarial review—framed as equivalent cost to a fractional dev team with no push-back on small UI changes. (from multi model review workflow zook)
SaaS displacement via vibe-coding
- Codex can replace a $30/mo SaaS (Superhuman) with a fully customisable agent-native email client running on Gmail CLI, replicating UX patterns while removing vendor lock-in. (from bentossell diy email client codex)
AI-built SaaS alternatives
- The pattern: identify a subscription SaaS with well-understood UX conventions, replicate those conventions via an agent-built CLI tool, and replace recurring cost with one-time build effort. Superhuman → Gmail CLI is the worked example. (from bentossell diy email client codex)
Hybrid Desktop AI Stack
- Tech stack: Tauri 2 (desktop shell, auto-update, cross-platform), React 19 + TypeScript + Vite (frontend), Python 3.13 + FastAPI + WebSockets (backend sidecar), SQLite (CRM), KuzuDB (profile graph), LanceDB (vector store), Playwright (experimental automation only). Thin installer ~100MB; heavy runtime (browser + vector libs + ONNX model) downloads once on first run. (from justhireme local first job intelligence)
model-training-vs-harness
- Core thesis: no amount of harness engineering, loop tuning, or prompt-magic ('adversarial review', more linters) can fix models' inability to maintain and improve codebase quality over time — this is a model-training/RLVR and benchmark gap, not a skill issue. (from why software factories fail)
benchmark-gaps
- There are currently no good benchmarks for measuring a model's ability to maintain codebase quality over time, even though models ace one-off coding and greenfield benchmarks — a key blind spot in frontier model evaluation. (from why software factories fail)
agentic-coding-quality-decline
- Faros AI report (since teams adopted AI coding tools ~Jan/Feb 2025): PR review quality dropped — more/longer comments, more PRs merged with zero review, incidents up, bugs per developer up. Treated as correlational, directionally valid signal, not proof. (from why software factories fail)
- HumanLayer went fully 'lights-off' (background agents, no code reading) in July 2025 and hit unsolvable issues repeatedly; after the third major incident (Nov 2025) the cofounder spent two weeks manually rewriting the codebase in VS Code to restore maintainability. (from why software factories fail)
- Reframes the real problem: it's not 'too many PRs,' it's too many bad PRs — a PR needing even 20% rework imposes real intellectual/emotional burden on reviewers, and AI oneshot PRs often trend closer to 50% needing rework. (from why software factories fail)
structured-agentic-workflow
- Proposed workflow to keep human-level code quality while moving 2-3x faster: four front-loaded phases before implementation — Product Review (problem/success criteria + HTML mockups), System Architecture (sequence diagrams, contract shapes, schemas), Program Design (call-stack diffs, file-tree diffs, type signatures), and Vertical Slices (build middle-out: API contract→mock→frontend→services→DB, testing at each step) rather than models' preferred horizontal/stack-order plans. (from why software factories fail)
- Distribution guideline for applying the 4-phase process: ~40% of tasks get oneshot or oneshot+light feedback; medium tasks combine product/system design into one doc without phase breakdown; only large/complex tasks get the full four-phase treatment (skipping product review for big refactors). (from why software factories fail)
planning-through-doing
- For visual/UI tasks, iterate on wireframes in .html files instead of markdown specs — agents can zero-shot complex interactions from wireframes, and can screenshot/interact with them via a dev-browser CLI or pull details from the dev console for extra feedback. (from planning through doing vs markdown)
- Moving a wireframe into the actual product is often a zero-shot task for the agent, depending on scope — visual artifacts transfer to implementation better than text specs. (from planning through doing vs markdown)
- For backend logic/API tasks, iterate using separate runnable scripts rather than markdown docs — execution logs (formatted, colored, or reported interactively) give far more feedback than prose, enable faster/deterministic comparison of iterations, and can be comment-rich to carry codebase context. (from planning through doing vs markdown)
- Agents rarely zero-shot full integration from a test script on the first try, but executable input still beats a wall of markdown text for quality of results. (from planning through doing vs markdown)
- For the app itself, prefer 'planning through doing': generate and regenerate code directly in the app rather than planning on paper first. Paper planning made sense when code was hand-written or agents needed many iterations to get syntax right — now LLMs often get it right the first time, so pull context (e.g., git status) only afterward if needed. (from planning through doing vs markdown)
- If cross-session text-based context is genuinely needed, use .mdsvx/.mdx-style formats rather than plain markdown — they support visual feedback and/or interactivity that improves reasoning about the files. (from planning through doing vs markdown)
reusable-primitives
- Claim: ~99% of software can be built from the same finite set of reusable elements/components rather than bespoke code each time. (from finite elements software reuse)
- Rebuilding common software elements from scratch repeatedly via AI/tokens each project is framed as wasteful — described as 'peak AI psychosis' — implying teams should invest in a reusable component library instead. (from finite elements software reuse)
agent-prompting
- A strong first prompt to a coding agent has five components: state engineering judgment/traps explicitly, paste raw errors/logs (not summaries), point to exact files/PRs, specify mode (research vs implement), and tell the story first then ground it in code context. (from how i work with coding agents)
- Check agent direction early (which approach, which file opened first) rather than reviewing the finished diff — passing tests can validate a wrong solution, and 90 early interrupts in a month were cheaper than discarding finished wrong work. (from how i work with coding agents)
- For large features, split work into three sequential modes: a no-code planning conversation for architecture/tradeoffs, then a spec with precise tasks, then parallel-wave execution — agents excel at executing precise specs but are mediocre at inferring vague intent. (from how i work with coding agents)
agent-governance
- Rule: agents never commit code. The human reads every diff and writes/batches commits themselves — this is framed as essential for retaining a mental map of the codebase, which in turn enables good prompts and early direction-checking. (from how i work with coding agents)
- Any mistake an agent makes twice should be converted into a written rule in the project's instructions file (e.g. 'never commit unless asked', 'branch names are fix/ or feat/ only', 'use pnpm, never npm') — described as the single highest-leverage artifact in the repo. (from how i work with coding agents)
context-management
- Clear agent context frequently rather than letting long conversations accumulate stale assumptions and outdated file state; a fresh agent with a strong first message plus instructions file consistently outperforms a tired agent with long memory. (from how i work with coding agents)
llm-writing-tics
- AI coding assistants (Claude, Codex-style agents) exhibit a recognizable set of stock metaphors in code commentary: 'load-bearing X', 'X in a trenchcoat', 'the codebase wants to', 'held together with hope/convention', 'archaeological layers/scar tissue' — useful as a heuristic for spotting AI-generated PR descriptions or comments. (from llm clanker phrases)
- A common LLM rhetorical pattern for hedged distinctions: 'X isn't just Y anymore', 'less about X, more about Y', 'not X so much as Y', 'X, but for Y' — these formulaic contrast structures appear disproportionately in AI-written explanations and summaries. (from llm clanker phrases)
- 'Actually let me' / 'Let me write this' repeated mid-generation is a distinctive self-correction tic seen in AI coding agent transcripts, signaling real-time plan revision during generation rather than upfront planning. (from llm clanker phrases)
model-overengineering
- GPT-5.6 Sol-class models exhibit a distinct failure mode: not incorrect code, but overbuilt code. A requested config change can produce a full framework; a bug fix can produce an unnecessary adapter layer. (from stop gpt5 sol overengineering)
- Overengineered AI-generated code often passes tests and CI cleanly, hiding the problem until someone tries to modify the logic later and finds multiple unnecessary abstraction layers in the way. (from stop gpt5 sol overengineering)
- The post promises a concrete fix for curbing GPT-5.6 Sol's overengineering tendency, but the fix itself was not captured in this excerpt—only the problem framing. (from stop gpt5 sol overengineering)
agent PR tooling
- Avoiding committed media in PRs (screenshots/videos) keeps repo size clean while still giving reviewers visual before/after diffs, addressing a common pain point in agent-generated pull requests. (from github user attachments pr media)
Voices
66 contributors
Nick
@nickbaumann_
codex @openAI | prev @cline | product of @UWMadison 🦡
Miles Deutscher
@milesdeutscher
Obsessed with AI. Tweets aren’t financial advice. Sharing early alpha in @mileshighclub_. Building @aiedge_.
Theo - t3.gg
@theo
Francois Laberge ✍️
@seflless
Founder Decode. Building the whiteboard for you & your coding agents.
brett goldstein
@thatguybg
founder/designer @microHQ | investor @launchhouse | ex-Google cognitive scientist | sprezzatura
Siqi Chen
@blader
🏗️ Love to build (@runwayco @sandboxvr @zynga) people love 💸 Investor @amplitude_hq @mercury @owner @elevenlabsio @meetgamma @sfcompute @turingcom++
George from 🕹prodmgmt.world
@nurijanian
Can I make everyone a great product manager? I will do my best | Get my product management OS + AI skills for Claude Code/Cursor: https://t.co/ngCnvp77SD
Thariq
@trq212
Claude Code @anthropicai. prev YC W20, mit media lab. towards machines of loving grace
Alex Finn
@AlexFinn
Founder/CEO of Henry Intelligent Machines PBC and Creator Buddy. Building a 100 trillion dollar economic engine
Vox
@Voxyz_ai
Andrej Karpathy
@karpathy
I like to train large deep neural nets. Previously Director of AI @ Tesla, founding team @ OpenAI, PhD @ Stanford.
Matt Van Horn
@mvanhorn
Co-founded June ("self-driving oven" acquired by @webergrills) & the co that became @Lyft. Building again, more soon. Vibe coding @slashlast30days research tool
OpenAI Developers
@OpenAIDevs
Official updates for developers building with Codex & the OpenAI Platform • Service status: https://t.co/kZwnwdYYEq
Peter Yang
@petergyang
Practical AI tutorials and interviews for busy people | Join 140K+ readers at https://t.co/XYKTmGVH14 | Product at Roblox
klöss
@kloss_xyz
AI Educator, Designer & Developer | @psychanon CEO Building AI-powered brands, workflows, and apps.
Josh Pigford
@Shpigford
✨ dabbler 🪀 https://t.co/BjH6uyEPsY 💬 https://t.co/KkRZyGkkMW 🚀 https://t.co/Zess2MJy7R 🤫 https://t.co/yCg2xUTyoz ◀️ https://t.co/m98SW4tKrF
Todd Saunders
@toddsaunders
CEO of @Broadlume, vertical SaaS for 4,000+ flooring retailers. Acquired 8 companies before selling to @Cynclyco. Previously @google. Long @townofwestfield.
Zara Zhang
@zarazhangrui
Builder. Dangerously skips permissions. Harvard’17. GitHub: https://t.co/KCuEajezlL YouTube: https://t.co/8xzbGWtf6w
Jonata Santos
@_jonatasantos
Building a portfolio of products at https://t.co/mFAVIoxTzS 🚀
Okara
@askOkara
ai lab building products for consumers: AI CMO, Private AI Chat
Ben Tossell
@bentossell
can't code, won't code. builder, investor. 3 under 3 👶. only invest in dev tools & infra
dex
@dexhorthy
GREG ISENBERG
@gregisenberg
I drop startup ideas daily. Host @startupideaspod. CEO: @latecheckoutplz we build companies like @ideabrowser, @meetLCA, @boringmarketer etc
Guinness Chen
@guinnesschen
Building codex at @openai, prev @stanford, @imbue_ai
Jason Zook
@jasondoesstuff
IWearYourShirt, BuyMyLastName, SponsorMyBook, BuyMyFuture, Teachery, Wandering Aimfully, now vibe coding to $100k in 2026! 🟩 🟩 ⬜️ ⬜️ ⬜️ ⬜️ ⬜️ ⬜️ ⬜️ ⬜️ (20%)
Michael Guo
@Michaelzsguo
Building AI agents and AI-native orgs. Demystifying AI in practice. EN/中文
Nabeel Hyatt
@nabeel
vc @sparkcapital 🐶 partner to @thebotcompany @discord @descriptapp @meetgranola @instawork @ridezum, prior bod @cruise @postmates Q 🍉 side-gig @BerkeleyTTL
Nathan Baschez
@nbaschez
Founder of @lexdotpage. AI scout at @trueventures. Previously: co-founder @Every, first employee @SubstackInc. Always: @SoniaBaschez
Nick Spisak
@NickSpisak_
| AI Transformation Engineer | Seven Figure E-Commerce Business Owner
Nico Bailon
@nicopreme
Building open source agentic tools, mostly for Pi coding agent.
nini
@nini_incrypto_
07大二在读|crypto➕ai Web3 创作者|@CHAINISLE_ZH 生态成员 网红孵化机构/商务合作 vx+ninio2o(付费社群/付费增长咨询/ip孵化陪跑)
shadcn
@shadcn
I own a computer / Working on https://t.co/HJcOr0AUAr & https://t.co/5FRvxukoY5 / @vercel.
Simon Last
@simonlast
Building @NotionHQ
Nick Khami
@skeptrune
currently doing things at Mintlify, prev. built a search API (trieve acq. YCW24), you should try to fail faster
Suryansh Tiwari
@Suryanshti777
Exploring AI & SaaS trends early Sharing what’s actually useful Helping builders turn ideas → products → traction – 📩 Open to collabs
tobi lutke
@tobi
Shopify CEO by day, Dad in evening, hacker at night, Aspiring comprehensivist. + qmd !
Vasu-Devs
@Vasu_Devs
AI Engineer • 21 • Open to work! Building: https://t.co/f64U8MDKtG
Adam
@_overment
agrim singh
@agrimsingh
tinkering @ https://t.co/DlHMufmuQ1 // https://t.co/tpc1Nwqrix // https://t.co/2mTTsnMgAY // ambassador @openai codex, @cursor_ai, @v0 // whisky guru & dj
AI Builder Club
@aibuilderclub_
Aniket Panjwani
@aniketapanjwani
I teach agentic coding to economists || PhD Economics Northwestern || Director of AI/ML @ Payslice || ex-MLOps @ Zelle
Balint Orosz
@balintorosz
Founder @craftdocs
Daax
@daaximus
Khairallah AL-Awady
@eng_khairallah1
James Bedford
@jameesy
engineering @zerion, building https://t.co/3RQYuZAcrt
Jamon
@jamonholmgren
Jason D
@JasonUXUI
Design Engineer for Startup & B2B Products ✨ | Let’s talk https://t.co/dKQZuCUncv
Justin Rands
@jayrizpop
brand engineer @clay
˗ˏˋ Jesse Hanley ˎˊ˗
@jessethanley
Marketer, self-taught developer, and founder of @Bento and https://t.co/lcsIohchEv. Designing a quiet family life in 福岡, Japan. DMs open if you need email help 🌿
Shahaf Antwarg
@kotevcode
Manthan Gupta
@manthanguptaa
ai research engineer • designing agent runtimes, memory & retrieval systems for autonomous agents • dms open
Matthew Berman
@MatthewBerman
caiden
@pipelineabuser
https://t.co/vomYheloEx - AI B2B growth systems https://t.co/kIJQjFhaiv - premium cold email inboxes https://t.co/j8g8jXTfL0 - private newsletter
Prasann Pandya
@prasann_pandya
In pursuit of building great products | Built @myreader_ai (100k+ users), @colive_planet | Nomad 🎒 | Yogi 🧘 | Reader 📚 | Minimalist | Very outspoken
Quentin Romero Lauro
@Qromerolauro
Co-Founder & CEO at Inspector (YC F25); previously @character_ai @Berkeley_EECS @cmuhcii
rahul
@rahulgs
head of applied ai @ ramp
Rexan Wong
@rexan_wong
✞ 18 / building b2b software / prev. content apps (500K+ users)
Sam Bernhardt
@samuelbernhardt
dabbling in design & engineering. prev @tokenterminal
Samuel Stroschein
@samuelstroschei
Savio
@saviomartin7
18 — Full stack web & IOS Developer. Building https://t.co/A6FuaY3m9h — Deploy your OpenClaw instance under 1 min, built for normal people.
Sawyer Hood
@sawyerhood
Software Engineer that is building weird shit in public, latest projects: https://t.co/6dQdN8KliO https://t.co/rQyv2zpSkj. Prev: @figma, @facebook.
Shiv
@shivsakhuja
Pontificating... / Vibe GTM-ing / Making Claude Code do non-coding things building a team of AI coworkers @ Gooseworks / prev @AthinaAI /@google / @ycombinator
Patrick Coddou
@soundslikecanoe
Teaching machines to work so humans can create. Operations / AI at AJF Growth (Meta Agency) Previously founder at @getsupply (acquired) Rarely serious.
Tom Osman 🐦⬛
@tomosman
Trevin Chow
@trevin
Vincent van der Meulen
@vinvan