Field guide / primary domain

Vibe Coding

71

sources in this field

Updated August 30, 2026

Current thesis

The shortest path to orientation.

This synthesis records claims and practices from the cited sources; reported outcomes and product capabilities have not been independently verified.

Vibe coding has matured into production-scale software where frontier models handle complex tasks autonomously through supervised orchestration. Ultracode mode in Opus 4.8 removes manual intervention by enabling Claude to invoke workflows independently; supervisory workflows delegate subgoals, route routine execution to cheaper models, and enforce quality gates. Two-model adversarial loops—one drafting, one reviewing—prove effective; GPT-5.5 consistently finds issues in both planning and code review. At ~$400/month for Opus 4.7 + GPT-5.5, end-to-end feature work costs equivalent to fractional dev teams. Design specs via DESIGN.md achieve 95%+ principal completion rates. Strong prompts engineer state traps explicitly, paste raw errors, and specify mode; written rules in instructions files have highest leverage. For large features, split into planning, specification, then parallel execution. GPT-5.6 Sol-class models exhibit a distinct failure mode: not incorrect code, but overbuilt code where a config change produces full frameworks and bug fixes produce unnecessary adapter layers. Overengineered AI code often passes tests cleanly, hiding problems until modification attempts. Avoiding committed media in PRs keeps repo size clean while preserving reviewer visibility. Humans read every diff before commit.

Evidence board

Claims worth carrying forward
01

GPT-5.6 Sol-class models exhibit a distinct failure mode: not incorrect code, but overbuilt code. A requested config change can produce a full framework; a bug fix can produce an unnecessary adapter layer.

02

Overengineered AI-generated code often passes tests and CI cleanly, hiding the problem until someone tries to modify the logic later and finds multiple unnecessary abstraction layers in the way.

03

The post promises a concrete fix for curbing GPT-5.6 Sol's overengineering tendency, but the fix itself was not captured in this excerpt—only the problem framing.

04

Coding agents can add before/after screenshots or videos to GitHub PR descriptions by POSTing media to GitHub's user-attachments upload endpoint, then embedding the returned URL directly in the PR body—no need to commit binary files or host media externally.

05

Avoiding committed media in PRs (screenshots/videos) keeps repo size clean while still giving reviewers visual before/after diffs, addressing a common pain point in agent-generated pull requests.

06

AI coding assistants (Claude, Codex-style agents) exhibit a recognizable set of stock metaphors in code commentary: 'load-bearing X', 'X in a trenchcoat', 'the codebase wants to', 'held together with hope/convention', 'archaeological layers/scar tissue' — useful as a heuristic for spotting AI-generated PR descriptions or comments.

07

A common LLM rhetorical pattern for hedged distinctions: 'X isn't just Y anymore', 'less about X, more about Y', 'not X so much as Y', 'X, but for Y' — these formulaic contrast structures appear disproportionately in AI-written explanations and summaries.

08

'Actually let me' / 'Let me write this' repeated mid-generation is a distinctive self-correction tic seen in AI coding agent transcripts, signaling real-time plan revision during generation rather than upfront planning.

Adjacent fields

Key voices

Latest evidence

Recent additions

All synthesized insights →

AI Builder Club

GPT-5.6 Sol Models Overengineer Code Even When Correct

A post identifies a specific failure mode in 'Sol-class' GPT-5.6 models: they don't produce wrong code, they produce excessive code—turning simple config changes into frameworks and bug fixes into adapter layers. Tests pass and CI stays green, masking accumulating complexity that surfaces later when someone tries to modify the logic. The post promises a fix but the solution content itself is not included in this excerpt.

Vibe Coding Needs context

Shahaf Antwarg

How I Work With Coding Agents

A workflow synthesis from a developer running 20 parallel coding agent sessions: role shift to tech-lead, isolated worktrees per task, five-part first-message discipline, early direction-checking over end-diff review, manual commits, and a living instructions file.

dex

Why Software Factories Fail: Harness Engineering Is Not Enough

Dex (HumanLayer) argues the 'lights-off software factory' (no human reads or writes code) fails because models cannot maintain codebase quality over time — a model-training/RL limitation, not a skill or harness-engineering issue. He proposes a four-phase human-in-the-loop process (product review, system architecture, program design, vertical slices) to move 2-3x faster safely.

Adam

Why Planning in .md Is a Waste of Time

Adam argues markdown-based planning is largely obsolete: for UI/UX use wireframes agents can iterate on visually, for backend logic use runnable/testable scripts instead of prose, and for the app itself plan by directly generating and regenerating code, since LLMs now often get syntax right on the first try.

Vasu-Devs

JustHireMe: Local-First Agentic AI Job Intelligence Workbench

JustHireMe is an open-source, AGPL-licensed desktop app (Tauri + React + Python/FastAPI) that addresses broken job search UX with a local-first, explainable AI pipeline: scrape → quality gate → rank → match (GraphRAG + vector) → generate tailored application materials. All data stays on-device by default, LLM usage is keyless-capable via Ollama/Claude Code CLI, and the architecture prioritizes human-in-the-loop control over blind automation.

GREG ISENBERG

Design.md + AI Skills: Consistent Startup Branding in One Hour

Google's open-source Design.md format captures typography, colors, and spacing in a single markdown file that agents reference to produce consistent outputs. Combined with reusable skill files (landing page, mobile, pitch deck), it creates a design system any AI agent can apply uniformly across all surfaces—replacing the common pattern of polishing one screen while everything else looks generic.