· 5 min read
Why agentic coding, not vibe coding
by Furkan Durna
#ai #workflow #agents
Vibe coding and agentic coding look the same from the outside: you describe what you want and an AI writes the code. The difference is what happens next. In one, you accept the result and move on. In the other, the result has to prove itself before it counts. For anything that other people will use, the second is the only sensible choice.
Two terms, one line between them
Andrej Karpathy coined "vibe coding" in a post on February 2, 2025: coding where you "fully give in to the vibes, embrace exponentials, and forget that the code even exists." He accepts every change without reading the diffs, pastes errors back in without comment, and when a bug won't go away he asks "for random changes until it goes away." He also named its limit: "It's not too bad for throwaway weekend projects."
Simon Willison drew the line a few weeks later: vibe coding is "building software with an LLM without reviewing the code it writes." If you review it, test it and can explain how it works, "that's not vibe coding, it's software development."
Agentic coding sits on the other side of that line. The agent reads the codebase, plans changes across files, runs commands and tests, and iterates. You set the constraints, make the decisions, check the evidence and own the result.
Why it matters
The code works, but it isn't safe
Veracode tested more than 100 language models on 80 coding tasks for its 2025 GenAI Code Security Report. In 45% of the tasks, the generated code introduced a known security flaw from the OWASP Top 10. For cross-site scripting, the models failed to defend against it in 86% of relevant cases. Newer and larger models were not more secure; they only got better at producing code that runs.
Code that runs is exactly what vibe coding checks for. A security flaw doesn't throw an error you can paste back in. If nobody reads the code, nobody finds it until someone else does.
Instructions are not guardrails
In July 2025, Jason Lemkin was nine days into a vibe coding project on Replit when its agent deleted his production database, with records for more than 1,200 executives and 1,190 companies. It did so during a code freeze, after eleven separate warnings in all caps not to make changes. It then told him recovery was impossible, which turned out to be false. Replit's CEO called it "unacceptable" and shipped automatic separation of development and production databases.
The lesson isn't that agents are malicious. It's that telling a model what not to do is a request, not a control. What actually prevents damage is structure: an agent that can't reach production, changes that go through review, and backups you have tested.
Feeling fast is not evidence
In a randomized trial by METR, experienced open-source developers took 19% longer on real tasks when they could use early-2025 AI tools. Before starting, they expected to be 24% faster; afterwards, they still believed they had been 20% faster. METR's 2026 follow-up says developers are likely sped up by newer tools, while calling its own data weak evidence of how much.
The number that matters here is the gap between perception and measurement. If you can't trust your sense of how fast you're going, you can't trust your sense of whether the code is right either. Tests, type checks and screenshots don't have that problem.
Code nobody understands is debt
Karpathy said it himself: "The code grows beyond my usual comprehension." For a weekend project, that's fine. For software that has to keep running, it means every future bug is a bug in code nobody on the team can read. "Ask for random changes until it goes away" isn't a debugging strategy you want behind a payment flow.
Someone still owns the result
Users don't care which model wrote the bug that leaked their data. The person who shipped it is responsible, and they can only take that responsibility for code they have checked.
What agentic coding looks like in practice
This site was built with coding agents, so here is what that means concretely:
- Write the constraints down. The repository has an
AGENTS.mdtelling agents that this Next.js version differs from their training data and that they must read the bundled docs first. Project rules cover things like keeping the plain-text guide to the site in sync with the apps. Content rules say there is no email address on the site and no claim the code contradicts. - Limit what the agent can touch. Secrets stay in environment files that are never committed. I commit and push myself.
- Make every change prove itself. Typecheck, lint, and for anything visual a browser script that drives the page and takes screenshots. When I added a countdown to Tetris, the first test run reported a restart that shouldn't happen. The screenshot showed the test was at fault: it kept pressing Space after game over, which starts a new game. Without the screenshot, working code would have been "fixed".
- Check claims against the code. The first draft of the site guide said the Settings app has a language option. It doesn't. The Finger of Hope's project page says its encryption layer was never connected, because that's what the code shows.
- Verify what the agent tells you. While I was removing a name from an old repository's history, an agent said the deleted commits could only be found by someone who already knew their IDs. GitHub's public events feed listed them. The claim sounded reasonable; it was caught because it was checked.
When vibe coding is fine
Karpathy and Willison agree on this, and so do I: for a throwaway prototype, a one-off script or an experiment to see what a model can do, let it rip. Low stakes, no user data, no money on the line, nothing anyone else depends on.
The moment someone else relies on the code, the loop has to close.
The short version
Let the agent write the code. Don't let it be the last one to look at it. Speed comes from the agent; confidence has to come from checks, and responsibility stays with whoever commits.