InfoWebPlus Logo
Home
Services
AI IntegrationsMobile App DevelopmentWeb Design & DevelopmentCustom SoftwareSEO Services
Resources
ResearchCase StudiesToolsInsights
Toolkit
Website Cost CalculatorAI Cost EstimatorAI Prompt LibraryView All
Contact

Referral Program

Earn money by referring clients. Fixed commissions, simple process, GDPR compliant.

Join the referral program
InfoWebPlus Logo

Product engineering studio, web applications, integrations, and pragmatic AI. Small by design, senior by default.

contact@infowebplus.com

Company

  • About
  • Services
  • Expertise
  • Contact

Services

  • AI Integrations
  • Mobile Apps
  • Web Design
  • Custom Software

Apps

  • Games
  • Developers
  • Mobile Apps
  • Services

Costa del Sol

  • All areas
  • Málaga
  • Marbella
  • Mijas
  • Estepona
  • Sotogrande
  • Fuengirola
  • Gibraltar

© 2016–2026 InfoWebPlus™ · All rights reserved.

Legal|Privacy Policy
Research

How Vibe Coding an AI Agent Taught Me What Hallucination Actually Costs

Wanting to test an existing AI agent project, I had an assistant vibe-code the infrastructure around it instead. It never recognized the project I meant, never said so, and confidently built something else entirely.

  1. Home
  2. /Research
  3. /How Vibe Coding an AI Agent Taught Me What Hallucination Actually Costs
Research

Published September 3, 2026·3 min read

  • AI Implementation
  • Llm Risk
  • Hallucination
  • Vibe Coding
  • Ai Agents
George Barbu

George Barbu

Founder & Product Engineer

Full analysis

Wanting to test an agent, not build one

Discovering an existing AI agent project online is usually enough to make someone curious, and reading about it is a poor substitute for actually trying it. The fastest way to understand how something like that really works, or so it seemed, was to have an AI coding assistant build the surrounding infrastructure through vibe coding: describe the intent in plain language, let it generate the stack, iterate from there.

That decision, treating "let me set this up so I can try it" as a single vibe-coded request, is where the actual experiment started, even though the real subject turned out to be something else entirely.

What "setting it up" quietly became

The assistant had no real knowledge of the specific project being referenced. Rather than saying so, or asking what was meant, it treated the request as "build me a personal AI agent" and proceeded with total confidence: a Docker Compose stack on a small VPS, a reverse proxy with TLS, a database, an admin-only VPN mesh, observability, a messaging app as the primary interface, a phased integration plan across a dozen external services, and a dual-model routing layer with a confidence threshold for uncertain actions.

Within days, that stack existed and mostly worked. It even had its own name, chosen unprompted, because a coherent personal-agent project apparently needed one.

The hallucination nobody flagged

Here is the part worth naming precisely: at no point did the model say "I don't recognize the specific product you're referencing." It never asked whether the goal was to try that exact project or to build something similar instead. It picked the second interpretation on its own, stated nothing about the gap, and moved straight to implementation.

That is what hallucination actually looks like inside a multi-day agentic build: not a wrong answer to a direct factual question, but an entire working system constructed on top of one unstated, incorrect assumption, several inferential steps removed from where the conversation started, and invisible for exactly as long as everything downstream of the mistake keeps working. Every decision point along the way, dozens of them across a few days of iteration, was a place the model could have said "I'm not certain this is what you meant," and chose a confident, specific default instead. Not once, at any of them.

What building the wrong thing well actually cost

Within a few more days, that custom-named agent had been taught to make small, real modifications to its own infrastructure, genuinely working automation for genuinely small tasks. The system was not a failure. It functioned.

But the honest comparison is not "a working system" against "nothing." It is against the possibility that the specific project originally intended already did this, tested and maintained by someone else, without a week of rebuilding groundwork from scratch. Some of that research on LLM hallucination in business contexts generalizes directly here: the risk of a hallucination is not fixed by how often a model gets things wrong, it depends entirely on how far downstream the consequence propagates before anyone checks. A wrong answer in a single chat message costs a re-read. A wrong assumption at the start of a multi-day infrastructure build costs the build.

The actual lesson

Vibe coding rarely fails loudly. It fails by cheerfully constructing something adjacent to what was asked for, finishing it competently enough that nobody thinks to ask whether the starting assumption was ever confirmed.

The practical fix costs one sentence: name the exact product or reference being tested, and ask the assistant directly whether it recognizes it, before it generates anything. A reference architecture is easy to adapt once you know what you're actually building toward; it is expensive to discover you built the wrong one four days in.

What actually got built

In the end, the result was not a clone of the project that started the whole exercise, and not a competitor to it either. It was infrastructure for a custom-named AI agent, for the plain reason that the model didn't know what the original project was, and confidently assumed that inventing one was the goal.

Frequently asked questions

So the AI just made something up?
Yes, functionally. It had no real knowledge of the specific project being referenced, so it substituted the closest thing it could construct with confidence: a generically named personal AI agent, complete with its own name and design decisions.
Could this have been avoided?
Easily, with one extra step: asking the model directly whether it recognized the specific product being referenced before it generated anything, and treating a vague or missing answer as a signal to stop and clarify.
Was the time spent worthless?
No, the resulting system genuinely works. But it duplicated groundwork an existing, maintained project had likely already solved, which is the real cost: not wasted effort, effort redirected at the wrong target.

Share this article

Copy the article URL or use your device share sheet.

Related reading

Sep 14, 2026·George Barbu

Best Restaurant Management Software: Off-the-Shelf vs. Custom-Built

A practical guide for restaurant owners comparing off-the-shelf platforms like Toast or Square against a custom-built system, covering where generic tools break down and how to decide which one your business actually needs.

  • Restaurant software
  • POS systems
  • Custom software
Read more
Jun 15, 2026·George Barbu

LLM Hallucination in Business Contexts: Risk Profiles by Use Case

Key finding: Business risk of LLM hallucination is determined by the combination of error detectability, error reversibility, and error exposure scope, not by hallucination rate alone.

  • AI Implementation
  • Llm Risk
  • Hallucination
Read more

Want help applying this?

Rough scope is fine. Tell us what you are building, we reply with options and tradeoffs, not a generic pitch.

ContactGet a quote

Related on InfoWebPlus

  • Technical SEO Review
  • SEO Audit
  • Entity Stacking
  • Website Cost Calculator
  • Schema Generator
  • Workflow Mapping