A $63 Mistake in Production: How an AI Agent's Loop Escaped Without Guards

An automated trigger fired every two minutes for 8 hours. 243 agent runs. 31M tokens. $63 wasted. The problem: nobody noticed until morning.

This is a case study in production safety when AI agents are on the loop — and a cautionary tale about human approval gates.

The technical breakdown:

  1. An Ash/Oban cron trigger polling GitHub every 2 minutes — seemed fine on review
  2. The check was level-triggered (state-based) instead of event-triggered
  3. GitHub review feedback is immutable history, not a dismissable flag
  4. So the agent kept trying to fix the same feedback, over and over
  5. No cost counter, no alert, no visibility — just a morning invoice

Why the existing guards failed:

  • Freshness check existed but lived on an unreachable code path
  • Attempt counter got reset to 0 on every pass (accidentally disarmed)
  • Branch order made recovery impossible once one reviewer approved

The real lesson: when you put an expensive actuator (LLM tokens) on a loop, “harmless redundancy” stops being a category. Every pass has to earn itself.

Full postmortem:

Built with Elixir/Phoenix/Ash, and the hardest part wasn’t the code — it was the code review.