A 21-day agent run migrated a key legacy module to TypeScript and React for under $500 in model tokens, while exposing the visual checks needed before production.
Read in full here:
A 21-day agent run migrated a key legacy module to TypeScript and React for under $500 in model tokens, while exposing the visual checks needed before production.
Read in full here:
The interesting part here isn’t just the <$500 figure; it’s that the 21-day run surfaced the visual checks needed before production. That makes evaluation look more like a pipeline question: how much context can the agent retain between migration steps, and which observations are worth carrying forward? In my own experiments I’ve been looking at Contextpress (GitHub - Taha-azizi/contextpress: Deterministic context compression for LLM chat, RAG, and agent pipelines · GitHub · contextpress · PyPI), a deterministic local compressor for message lists, as a way to drop duplicate tool output while preserving decisions and constraints. I’d want to compare migrations by replaying with a fixed budget and tracking review findings, retries, and regressions—not token spend alone. The visual checks seem especially valuable because they expose failures that a clean typecheck can miss.
The <$500 headline is only meaningful once you know whether those 21 days were one long cached session or many cold starts. A cache miss re-writes history at about 2× base input instead of reading it at 0.1× — a 20× swing on the prefix — so the same migration can land closer to $200 or $800 depending on TTL gaps and how often the agent compacted. Parallelizing steps helps only if you don’t copy the growing context into each worker; otherwise N workers can cost more than one patient agent. I wrote up that parallel-agent math with current Anthropic multipliers here: Claude Code agent teams and subagents: the real token cost · AI//COST (disclosure: I work on that site).