An Empirical Study of Harness Design for Coding Agents

The result that deterministic elision before summarization gives the best efficiency matches something I keep hitting in practice: token reduction and dollar reduction aren’t always the same thing when cached input is discounted. I’d want the next experiment to log rendered prompt size, cached vs uncached input, and task-quality retention side by side when comparing a deterministic compressor to a summarizer. For models like Jev with input-only pricing, a deterministic pre-call transform makes the savings much easier to attribute.