Discussion about this post

User's avatar
TOM@Future3labs's avatar

The harness-leverage result suggests a useful procurement lesson: benchmark the model and orchestration layer as a pair. A model that looks expensive in isolation may win on completed-task cost if it needs fewer retries, produces shorter tool traces, or benefits more from caching and decomposition. I would also distinguish token reduction from economic efficiency—external tool latency, human approvals, and failure recovery can dominate the bill. Have the authors published per-task variance? The average improvement is strong, but the tasks where harness design hurts quality may be the most operationally important.

3 more comments...

No posts

Ready for more?