{{POST_BODY}} Replace this with the full post.
An agent that fails loudly is a good day. The expensive failures are the quiet ones: a confidently wrong source, a loop that costs money, a slow degradation after an index refresh.
Scoring per stage instead of per answer is what makes these debuggable, because it tells you which part of the system moved.