What do you consider a fair comparison between two AI agents?
Suppose two agents are given the exact same task.
Agent A finishes in 20 seconds with 15 tool calls.
Agent B finishes in 45 seconds with 5 tool calls.
Which one would you consider "better"? I'm curious how people actually balance performance, reliability when evaluating agents.