In the ARC-AGI-3 benchmark, using a task-engineered harness is considered "cheating" because it shifts the evaluation from testing the AI model's native intelligence to measuring the human engineer's scaffolding.
> In the ARC-AGI-3 benchmark, using a task-engineered harness is considered "cheating" because it shifts the evaluation from testing the AI model's native intelligence to measuring the human engineer's scaffolding.
That's the claim. I'm not getting any kind of clarity on this issue as much as it seems it might be clear. People are debating this.
Your thoughts?
systems acing the entire suite,