u/mangoavococo

How do you set up evals when you want them to run against real dependencies?

Perhaps more of a noob question, but what's a smart way for me to set up evals when I want them to run against dependencies that come up in real app scenarios, like feature flags, real traffic, diff services? How do you test agents that call multiple real tools/APIs? I can't have an eval run issuing 40 actual refunds and printing 60 return labels.

reddit.com
u/mangoavococo — 6 hours ago