If someone asked you to prove a human has been supervising your automated system, what would you actually send?

Not logs I have logs. I mean something a person outside the team could read and come away convinced the supervision was real rather than theoretical.

Has anyone actually been asked for this, by an auditor, a customer, or your own legal team? What did you send, and did it hold up?

reddit.com
u/JuniorLeg6988 — 3 days ago

If someone asked you to prove a human has been supervising your automated system, what would you actually send?

Not logs I have logs. I mean something a person outside the team could read and come away convinced the supervision was real rather than theoretical.

Has anyone actually been asked for this, by an auditor, a customer, or your own legal team? What did you send, and did it hold up?

reddit.com
u/JuniorLeg6988 — 3 days ago

Approval logs can contain "approved" rows where no human was involved

I was reading through the docs for a human-approval service and noticed a few ways a decision ends up logged as approved without a person seeing it. A dev/test flag that auto-approves, server-side rules that resolve below a threshold, and timeouts that fall through to a default.

All legitimate features. But nothing in the log distinguishes those rows from a real human decision, so "we had 400 approvals last quarter" doesn't mean 400 people looked at something.

If you have an approval step on an automated system. Have you ever checked how many of yours actually reached a human? Curious whether people track this or whether it's the kind of thing nobody looks at until someone asks.

reddit.com
u/JuniorLeg6988 — 3 days ago

For those who don't let ai touch prod or money has this line ever moved? (I will not promote)

I keep seeing people say they’re comfortable using AI for a lot of things, but they still want a human involved anytime money is being spent or production code is being shipped.

I’m curious about people who used to feel that way but have since loosened the reins a bit.

Has anything in your workflow gone from “a human has to approve every single one of these” to “this is reliable enough that we just let it run”?

If so, what actually changed your mind?

Was it just seeing hundreds or thousands of clean runs? Better logging/auditing? Putting tighter limits around what the AI could do? Or was there some specific moment where you realized the human approval step wasn’t really adding much anymore?

And on the other side, are there things that have never crossed that line for you and probably never will?

I’m also curious if there’s anything currently stuck behind human approval that you genuinely wish you didn’t have to review anymore. What is it, and what would need to change before you’d be comfortable automating it?

reddit.com
u/JuniorLeg6988 — 3 days ago

For those who don't let agents touch prod or money — has that line ever moved?

Lots of people say they keep a human on anything that spends money or writes to prod. Curious about the ones where that changed.

Has anything ever graduated from "human approves every time" to "it just runs"? What made you comfortable — a volume of clean runs, an audit, a specific incident that didn't happen? Or has nothing ever crossed, and the line is where it is permanently?

And if something is stuck on the human-approval side that you'd genuinely rather not be reviewing, what is it?

reddit.com
u/JuniorLeg6988 — 3 days ago

Where do you still keep a person in the loop on customer decisions, even when a tool could technically make the call?

Approving a refund, waiving a fee, closing a complaint in the customer's favor, whatever counts as "the call" wherever you work.

There's a difference between a tool that drafts something for a person to send, and one that could actually just do it, no one clicking "approve" first. I'm asking about the second kind.

If your team has anything like that: what's the thing you still won't let it do alone? Did something specific happen, or nearly happen, that's the reason, or is it more that even a correct call still feels too risky to let go through untouched?

It's a bit like how a self-driving car can handle almost everything but someone still needs to be ready to grab the wheel. Trying to find that same moment for customer service decisions. What would have to be true before the easy, obvious ones went through without a check?

Not selling anything, genuinely just trying to understand where people draw that line and why. Would love to hear from anyone who deals with this, whether you're the one approving it or the one cleaning up after.

Thanks!

reddit.com
u/JuniorLeg6988 — 4 days ago

What customer-facing decision do you still make a person approve, even when a tool could technically do it on its own?

Not just refunds, could be a credit, a waived downgrade, a retention offer, closing out an escalation in the customer's favor, whatever counts as "the call" on your team.

There's a difference between a tool that drafts something for a person to send, and one that could actually just do it, apply the credit, close the ticket, issue the resolution, no one clicking "approve" first. I'm asking about the second kind specifically.

If your team has anything like that, what's the one thing you still won't let it do without someone checking first? Was there a specific time something went wrong, or nearly did, that's the actual reason, or is it more that even when it's right, letting it go through on its own still feels too risky?

It's a bit like a self-driving car that can handle almost everything but someone still needs to be ready to grab the wheel. Trying to find that same moment for CS decisions specifically. What would actually have to be true before the easy, obvious ones went through without a check?

Not selling anything, genuinely just trying to understand where people draw that line and why. Would love to hear from anyone who owns this, whether you're the one approving it or the one who has to clean up after.

Thanks!

reddit.com
u/JuniorLeg6988 — 4 days ago

What return or refund call do you still not let happen automatically, even the ones that seem obvious?

Not the drafting of a response, the actual decision, approving the refund, granting the A-to-z claim, issuing a replacement instead of a refund, deciding a return doesn't need the item sent back. The stuff that currently needs a person to say yes.

If you've got any part of that automated, or looked into it: what's the one thing you still keep a person on, even for what looks like an easy, obvious case? Was there a specific time it went wrong, or almost did, or is it more that even a correct call still feels too risky to let go through untouched?

Return abuse alone is enough of a headache without adding software making the call on top of it, so I'm curious where people who've actually dealt with this draw the line.

Not selling anything, genuinely just trying to understand where that line sits and why. Would love to hear from anyone dealing with this, whether you're the one making the call or cleaning up after a bad one.

Thanks!

reddit.com
u/JuniorLeg6988 — 4 days ago

(rant) Disapointed By Life and my career(I will not promote)

Basically what the title says. I got a math degree and I have a psychology degree as well. I have been building things using computers since I was 7 years old(I'm 28 now) I have a mediu page. I get a lot of readers every time. Job recruiters contact me a lot. I have built a trading system that i'm happy with. I can solve hard problems. But I"ve been unemployed for 6 months. I've sent in lots of aplications, I talk with recruiters every day. I did a contract that lasted for a little while but it ended. honestly I jsut don't trust the job market anymore. As a college student before I started building the trading system I had a lot of ideas for things to build. But once I started building the trading system that was my only focus until I felt like I had figured it out sufficently well. Now I've been searching the internet like a mad-person looking for problems that Need to be solved. Some kind people on here were saying that I should try and solve problems that occur to me but honestly other than not having a good job I feel like my life is pretty good and i'm out of ideas. so I try and idk disect hard ai-related problems and I keep running into there being big companies already doing the same thing or problems just not being important enough to be solved. I really believe that if somebody is going to build smething it should be a lot better and not just a little bit better. I think what Peter Thiel was saying about looking for problems to solve not buisness ideas was 100% correct but i'm struggling to find real problems to work on. I have spent days and days looking into this and i'm starting to feel like a spammer posting on reddit asking people about their problems. I have the skills and the motivation but no real problem to solve that feels important and meaningful. idk I was going to make a good open source ai enabled writing app just to make a point that I think that it's stupid to make people pay a subscription for this. but idk other than making an open source repo to be like stop ripping people off. I have no clue what to build. Thanks for bearing with me through this rant.

reddit.com
u/JuniorLeg6988 — 5 days ago
▲ 3 r/mlops

What's an action you still won't let an AI agent perform autonomously in production?

I'm specifically interested in agents that can do things, not just generate answers.

If you have an agent that can technically execute some action — modify a database, issue a refund, deploy code, change infrastructure, update a CRM, send something externally, etc. — but you still require a human to approve or perform it, what's stopping you from giving the agent autonomy?

I'm especially curious about cases where the model itself is capable enough, but the surrounding system isn't trustworthy enough.

Was there a particular failure you were worried about or actually experienced? And what would you need to be able to verify/guarantee before you'd remove the human approval?

Not selling anything. I'm trying to understand where the boundary between “agent can do this” and “we trust an agent to do this” actually sits in production systems.
Thanks!!

reddit.com
u/JuniorLeg6988 — 5 days ago
▲ 0 r/sre

What's an action you still won't let an AI agent perform autonomously in production?

I'm specifically interested in agents that can do things, not just generate answers.

If you have an agent that can technically execute some action — modify a database, issue a refund, deploy code, change infrastructure, update a CRM, send something externally, etc. — but you still require a human to approve or perform it, what's stopping you from giving the agent autonomy?

I'm especially curious about cases where the model itself is capable enough, but the surrounding system isn't trustworthy enough.

Was there a particular failure you were worried about or actually experienced? And what would you need to be able to verify/guarantee before you'd remove the human approval?

Not selling anything. I'm trying to understand where the boundary between “agent can do this” and “we trust an agent to do this” actually sits in production systems.

Thanks!!

reddit.com
u/JuniorLeg6988 — 5 days ago

What's an action you still won't let an AI agent perform autonomously in production?

I'm specifically interested in agents that can do things, not just generate answers.

If you have an agent that can technically execute some action — modify a database, issue a refund, deploy code, change infrastructure, update a CRM, send something externally, etc. — but you still require a human to approve or perform it, what's stopping you from giving the agent autonomy?

I'm especially curious about cases where the model itself is capable enough, but the surrounding system isn't trustworthy enough.

Was there a particular failure you were worried about or actually experienced? And what would you need to be able to verify/guarantee before you'd remove the human approval?

Not selling anything. I'm trying to understand where the boundary between “agent can do this” and “we trust an agent to do this” actually sits in production systems.

Thanks!!

reddit.com
u/JuniorLeg6988 — 5 days ago

How is everyone handling agent regression testing in CI without going crazy?

Hey everyone,
At my last project, we spent hours every week manually spot-checking agent runs because every minor model tweak or context update seemed to silently break tool calling downstream.
Traditional unit tests don't fit because LLMs are non-deterministic, but most eval frameworks only grade the final text response rather than the intermediate tool-call trajectory (did it pick the right tool, pass valid parameters, and recover if an API errored?).
I’m working on better tooling around automated agent regression testing and deterministic tool validation in CI/CD, and I’d love to know what your current setup looks like:
How do you test whether a prompt/model update broke your agent’s tool calling before shipping to prod?
Do you run tests in GitHub Actions/GitLab, or is QA still largely manual / ad-hoc?
What’s the single most frustrating part of your current agent eval setup?
Appreciate any insights or horror stories from your production setups!

reddit.com
u/JuniorLeg6988 — 6 days ago

How is everyone handling agent regression testing in CI without going crazy?

Hey everyone,
At my last project, we spent hours every week manually spot-checking agent runs because every minor model tweak or context update seemed to silently break tool calling downstream.
Traditional unit tests don't fit because LLMs are non-deterministic, but most eval frameworks only grade the final text response rather than the intermediate tool-call trajectory (did it pick the right tool, pass valid parameters, and recover if an API errored?).
I’m working on better tooling around automated agent regression testing and deterministic tool validation in CI/CD, and I’d love to know what your current setup looks like:
How do you test whether a prompt/model update broke your agent’s tool calling before shipping to prod?
Do you run tests in GitHub Actions/GitLab, or is QA still largely manual / ad-hoc?
What’s the single most frustrating part of your current agent eval setup?
Appreciate any insights or horror stories from your production setups!

reddit.com
u/JuniorLeg6988 — 6 days ago

How is everyone handling agent regression testing in CI without going crazy?

Hey everyone,
At my last project, we spent hours every week manually spot-checking agent runs because every minor model tweak or context update seemed to silently break tool calling downstream.
Traditional unit tests don't fit because LLMs are non-deterministic, but most eval frameworks only grade the final text response rather than the intermediate tool-call trajectory (did it pick the right tool, pass valid parameters, and recover if an API errored?).
I’m working on better tooling around automated agent regression testing and deterministic tool validation in CI/CD, and I’d love to know what your current setup looks like:
How do you test whether a prompt/model update broke your agent’s tool calling before shipping to prod?
Do you run tests in GitHub Actions/GitLab, or is QA still largely manual / ad-hoc?
What’s the single most frustrating part of your current agent eval setup?
Appreciate any insights or horror stories from your production setups!

reddit.com
u/JuniorLeg6988 — 6 days ago

Would you pay for continuous verification of an agent's read-only SQL, or just keep the script?

Setup: an LLM agent takes a business user's question, writes a SELECT against the warehouse, runs it, shows the result. Read-only — it never writes to the database.

The failure I care about: the query runs fine and returns real rows, they're just the wrong ones. Wrong join key, wrong filter, stale schema understanding. Nothing errors, the number looks plausible, it's wrong.

Having a second LLM write its own query and compare doesn't really work — if the first agent misread the schema, the second one reading the same schema tends to misread it the same way, and you get confident agreement on a wrong number. What seems to actually work is deterministic checks that don't interpret the question at all: does the sum across a breakdown match the unsegmented total, does row count stay in an expected band, do foreign keys still resolve one-to-one. A wrong join changes cardinality, so a fan-out check catches a lot of it in plain SQL with no model involved.

I'm considering building this, and I want to be straight about where the line would be:

Free / open source — the check engine. Point it at your DB, it reads FK constraints and auto-generates structural checks, you run it against your test cases, get pass/fail. Runs when you run it.

Paid — the same checks running continuously against production traffic: results stored over time, dashboard, Slack alerting, and detection of when a schema change has made an existing check stale so you know which assertions to update.

What I actually want to know:

  1. How do you catch "ran fine, wrong rows" today, if at all?
  2. Do your checks run pre-ship, or continuously against production?
  3. When a schema change breaks them, do you find out from a migration diff, or because numbers came out wrong?
  4. The honest one: is that paid tier worth expensing, or would your team keep the homegrown script regardless? If it's worth paying for, what's the feature that tips it?

Genuinely fine with "we'd keep the script" as an answer — that's the thing I'm trying to find out before building.

reddit.com
u/JuniorLeg6988 — 6 days ago

How are you evaluating agents that write SQL against live databases?

I've been digging into agent evaluation for setups where the agent writes and runs SQL against a live database (Snowflake, BigQuery, etc.) and shows results to users.

The failure mode that seems underserved: the query executes fine and returns real rows just the wrong ones. Wrong join, wrong filter, stale understanding of the schema. Nothing errors, the output looks plausible, but it's wrong. Static eval sets with prewritten "golden" answers don't hold up here, because the correct answer changes as the data changes.

Interestingly, LangSmith has a cookbook recipe for exactly this storing labels as queries the evaluator runs at eval time to fetch current ground truth but it's DIY: you build and maintain that evaluator yourself. As far as I can tell, none of the major platforms (LangSmith, Braintrust, Arize) ship live data verification out of the box; online scoring generally falls back to reference-free LLM as judge.

I'm considering building a dedicated tool for this: connect your DB and your agent, and the evaluator independently queries the database to verify each output against what's actually there right now. Before I build anything, I want to know if this is a real problem for other people:

  1. If your agent queries a live DB, how do you catch "ran fine, wrong data" failures today?
  2. How often does that actually bite you in practice?
  3. What's your current eval stack LangSmith, Braintrust, Arize, custom scripts, nothing?
  4. Would you pay for this as a product, or just have Claude Code write you a one-off eval script?
  5. If you'd pay, what would make it worth it? If not, why not?

Not selling anything. Trying to figure out whether this is widespread before building...

Clarification: read-only queries. The agent isn’t writing to the database, it’s translating user questions into SELECT queries and showing the results.

reddit.com
u/JuniorLeg6988 — 7 days ago

How are you evaluating agents that write SQL against live databases?[I will not promote]

I've been digging into agent evaluation for setups where the agent writes and runs SQL against a live database (Snowflake, BigQuery, etc.) and shows results to users.

The failure mode that seems underserved: the query executes fine and returns real rows just the wrong ones. Wrong join, wrong filter, stale understanding of the schema. Nothing errors, the output looks plausible, but it's wrong. Static eval sets with prewritten "golden" answers don't hold up here, because the correct answer changes as the data changes.

Interestingly, LangSmith has a cookbook recipe for exactly this storing labels as queries the evaluator runs at eval time to fetch current ground truth but it's DIY: you build and maintain that evaluator yourself. As far as I can tell, none of the major platforms (LangSmith, Braintrust, Arize) ship live data verification out of the box; online scoring generally falls back to reference-free LLM as judge.

I'm considering building a dedicated tool for this: connect your DB and your agent, and the evaluator independently queries the database to verify each output against what's actually there right now. Before I build anything, I want to know if this is a real problem for other people:

  1. If your agent queries a live DB, how do you catch "ran fine, wrong data" failures today?
  2. How often does that actually bite you in practice?
  3. What's your current eval stack LangSmith, Braintrust, Arize, custom scripts, nothing?
  4. Would you pay for this as a product, or just have Claude Code write you a one-off eval script?
  5. If you'd pay, what would make it worth it? If not, why not?

Not selling anything. Trying to figure out whether this is widespread before building...

Clarification: read-only queries. The agent isn’t writing to the database, it’s translating user questions into SELECT queries and showing the results.

reddit.com
u/JuniorLeg6988 — 7 days ago

How are you evaluating agents that write SQL against live databases?

I've been digging into agent evaluation for setups where the agent writes and runs SQL against a live database (Snowflake, BigQuery, etc.) and shows results to users.

The failure mode that seems underserved: the query executes fine and returns real rows — just the wrong ones. Wrong join, wrong filter, stale understanding of the schema. Nothing errors, the output looks plausible, but it's wrong. Static eval sets with pre-written "golden" answers don't hold up here, because the correct answer changes as the data changes...

Interestingly, LangSmith has a cookbook recipe for exactly this — storing labels as queries the evaluator runs at eval time to fetch current ground truth — but it's DIY: you build and maintain that evaluator yourself. As far as I can tell, none of the major platforms (LangSmith, Braintrust, Arize) ship live-data verification out of the box; online scoring generally falls back to reference-free LLM-as-judge.

I'm considering building a dedicated tool for this: connect your DB and your agent, and the evaluator independently queries the database to verify each output against what's actually there right now. Before I build anything, I want to know if this is a real problem for other people:

  1. If your agent queries a live DB, how do you catch "ran fine, wrong data" failures today?
  2. How often does that actually bite you in practice?
  3. What's your current eval stack — LangSmith, Braintrust, Arize, custom scripts, nothing?
  4. Would you pay for this as a product, or just have Claude Code write you a one-off eval script?
  5. If you'd pay, what would make it worth it? If not, why not?

I just want to figure out if this is a widespread problem before building a fix! Thank you!!

Clarification: read-only queries. The agent isn’t writing to the database, it’s translating user questions into SELECT queries and showing the results.

reddit.com
u/JuniorLeg6988 — 7 days ago

What problem makes you think Omfg this sucks there must be a better way? (I will not promote)

I’m an engineer and I like building software/AI systems, but I’m trying very hard not to build a solution in search of a problem.
So instead of asking for startup ideas, I’m curious about problems people actually deal with.
What’s something in your work that is genuinely annoying, expensive, repetitive, error-prone, or takes way more time than it should?
I’m especially curious about things like:
processes that still involve a lot of spreadsheets/copy-pasting/manual checking
information scattered across different systems
software you pay for but still have to work around
things you regularly have to verify or reconcile by hand
tasks where you’ve thought “why the hell isn’t this automated?”
I don’t have a product to sell. I’m trying to understand real problems before deciding whether anything is worth building.
If you have one, I’d be interested in what the problem is, what you currently do to deal with it, and whether you’ve already tried software that was supposed to solve it.
Thank you!!

reddit.com
u/JuniorLeg6988 — 8 days ago