u/FirstStringPM

I’m trying to figure out if AI-generated legal work needs a “trust layer”

I’m trying to figure out if AI-generated legal work needs a “trust layer”

I’ve been spending the last few months talking to lawyers about how they’re actually using AI in their practices as someone who works in knowledge and innovation at a big law firm. The one thing that keeps coming up is that AI is getting very good at producing legal work product like briefs, memos, demand letters, contract analysis, research, etc.

But the workflow often still looks like (i) AI generates (ii) lawyer reviews (iii) lawyer hopes they caught everything.

As you can imagine, the consequences of missing something can be very different in legal than in most other industries.

A hallucinated citation, incorrect legal proposition, missed requirement, or unsupported statement can potentially mean wasted time, malpractice exposure, sanctions, a bad client outcome, or simply a lawyer putting their name on something they don't fully trust.

That got me thinking. What if legal AI needs an operational “trust layer” around the work product rather than just another AI that generates the work?

I’ve been building a very early prototype called Argos to explore that idea:

https://argos-eval.com

The basic idea is to evaluate AI-generated legal work before it gets used. Checking things like citations, factual support, requirements, consistency, and potential issues that a lawyer should review. The indirect value here is that each law firm is basically building their own legal benchmark for their own workflows. Not some general legal benchmark that could be found on the internet.

But I don't actually know yet if this is a real problem worth building a company around.

So I'm much more interested in hearing from lawyers and legal tech people here than getting people to sign up.

A few things I'd especially love to know:

- If you use AI to produce legal work, what are you most worried about missing?

- What do you currently do to verify AI-generated work?

- Are there certain practice areas where this is a much bigger problem?

- Would an automated “second set of eyes” actually be useful, or is this just adding another layer of review?

- If you think this is a bad idea, I'd genuinely like to know why.

If you're willing to talk for 15–20 minutes about how you actually use AI in your practice, I'd also love to hear from you. No sales pitch. I'm trying to figure out whether I'm solving a problem that lawyers actually care about.

I also have a high fidelity prototype that I could demo with you.

Looking forward to the engagement.

team@argos-eval.com

u/FirstStringPM — 1 day ago