u/Rough_Practice7631

Are AI companies still buying training, post training data or hiring their own data annotator, synthetic data specialists etc..?

Not sure if this is the right sub-reddit, but I'm wondering if anyone has knowledge in actual trend within AI companies, especially the bigger labs. No doubt they still have agreement to collect and buy new data for training their models, but is the trend going down?

Similarly for alignement, post training and fine tuning are they actually buying? Or is there a shift towards internalizing the capabilities? I'm seeing big labs hiring for synthetic data generation, sometimes even data annotations... or the other hand I'm also seeing startups getting traction by focusing more the infra for doing RL, alignement etc.. than the data itself (they call it "AI gym" or "world model").

What are you thoughts on this and where do you think the "traditional" data-selling industry is going?

reddit.com
u/Rough_Practice7631 — 7 days ago

I partnered with a lawyer who saw the same thing killing startups over and over. We built the thing he wished those founders had.

I'm the technical cofounder, my partner spent years dealing with founder disputes as a lawyer. What he kept seeing wasn't that founders lost in court. It's that they lost their startup while waiting for court.

Basically, cofounder conflicts freeze everything. The product stops shipping, the team can't hire, the round can't close, and a startup that stalls for that long is usually already dead by the time a judge rules. Even the "winner" walks out of a company that no longer exists.
The signed founder agreement doesn't prevent this. It just tells you who was right, eventually. That's the gap we're trying to close.

We're building Goodvernance to flip that. Instead of paying a lawyer several thousand dollars to draft a founder agreement you will most likely never read, you answer a few plain questions and get a structured agreement, before incorporation, in language anyone can understand.

What's live now: vesting is tracked, "what happens if a cofounder leaves" is a simulator that lets you see concretely what could happen, IP and assets are visible, amendments are versioned. Where we're heading: self-execution, where the terms just happen instead of needing to be demanded.

We also put a bunch of free notes on the community page (equity splits, vesting, 83(b), IP, Delaware), useful if you're thinking about launching your startup (or even if you already have).

Site: https://www.goodvernance.com/
The essay behind the vision: https://www.goodvernance.com/goodfounders/from-promises-to-execution
Founder notes: https://www.goodvernance.com/goodfounders#notes

Genuinely want feedback and any thoughts? Would you actually use something like this?

u/Rough_Practice7631 — 1 month ago

What I learned discussing with a lawyer dealing with startup founder drama.

TL;DR: Skipping pre-incorporation founder agreement kills a ton of startups, every week.

I recently had a discussion with a lawyer who handles startup founder disputes every week. The biggest thing I learned: the vast majority of co-founder breakups he handles happen because teams left equity, roles, or exit scenarios ambiguous prior to formal incorporation. They assumed they would "figure it out later," which turned into mis-alignment or conflict once things got real. It's surprising to see how many ideas die on a weekly basis because founders felt an early agreement was too awkward or premature.

Curious to know from the founders here who actually signed a founder agreement before your official corporate filing: How early did you draft it, and what do you wish you had included or excluded in hindsight?

reddit.com
u/Rough_Practice7631 — 1 month ago
▲ 8 r/Rag

Are you using embedding LLM models at scale? What are the best practices you follow to optimize throughput ?

By at scale, I mean at least 50k documents (e.g., 10-20 pages each) with relatively high frequency (say daily). Curious to know what empirical findings you have and what is the SOTA on this?

reddit.com
u/Rough_Practice7631 — 1 month ago

Are you fine tuning LLM or SLM ? If so, why and what data do you use?

I'm curious to know what are your use cases for fine tuning LLMs or SLMs, i.e., is it to teach domain knowledge / enforce style or constraints / save on cost (with SLM) ... ?

And for those who do fine tune, what data are you using ? Is it mostly open source or do you buy datasets ?

Thanks for sharing your thoughts on this,

reddit.com
u/Rough_Practice7631 — 2 months ago