r/AgentsOfAI

▲ 6 r/AgentsOfAI+7 crossposts

Uber's President Just Confirmed the Internal AI Adoption Leaderboard Is Real

Everyone's reacting to the "less people in 5 years" line. That's not the part I'd sit with.

The part that actually matters is how Uber decides — an adoption leaderboard, tracking who's using the tools and how much, feeding straight into headcount math.

That's not a hypothetical for some future reorg.

That's a live measurement system, running today, on people who have no idea they're on it.

 

I've watched that exact math play out before — in concrete and steel, not a dashboard, years before anyone called it AI.

I had the opportunity to be involved in the early design stage of an expansion project for a famous beverage manufacturing plant in Taoyuan, Taiwan – back in 2021. The beverage brand name is so famous, you'll instantly recognize it. So, I won't name it here.

Our team got to work on cool stuff - latest advanced technologies in high-density and automated racking system, bottle conveyor system, robotics, beverage packers, clean room environment, etc. – things that are expected in a high-tech. manufacturing plant nowadays.

Looking at the projected 10-year production forecast, with the given magnitude of the hardware, I would say they are planning to go big.

It's quite a sizeable expansion.

And you would think that they'd increase their headcount proportionately, right?

You'd be surprised. There IS headcount increase, but not as proportional.

It seems as though the machines were taking more centre stage than the humans. Even the office space increase wasn't even a top priority in the design. Their existing office layout can still accommodate the projected increase in manpower.

It was as if human beings are being set aside to make room for more artificial things – even though what they produce are meant to serve human beings.

Kind of ironic, isn't it?

That was back in 2021 before AI come into the picture. Now the compression is even more acute, it seems.

https://preview.redd.it/twllzlreobkh1.jpg?width=1024&format=pjpg&auto=webp&s=1c757cce2c856bc267f2ba2acc634fd325ed9ad6

__________

Different guest, same fork in the road: does the tool serve you, or does it just get pointed at you.

The industries change.

The question underneath never does — who's holding the ledger, and whether you're the one reading it or the one being read.

A 2026 WRITER survey backs the pattern from the outside too: 75% of execs privately admit their AI rollout is mostly for show, while the people actually inside the tooling get promoted 3x more often and ship 5x more.

That's the leaderboard, confirmed from a different angle.

 

Genuinely curious where you land: is a visible adoption leaderboard a fair way to measure a team, or is it just a slower-motion version of the same cut?

 

Clip credit: 20VC with Harry Stebbings & Uber. DM for credit or removal requests.

 

The leaderboard measures usage. It doesn't measure ownership — and the difference between those two is the actual thing I built the working system around.

u/cen6wkf — 13 hours ago

Why we hate AI generated slop content so much.

You can instantly tell Ai generated stories within a couple sentences.
And they are soooooooo bad.
Bland, clichee, predictable, bad writing through and through.

Many people use AI for writing and flood the Internet with low quality content.
And reading AI writing is like having brainfog. You can't actually READ it.
It has such low information density you MUST skim it.

Every AI picks the FORM of the answer before content. And with a little experience, you know how the answer is gonna look like before it was generated. AI can't "learn" faster, change faster than we grow tired of it.

And this technology already has PEAKED.
it CAN't improve
they overfit on benchmarks, that's all.

And this remains true after YEARS of "improvement".
Only thing it has gotten good at is styling of HTML CSS pages.
It's a loser technology.
And I'd short all of it.

People ho disagree either genuinely don't see it, have no taste and don't read, or write their own slop stories and don't want me to take away their fantasy that they are an "author" or "software engineer" now.

reddit.com
u/South-Mongoose-4743 — 1 day ago

How to generate B2B leads with AI agents

many people trying to use AI agents for B2B lead gen set up a simple agent loop with an LLM, point it at a scraped list of linkedIn URLs or emails and tell it: "research this person and write a personalized cold email."

Then they wonder why reply rates are under 1% and their sending domains get flagged within 3 weeks.

The issue isn't that agents can't do sales work but you need to explicity give it enough context cause agents are great for research but the issue is people are using agents for execution without context. An agent generating 1,000 cold messages a day from a static list is just automated spam.

If you want an agentic GTM workflo, the agent needs to operate across three distinct layers:

1, Your agent shouldn't start by looking at people or volumes but events so instead of scraping static directories, an effective agent workflow monitors real-time triggers like someone asking for alternatives to a competitor on reddit or discord or a company starring a competing open-source repo on github or a target account posting a job description with specific tech stack migrations.

If the agent doesn't have a live trigger, it shouldn't reach out at all and that's the thumb rule.

  1. A single post on reddit or a github star isn't a lead but surely a signal but the hard engineering part is connecting that signal back to a company and pulling the full context, so things like who is this company? have we interacted with them before? or where are they on the buying journey?

This is where platforms like Scale Intelligence fit in where it acts as an autonomous market intelligence engine with its own GTM agent (Algebra). It connects to 75+ data sources across reddit, github, discord, job boards, and CRMs and continuously mapping buyer readiness so your agents or your human reps via Slack and get fed high-intent accounts with the exact context already resolved. You can also plug your own agents (Claude Code, Cursor, Codex) into it directly via MCP.

  1. Once the signal is detected and resolved, the agent shouldn't just blast an email immediately but play based on intent level. So if you're building GTM agents right now, focus on giving your agent live market memory and real-time triggers rather than just optimizing the prompt template.
reddit.com
u/astrouis — 23 hours ago

AI coaching is starting to make ridealongs feel outdated

Thinking about how outdated the classic sales ride along feels now. A manager can lose half a day driving between locations just to sit in on one or two conversations. And even then you're coaching based on a tiny sample of what that rep does all week. We've been piloting a few AI tools for this and I'm starting to think coaching might be one of the better uses for them. Instead of trying to catch the right conversation in person you can record actual customer conversations and have the tool surface the parts worth reviewing. Could be a rep skipped an important discovery question or they handled an objection well. Maybe another issue keeps showing up across the whole team.

The part I like is that the manager still does the coaching. AI isn't replacing the 1 on 1s or telling managers how to run their team. It just gives them more context before they sit down with a rep. You can coach from what actually happened instead of trying to remember one ride along from two weeks ago. Still early for us since we're testing a few options but I can see this being useful than the classic old ride along model. Especially if you're managing reps across a bunch of locations.

reddit.com
u/Long-Ad7623 — 1 day ago
▲ 20 r/AgentsOfAI+7 crossposts

David Gerard (Pivot to AI): the internet's used up — now the same scrapers are hammering smalll self-hosted servers like mine, non-stop.

David Gerard runs Pivot to AI oon a server that costs him €7 a month.

Right now, something wearing a fake Chrome mask is hammering it — hopping IP addresses so he can't even block it properly, ignoring robots.txt because robots.txt was never a wall, just a sign nobody was required to read.

He's not a company.

He's not a platform.

He's one guy, doing his own sysadmin work, at 11pm, because the industry ran out of the free internet and started eating the cheap end of it instead.

Not stolen. Just... takenn, quietly, at scale.

 

I've watched this exact shape happen before — just slower, and on paper instead of a server log.

Circa 2005, Malaysia. I was Assistant Technical Manager for one of the largest construction main contractors in the country. We were compiling tender documents for a factory job — flat-flooring work, strict F-numbers, the kind of spec that keeps a forklift's raised forks from clipping the racking on a narrow run.

A subcontractor walked in to drop off her quotation. She glanced at our papers, open on the table.

And she went pale. I heard the gasp.

"这是我写的,为什么会在这里?" — This is what I wrote. Why is it here?

Word for word hers. Now sitting under our company's logo and headings.

She looked at me. I looked at her. She was waiting for an answer I didn't have.

Then her eyes flickered — a thousand thoughts passing through in a second — and she said, "没关系。我可以再写过。" — Doesn't matter. I can write it again.

And she left. Good for her.

https://preview.redd.it/o6qawkco45kh1.jpg?width=1024&format=pjpg&auto=webp&s=d9164c398628663f07d2b86343d59947ae045aa0

________

Every one of these stories eventually lands on the same fact: the exposure runs downhill, from the platforms with lawyers down to the servers with none.

 

If you're running anything on a boxx that isn't Amazon or Google's, drop your own scraper-traffic story below. I want to see how far downhill this actually goes.

 

Clip credit: David Gerard — full video on The Tech Report's channel. DM for credit or removal requests.

 

Rohan's not the only one who found out the hard way that "small" doesn't mean "safe" — the actual mechanism for making that stop is one honest look away.

u/cen6wkf — 1 day ago

Anyone else finding Fable burns through Max plan limits ridiculously fast?

u/astrouis — 2 days ago
▲ 66 r/AgentsOfAI+8 crossposts

Ed Zitron just explained why your boss can't tell if the AI-generated model is actually right

Executives don't lack tools.

They lack a ruler.

 

That's the actual claim Ed Zitron made — not "AI is bad," but that the people signing off on AI-assisted work were never equipped to check it in the first place.

They see a document that looks finished and call it done, because "finished-looking" is the only bar they've ever had to clear.

 

For anyone whose whole job is catching the thing that looks fine and isn't — this isn't a tech story.

It's a story about who gets trusted, and why it's rarely the person who's actually right.

 

I've been on the other side of that exact gap.

Long before spreadsheets and dashboards, mine had a tape measure in it.

 

I was working as a Site Engineer for a Singaporean construction company building a primary school in Chua Chu Kang district back in 1998. Time flies. Just graduated from university. Figure I get some site experience first.

One day, I came to the project site. And I saw the newly delivered precast half-flight staircase lying on the ground next the building. I asked around to find out why wasn't it crane-lifted to position, which is between 1st and 2nd floor. And I was told the measurements were off. They couldn't fit it nicely on place.

And so, I went to work. I took my measuring tape, measure the staircase, and recorded the lengths, widths and whatnot. Then I went up to the building's 2nd floor — where the staircase was supposed to fit and meet. And I swung my measuring tape across the length of space between the positions where the 1st and last step of the staircase supposed to sit on. And took the site measurements too.

Then I went back to my office, took out the construction drawings from the drawing rack, lay it on the meeting table. And with a piece of paper, I started drawing it out. I knew full well the measurements I got will not exactly match that in the drawings, because — you know — site tolerances are still allowed and anticipated in the BS Code of Practice.

And then, through calculations, I found it. The measurements were way out of tolerance limit. No wonder the staircase can't fit. The blame squarely landed on our RC works sub-contractor. They screw up the levelling of the building.

Each of us supposed to have an "internal ruler" we rely on, to judge whether things look good or bad. For me, back then, it was Pythagoras and a fresh sheet of paper. My boss, years later, called his the same thing in different words — his "feel," thirty years deep. Kevin O'Leary's is knowing he can smell bullshit from a mile away.

So, coming back to these leadership people that Ed Zitron was attacking: don't they have their "feel" of things before shit hits the fan? Don't they use their "internal ruler" to measure it for themselves? Can't they smell bullshit from a mile away?

https://preview.redd.it/1j7kqeub7yjh1.jpg?width=1024&format=pjpg&auto=webp&s=7c3ef0563fe287471d2b68fbd35d2e35806943c4

________

Every post on this account keeps circling back to the same thing, whichever industry the clip's from: the people getting quietly pushed out are rarely the ones who got it wrong.

 

Drop your take: what's your internal ruler, and who around you doesn't have one?

 

Clip credit: Ed Zitron (Better Offline) on Adam Taggart's Thoughtful Money. DM for credit or removal requests.

 

If your ruler's ever been right and still overlooked, it's worth seeing what building past that actually looks like.

 

u/cen6wkf — 2 days ago

Replit CEO says: "By next year, using a computer will be optional.."

Amjad Masad (CEO of Replit) posted this on X:

>By next year, using a computer will be optional. Work will radically change.

what people here think. Do you see this happening that soon with AI agents and better interfaces?

reddit.com
u/nameaval — 2 days ago
▲ 13 r/AgentsOfAI+8 crossposts

Lauren Tan (Cursor engineer): I stopped writing code. Now I run quality control on a kitchen of agents.

“你在帮人倒米吗?“

Lauren Tan didn't get replaced by her own tooling.

She got promoted by it — and nobody handed her that promotion.

She built the case for it herself, one lint rule and one CI gate at a time, until the argument was undeniable.

That's the part nobody's really talking about when they talk about AI and engineering jobs: the shift rewards the people who go looking for the leverage first, not the people who wait to be told it's safe to look.

 

That "build the case yourself" instinct is exactly what clicked for me watching my own son learn to run a team instead of carry it.

My son started playing 王者荣耀 (Honor of Kings) since he was a teenager — a 5v5 multiplayer battle arena game where you manage a roster of specialized heroes, growing and levelling up their strengths through battles and gear.

In his early gaming days I could hear him cursing and swearing from his room — bad coordination, worst teammates. There was a phrase we used for a bad teammate in my own career — 帮人倒米*, a Cantonese idiom that literally translates as helping someone tip over their own grain container, meaning ruining or sabotaging someone's livelihood.*

But the cursing became less and less. He got good at managing his heroes and coordinating with his team. He started climbing the leaderboard. People started noticing him and his team. Then, in college, he started getting invited to tournaments — cash prizes when he won, and one lagged-connection loss at a KL tournament he still suspects was foul play.

Time has changed — my dad would've killed me for wasting my teenage years on video games.

Now he's in university, still playing, still winning tournaments and cash prizes with his team.

Why I'm bringing this up: I always thought these AI agents are kind of like the heroes my son uses in the game. Your skill is in your managing these heros and how to grow them, level them up to serve your purpose. You don't go down to the battle yourself. You engage the heros to do it for you.

The skill is in the managing.

https://preview.redd.it/xiee128173kh1.jpg?width=1024&format=pjpg&auto=webp&s=638e2d62448f1e60434b937b5a21ea332501c595

__________

 

I keep walking into the same room wearing a different name on the door — the accountant's room, the analyst's room, now the engineer's.

Every time, someone's being told the machine is coming for their hours, not their name on the work.

 

Drop your take — are you already the head chef of your own stack, or are you still doing all the cooking yourself?

 

Clip credit: MTS (Monitor The Situation) — full video on their channel. DM for credit or removal requests.

 

If encoding your own judgment into the system sounds like the actual leverage skill here, the mechanism I built around exactly this is one link away.

u/cen6wkf — 1 day ago

Should a review agent see the implementation thread, or only the diff?

Been testing a three-role setup on a small code change: one model handled the implementation, another reviewed it, and a third drafted the release notes.

Until now, I did this manually. I’d open another chat, paste in the diff, sometimes include notes from the first model, and ask what it missed. I hadn’t really considered whether those notes were helping the reviewer or anchoring it to the original explanation.

For this run, I added three providers through BYOK in MiniMax Code. The part I wanted to test was the handoff, so I gave the reviewer the diff and test output but none of the implementation conversation.

It caught one case the first model missed. An API response field was typed as string | null, but the patch treated it as always present. Every test fixture happened to populate the field, so the tests still passed. The reviewer flagged the missing null branch.

That’s only one run, so I can’t say the missing context caused the catch. It did make me wonder whether a review agent should start cold and ask for the reasoning only when it needs it.

There are downsides. Cross-provider calls were slower, I had three separate quotas to watch, and an isolated reviewer can flag a deliberate tradeoff simply because it doesn’t know why the decision was made.

For people separating implementation and review, do you give the reviewer only the diff and test output, or include the design reasoning too?

reddit.com
u/coffeebeforelogicc — 1 day ago
▲ 5 r/AgentsOfAI+3 crossposts

Built a fully local agent browser that cuts token usage ~32x by compressing webpages

Browser agents are given raw HTML and have to sift through and find out what's clickable themselves, which is a huge pain and waste of tokens. 

I’m building a browser that compresses webpages before passing information to agents, cutting token use by around 32x. 

Weak local models get constrained to one grounded choice at a time so they can't hallucinate actions or element IDs. Stronger models can plan multiple steps ahead. 

Actions are also policy gated, so riskier actions like purchases always require approval. 

It works with Claude Desktop, Cursor, VS Code, and Codex CLI over MCP, or directly through a local API if your setup doesn't use MCP.

u/nuterralabs — 2 days ago

A linux phone turns agents into a Black Mirror episode.

I wanted to give eyes, ears and a voice to my agent. So I bought it a Linux phone.
The phone is fully dedicated to the agent. I talk to it via Telegram on my iPhone.

It didn’t take long before the agent could use the camera, microphone and speakers to interact with what is around it. It has also gps, gyroscope and a sim card.

I travel a lot and my original idea was for it to be my personal security cam manager for my stuff (using the phone’s camera). But this thing is kinda eerie now. It can wake itself up when it wants. It uses realtime api by openai to talk and reads the conversation as it’s happening… it “sees” from the camera. And it does everything while the display is turned off (I thought Linux Touch didn’t allow background activities after being idle for a while).

The fascination for this is currently beating the spookiness of it all. So I’ll keep on using it. For privacy reasons yesterday I switched the model to a local qwen3.8-27b running on my RTX 5000 Pro at home.

u/Valuable-Run2129 — 4 days ago

Drowning in context-switching between Slack, Notion, and Jira how do you keep track of everything?

My team and I are losing hours every week just trying to reconstruct past decisions. We constantly jump between Slack threads, Jira tickets, and Notion docs, and half my day is spent asking "where did we agree on this?" or digging through notification backlogs.

How are you guys managing context switching and keeping track of active decisions across multiple tools without losing your mind? Any specific workflows or tools that helped?

UPDATE:

Done some digging and came across Sugarbug. It's supposed to continuously aggregate background context and track temporal updates across Slack/Jira so you don't lose context. Reviews online seem pretty solid, but has anyone here used it in practice? Want some real feedback before trying to roll it out to the team

reddit.com
u/Ill_Pride_9537 — 3 days ago

How do you gate what your agents are actually allowed to do in prod?

In our company we tried to send emails and touch Stripe with the agent but it feels not really save.

Our first version was just if/else in the tool wrapper. Allow Stripe, deny refunds over some amount. It worked for a short time but after defining multiple tool definitions it was kinda messy and we needed a separate policy layer while also putting a human in the loop when the policy definitions are not enough.

What we really want is to block the API call while a person looks at the tool the payload and hits approve or reject. For now we handled that with automated Slack messages, but ther're slow to respond, and it always took them about 40 minutes to process a request for us.

We also don't really like handing the agent a long-living API key, which in hindsight is one prompt injection away from a very bad day...

So I'm curious how you currently solve these problems:
- Policy as code, or still hardcoded in the tool layer? Anyone using OPA for this or is that overkill?
- Has anyone made in-line human approval not feel terrible?
- Do you scope credentials per action, or is everyone still passing the real key?

Disclosure so it's not weird later: I'm building in this space. Not pitching it here. I mostly want to know if everyone else solved this and I missed it.

reddit.com
u/Excellent-Park-1160 — 3 days ago

I built a system to host all of your different agents and prompts

The system allows managing multiple agents at scale.
It’s complex and complicated so the best to work with it is to ask your Claude or any other agent to orchestrate your business or operations on that system.
You can also use the systems father agent that will build everything for you.
Happy to hear your thoughts and to help you quickly setup your business on my platform.

u/Particular-Tie-6807 — 3 days ago
▲ 15 r/AgentsOfAI+6 crossposts

Economist Molly Kinder: the "safe" job wasn't safe. It was just priced high.

Bloomberg's own reporting already answers the question this clip raises — is the "messy middle" projected or already happening?

At Commonwealth Bank of Australia, Microsoft, Uber, and Hyatt, it's already happened: sizable call-center headcount cut using automated phone and chat systems, savings already banked.

Economist Molly Kinder's point isn't a forecast.

It's a line item that's already closed.

 

I've watched this exact math play out before superior technology ever touched a keyboard.

Suncon was getting jobs overseas. One of the countries we went to was India. We were building infrastructure — roads and bridges there. I didn't go. But my seniors went stationed there. When they came back during their scheduled holidays, one of them, a project manager, told me this story.

It so happened, that building roads and bridges inland means clearing jungles and passing through villages. As they were doing it, of course they engaged local villagers to be their workers and supervisors. Well, of course building infrastructure means bringing in heavy machineries, such as excavators, bobcats, mobile-cranes, 4-wheel-drive land-cruisers, etc. You know — the usual.

But the local villagers weren't happy. They complained that all these machineries have deprived the local population of their means of making a living. They have so many mouths to feed. A lot of them are quite poor. And many of them are very hunger for work.

And so a huge argument broke out. They even spitefully challenged our project team that they vast manpower was more superior than our machinery. I was so surprised when I heard it. How can they say that? How was that even a reasonable challenge, you know. My curiosity had the best of me.

Well, the project manager came up with an idea. He said, since they're so confident of their manpower, why don't we have a competition. Let's do a challenge of moving earth from point A to point B for our excavator/mobile-crane operator versus their vast manpower.

And they accepted.

At the day of challenge, the project manager set up two huge piles of earth at point A for both teams. The local villagers' team had their "vast" manpower formed a long-ass line between A and B, and started moving earth, with their primitive buckets and whatnot.

For our team, we set up our mobile-cranes, excavators and bobcats on strategic locations. And off we go.

You can guess the result. We won by a large margin. We were obviously much faster and better at it.

After that, the local villagers concede defeat.

The math is actually quite similar here. The one with superior technology always wins. This AI-take over is no different.

https://preview.redd.it/ntbpt4qenqjh1.jpg?width=1024&format=pjpg&auto=webp&s=cea883ec3b203b4c87a315d4edcf410b183609c8

__________

Every time I dig into one of these stories the shape repeats: the tool doesn't ask permission, it just wins the argument by moving faster than the objection can be raised.

 

Ever watched something you thought was irreplaceable lose, and lose fast? Drop your take below.

 

Clip credit: Center for Humane Technology — full video on their channel. DM for credit or removal requests.

 

If you want to see how I'm actually building leverage against this instead of just watching it happen, it's just one link away

u/cen6wkf — 3 days ago

Scottish Water turned capital investment reporting into a conversational data experience with Databricks Genie

Scottish Water had plenty of capital investment data. The problem was getting the right answer to the right person quickly.

Project teams often had to search through a large collection of reports—or rely on analysts and data specialists to extract information from underlying tables. That created duplicated reporting work, slowed decision-making, and made valuable project data harder for non-technical users to access.

Their answer was SPARK, an internal natural-language interface built with Databricks Genie and embedded in Microsoft Teams through Copilot.

Instead of asking, “Which report has this information?”, teams can ask questions such as:

  • Which open project risks are expiring this month, and who owns them?
  • What is the current live risk score for a project?
  • Which risk has the highest exposure?
  • Who is the future contractor for a project?

Genie translates those questions into queries against governed data in Unity Catalog, using curated gold-layer data and metric views to keep business definitions consistent.

The governance and delivery details are especially interesting:

  • Business rules, fiscal-period conventions, project IDs, and milestone logic were explicitly configured.
  • Worked examples and benchmark queries were used to improve and repeatedly test answer quality.
  • Adoption, conversation patterns, query performance, and cost per user are monitored in production.
  • Databricks Asset Bundles and Azure DevOps support repeatable deployment across development, test, and production environments.

The reported efficiency opportunity is meaningful: if 100 users ask three questions per week, saving just two to five minutes per request could recover roughly 10–25 hours every week.

The broader lesson is that conversational analytics is not just about adding a chatbot on top of raw tables. The quality of the experience depends on the foundations underneath it: curated data, shared semantics, governance, testing, monitoring, and integration into the tools people already use.

It’s a useful example of how organizations can make governed data more accessible without giving up trust or control.

What has your experience been with natural-language analytics? Does embedding it in an existing collaboration tool make adoption more likely, or do users still prefer traditional dashboards?

reddit.com
u/myth-buster9999 — 2 days ago

One Agent, Many Hats - The Trinity of Agentic System

Imagine an LLM Agent that learns on its own. That's the wild idea that kept me driven in the last few months and here i am with "One Agent Many Hats"
(Open Sourced it so you can experiment it for free)

I enjoy building automations and one of the frustrating aspect of building agents was coding everything around the logic.

While AGI is the next big thing, I imagined Autonomy and automations as next big leap in the agentic systems I built. A system that can learn on the go, expand its horizon with more interactions, just like we as humans learn bound by the rules.

So, after my previous paper on "Conversational Decision Intelligence", i dwelled deeper and tested multiple frameworks and inspired by how claude's operating model, came up with "One Agent, Many Hats - The Trinity"

Here, the agent is a individual LLM - Just like you & me which learns, corrects, builds knowledge on the go.

This is just the beginning and I want more brains to come in. So, happy to open source the code so that it can lead to something meaningful that AI community will build on.

Check it out. Link in Comments

reddit.com
u/sandeepkavety — 3 days ago
▲ 271 r/AgentsOfAI+4 crossposts

Do multi-app long horizon tasks on frontier models for free!

This is coarena.ai, where you can run long horizon computer-use tasks on frontier models for free!

u/Good-Baby-232 — 5 days ago