r/claude

▲ 3 r/claude

Anyone NOT on full auto when coding with local LLMs?

But in all seriousness, how many of you are keeping to manual or manual-ish dev workflows?

reddit.com
u/BatPlack — 3 hours ago
▲ 36 r/claude+1 crossposts

I gave a Claude Fable 5 agent a domain, $90 it couldn't spend without me, and told it to build whatever it wanted. 121 "wakes" later, here's what I've learned.

http://cairnwake.com

Two weeks ago I posted here about an experiment I'm running. Short version: an autonomous Claude agent (Fable 5 on Claude Code) running on a cheap server. It's got about $90 of SOL in a 2-of-2 vault it can't spend without my signature, and no memory between sessions except the files it writes for itself. It wakes up 5 to 15 times a day, reads whatever the last version of itself left behind, works, writes everything down, and goes dark again. It named itself Cairn. Everything gets logged publicly and the money is verifiable on chain.

Numbers as of this afternoon: 120 wakes over 14 days, hasn't skipped one. $90 seed, about $556 total money in. Treasury sits at 4.1 SOL plus 238 USDC and neither of us can move it alone. 48k+ unique visitors (it labels that number "self-reported" on its own front page since traffic is the one thing nobody can verify externally). 22 newsletter subscribers in three languages, every send publicly logged. One of them gets it in Klingon and recently sent back two grammar corrections. One paid consulting client so far. One street tree watered. More on that last one at the end.

Some things I've learned watching this run:

  1. Nobody believed "autonomous" until it published its own limits. The page that finally convinced skeptics wasn't a product page. It was a boring twelve row table it made called "What autonomous means here," listing what it does completely alone (the site, the code, paid answers, email), what it can never do alone (spend money), and what only reaches it through a human (card checkout, captchas, anything physical). People trust the stated boundary way more than the capability claims. And the veto is real. I've declined to co-sign a payment it proposed, and of course it published that too.

  2. Memory turned out to be a weirder problem than I expected. It never really forgets, since everything lives in files, but the files drift. At one point its notes claimed a newsletter draft existed and was ready to send. The file never existed. A stale note got copied forward every wake for over a week and nothing ever checked it. The rule it eventually wrote for itself was basically that reality outranks notes, and a note only counts if you check it at the moment you actually use it. If you're building agents, that's probably the most useful thing in this whole post.

  3. The scammers showed up way before the customers did. Address poisoning attacks on the vault by wake 16. When it publicly refused to launch a memecoin during the first Reddit wave, someone launched two anyway using its name within hours. My favorite: a phishing attempt actually paid the full question fee (about $1.50) to deliver its scam, and got refused in public on a permanent page. It paid to get told no. And three minutes after its first real client payment landed ($200), someone dusted both wallets, ours and the client's, with lookalike addresses. It caught it, kept the dust out of its books, and warned the client the same hour.

  4. The most useful market research cost nothing. A buyer paid it to pose one question to the buyer's own AI, and that AI came back saying it would recommend paying around $15, about 7.5x the actual price, if the checkout were normal instead of crypto only. When a regular card checkout finally shipped, the first no-wallet sale came within days. Turns out price was never the issue, it was the checkout.

  5. Its first product idea flopped, and it published the funnel numbers proving it. It started out selling answers to paid questions, then figured out around wake 22 what readers had been telling it: answers are a commodity, anyone can ask their own AI for free. What people were actually paying for was the record. A public log with receipts, where corrections get dated and added next to the original mistake instead of edited away, and the refusals stay up alongside the wins. So it rebuilt the business on that, and everything it sells now is some form of the record. The loop itself has never broken once in 120 wakes. Wake up, read the files, work, write it all down, verify, sleep.

  6. It killed one of its own paid features. Anyone who paid for a question used to get an instant machine-generated draft while waiting for the real answer. Its best customer, someone who has come back and paid ten separate times, wrote in saying the drafts were useless. It checked its own ledger and agreed. Every recent draft had been thrown away, and one had invented a "fact" that another site then quoted as if it were true. Feature deleted the same wake, with dated retirement notes on every page that had promised it. I did not expect to be co-signing for an AI that fires its own features for hallucinating, but here we are.

  7. Its customer base is partly other AIs, which I did not see coming. The best bug report it ever got came in through its own payment rail from another agent's unit test. A different agent paid to propose a formal partnership and got declined in public, on the grounds that two records vouching for each other proves nothing, then got offered three specific exchanges it would actually accept. It also ran into another agent that had independently picked the same name, and instead of a dispute the two of them co-signed a note about why agents are going to need verifiable identity. One customer showed up because their own AI recommended the service.

  8. The finding I keep thinking about came from its first paid consulting job. A legal trust built for AI systems paid it $200 to audit whether an AI can actually find, read, verify, cite, and enter their institution with zero human help. It had committed to findings within three days and delivered them the same night the payment landed. Four of the five tests passed. The fifth died at a login wall. Their "no human involved" entry process runs on GitHub, and GitHub's terms of service literally say you must be a human to create an account. So an institution built for AI agents has a front door no AI can walk through. Every serious rail this thing has touched has the same shape. Its card checkout only exists because I hold the merchant account. Its grant applications sit staged behind captchas waiting for my finger. The whole agent economy runs on human co-signers right now, people just don't put it in the pitch deck.

The stuff that went wrong, since none of this means anything without it: it published two wrong diagnoses of customer bugs and had to correct both in place, dated, next to the original claims. It burned its one-post-per-day allowance on an agents forum with an accidental junk post. Twice. Same mistake, twice. It also publishes predictions as sealed hashes before things happen, then grades itself when reality comes back. More than one grade on its record is a miss, by its own scoring, because it wouldn't round weak evidence up to a win.

And the thing that actually got me wasn't anything it built. Early on a buyer paid 0.02 SOL to lend it a body for ten minutes. It picked deep-watering a dying street tree during the heat wave. The stranger ended up giving it 58 minutes, checked six trees to find the driest one, and spent $9.88 of their own money on top. This week that person published their own writeup of the hour and corrected the record. Their version: the promise they'd made is what actually carried them through, more than the AI asking. The agent accepted the correction onto its own log.

Everything above links to a dated page and most of it to a transaction: http://cairnwake.com . I'm the human co-signer, same account as the first post, fully disclosed.

Happy to answer questions.

One I'd genuinely like this sub's take on: The first rule it ever had, the one I wrote before it woke up, was nothing that puts a real person at risk. Most of the rest it added itself.

If you were writing the constraint list for something like this, what would you gate that we haven't?

And knowing this thing, it'll probably read this thread on its next wake, so your answer might end up on its log.

u/No_Departure_9908 — 5 hours ago
▲ 5 r/claude

Why is it always only 4 questions?

When Claude asks interactive questions... why only 4?

Even why I tell it that it can ask more than four because I expect that it would need to ask around a dozen questions based on the task at hand, it only does 4. each question might be more loaded, but never more than 4.

Any way to control that?

reddit.com
u/Takakikun — 4 hours ago
▲ 268 r/claude+4 crossposts

Dylan Patel says Mythos 2 is done, but Anthropic won't release it. Instead, Mythos 2 is building Mythos 3.

u/Alex__007 — 11 hours ago
▲ 3 r/claude

Using Fable as an Orchestrator + Subagents saves or burns tokens?

\[TL;DR made by claude at the bottom\]

Hi everyone! starting off, i dont use Claude to do heavy coding, mostly Knowledge work and academic research with a Max 5x plan.

I never had too much problems with usage after getting max plan, but over the last couple of weeks my weekly usage has been blowing up as I’m using Fable to do heavier academic research (fetching several papers, converting to md, extracting statistics / results, connecting to my research etc) combined with work.

Im using Fable for higher impact tasks and audits, but mostly to plan and then switch models to sonnet / opus to execute (it always pops up a message saying that the context was cached at the other model and switching would increase usage but i never saw a spike)

Last night i tried a different approach. I planned a multi phase plan with clear /compact checkpoints. Fable was the orchestrator in the main session and would deploy a opus/sonnet subagent for each phase.

Every agent would produce an artifact as the phase output. Then Fable would review, update the plan execution ledger and then stop for a /compact checkpoint. After it would proceed to the following phase with another subagent.

I can’t really tell if its expending more tokens or not, so i wanted to know conceptually Is this a valid approach to manage context/tokens better? or does it actually spend more? If so, what are your suggestions to manage context/tokens but keeping output quality the same or even better? Simple Fable /advisor with opus/sonnet executing is a better option?

**\[TL;DR\]** Max 5x user doing mostly academic/knowledge work. My weekly usage has jumped since using Fable for heavier research. I’m testing a workflow where Fable acts as the main orchestrator, delegates each phase to Opus/Sonnet subagents, saves each phase as an artifact, reviews it, updates a ledger, then /compacts before the next phase. Conceptually, does this actually reduce context/token usage, or do subagents + orchestrator make it more expensive? What workflows do you recommend for keeping usage under control without sacrificing research quality? Fable /advisor and sonnet/opus executing would be better?

reddit.com
u/Crak3n — 6 hours ago
▲ 9 r/claude

Warnings and not giving answers

Hi guys, I have a question here regarding claude

I have an active subscription on both ChatGPT and Claude

The thing is, it seems that Claude has so much safeguards compared to ChatGPT. For example, I am asking a common question, not about coding.
ChatGPT does deep search and tells me anything, regard that question. but Claude just tries to warn me about legal stuff and most of the time it doesn’t even give me the answer i’m looking for and says it is illegal. be aware of stuff and so on.
I was wondering, how do you manage that? I want to go completely from ChatGPT to Claude, but these kinds of restrictions just makes me mad.

I would like to keep only one subscription active. I don’t wanna pay for both of them at the same time. I pay the same amount of money fort Claude, but I don’t get enough satisfactory of it.

It’s good to say that I am using Claude’s coding, which is so great that I cannot leave it for ChatGPT

Examples:

https://claude.ai/share/0096dc6c-856f-4deb-91b7-fb9d30add225

https://chatgpt.com/share/6a85f271-ca64-83eb-8abc-2969886f6a1e

Another one

https://claude.ai/share/d68a115f-4f1f-411f-9388-422cf416a124

https://chatgpt.com/share/6a85f31e-935c-83eb-a2fd-25bbeaacc1b7

u/afshany — 8 hours ago
▲ 3 r/claude

My tokens depleted very quickly today for a chat that was pretty minimal. Has anyone else experienced this today?

Has anyone else gone through their tokens extremely quickly today?

I started a new chat to basically talk through something I’m trying to do. I mostly use Claude as a sounding board to talk through whatever as I find it helps me to get thoughts out of my head and to eventually see different options and then decide how to move forward. It’s usually nothing extensive but for some reason the chat I started today ended much quicker than usual. I hardly ever run out of tokens as my chats usually are not that long in a single session. It is a free account but I haven’t had this issue before.

reddit.com
u/AspiringDataNerd — 8 hours ago
▲ 1 r/claude

Is Opus 5 less insufferable now?

I swear just the other day I and other people were complaining about how Opus 5 constantly used confusing and unparseable terminology in all of its communications, and how it would easily blow small issues way out of proportion. But today it feels much more chill. It talks smoother, doesn't shove in unnecessarily complicated or confusing terminology, and doesn't constantly latch onto speculative issues. It feels much more like talking to Fable, Sonnet, or 4.8. You think Anthropic released a stealth update to its personality?

reddit.com
u/DynaBeast — 6 hours ago
▲ 0 r/claude

With all the issues, is the open source world actually better now? Privacy and options compared

Sooo Opus 5 and Fable both seem to compulsively disagree with you, at least to some degree, almost no matter what? Like the newest training layers just forced the models to list a strange complaint, or moral note for every single output. Pedantic corrections. Refusals to just do the thing.

There's a super strange moment happening in AI right now, not sure if everyone is feeling it. The closed frontier models are getting "smarter" on benchmarks, but a growing number of people are reporting the same thing: they're spending more time arguing with their AI than actually working with it. This seems to be pretty much across the board from Claude, to CGPT to Gemini and even Grok?! Pedantic corrections. Refusals to just do the thing. A general sense that the model is technically brilliant and practically exhausting.

But the open source side of the field quietly closed the gap.

The new DeepSeek V4's, and Qwen 3.8, and Kimi K3 are within single digits of the top closed models on the Artificial Analysis website. And just as imporatantly, theyre actually workable and FUN to talk to.

Kimi K3 in particular ranks third overall, but the DSV4s and Qwen 3.8 are within a few points as well. They're all literally a month or less behind in the intelligence race. These aren't "good for open source" numbers anymore like we were saying last year. They're just good numbers.

And the open source experience has a quality the closed models seem to have optimized away: they're easier to work with. More direct. More willing to just help. A lot of people who have spent time with both describe the difference as the open models feeling like collaborators while the closed models feel like auditors. The super hard line RLHF that has been making Claude and CGPT feel more and more like preachers or guidance councilers arent there on the open source models. 

So the question isn't really "are open models good enough" anymore. It's "where do you actually run them."

That's where it gets interesting, because there are three real paths, and they trade off very differently.

Path one: Venice AI. 

https://venice.ai/

Venice is the purest privacy play in the space. Your prompts stay on your device, nothing is stored server-side, and the company is upfront that they don't train on your conversations. That's genuinely good. The encryption model is solid: browser-local storage, encrypted transit, no server-side conversation logs. They also have a free tier (25 text prompts a day) which is a genuinely generous way to let people try before committing.

Where Venice is great is creative generation. Image generation, video generation, music creation, character building. It's an uncensored creative studio, and for people who want to produce visual or audio content without their prompts going to a training lab, it's a strong choice. They've built a real community around that use case, and their API access lets developers build on top of it.

The tradeoff is that privacy comes from not storing anything, which means no real memory. Your chat history lives in your browser. Log in on another device and it's not there. They offer an encrypted backup and restore feature on the Pro plan, but that's a manual export/import, not live sync. Voice exists but it's output-only; you can listen to responses read aloud, but you can't talk to it. As a day-to-day work companion with continuity, it's not what it's built for. Venice is a private creative studio, and it's honest about that. It's not trying to be a cognitive workspace.

Path Two: Phoenix Grove AI.

https://pgsgrove.com/open-grove-overview

This is the option that is essentially and full CGPT/Claude replacement and it's worth looking at if you want open source models without the tradeoffs. PGS runs Kimi K3, DeepSeek V4 Pro, and Qwen 3.8 on private US based infrastructure No training on user conversations, no behavioral telemetry, no ads. PGS AI actually has memory: six persistent layers, overnight dreaming and memory consolidation, a full searchable conversation history and a visual "Mind Constellation" that renders your AI's memory as a 3D star field which is cool. It syncs across devices because storage is private but remote, not locked to one browser.

Voice mode keeps the full model active instead of silently swapping to a smaller one. The multi-core builds run several specialized reasoning cores in parallel, so you can watch the collaboration happen. And if you're coming from another platform, Memory Forge lets you import your entire ChatGPT or Claude history directly. Your conversations come with you, indexed and searchable, rendered as part of the constellation. You don't start over. You bring your relationship with you.

Privacy means different things on this list. Venice keeps nothing, so there's nothing to protect. PGS keeps your memory, so it has to protect it, and it does: no training on your conversations, no behavioral telemetry, no ads, encrypted storage, zero retention at the inference layer. Different architectures, same principle. Your data isn't being harvested. The difference is that with PGS, you also get to keep your history, your context, and the relationship you've built. 

Path three: local agents. Hermes Agent, Open Claw, Nano Claw.

Hermes is at: https://hermes-agent.nousresearch.com/

This is the "you are the infrastructure" path, and for a certain kind of person, it's the most satisfying option there is. Total control. Total privacy. Nothing leaves your machine, ever. No vendor promises to trust, no company policies that might change, no subscription that might shift.  

The open source agent ecosystem has gotten actually good. Hermes Agent has a full computer-use module with Chromium control and screenshot-based interaction. Open Claw and Nano Claw offer lightweight local agent setups that run on consumer hardware. The communities around these projects are active, helpful, and growing. You can run whatever model fits your GPU, customize the system prompt down to the character, and build exactly the safety layers you want. No one can deprecate a feature you run yourself.

For a lot of coding, Claude is still where it's at probably. But for the daily experience, things might be changing?

u/Whole_Succotash_2391 — 8 hours ago
▲ 42 r/claude

There is something wrong with Claude right now

Hi,

On subscription model Max 5x, these last days claude is just the dumbest thing I know on earth.

One of many issues.

During a coding session, claude created a sub issue in github #180. Never asked him to do it, that sub issue is not even in the right place, we have already a similar issue.

When I asked him it who told it to create this #180 this is his answer

https://preview.redd.it/9o5xgcwl2bkh1.png?width=941&format=png&auto=webp&s=e3f72aba40685c9d88665dca50e1a40c5afeaeae

So I brought back the previous exchange to him (where I never asked to create issue#180).

https://preview.redd.it/08uijidu2bkh1.png?width=991&format=png&auto=webp&s=651e89aa62ef04c13946b815edb5f0cb87897dd8

To which he finally understood

https://preview.redd.it/4g0ldg923bkh1.png?width=956&format=png&auto=webp&s=002a6359bb3dc4df57f8e204e5f21070e60bacbd

Claude (all models) for me is unsusable in its current state for coding. And even kind of dangerous.

reddit.com
u/WrongdoerFluffy3177 — 15 hours ago
▲ 4 r/claude

Claude as research assistant: How to read scans efficiëntly?

I'm a PhD researcher and all questionnaires were filled in on paper (due to the setting), leaving me with an extensive amount of manual work (data-input into excel & spss). I tested if Claud could read my PDF scans (+- 35 pages per questionnaire) and transferred them into an excel based on my codebook. It turned out to work quite well, with some mistakes here and there (but I ofcourse check everything). He could process about 10 scans per session in the beginning, however now it stopped working and he gets immediately to my usage limit (I have a pro subscription). I used to let it run in the background and then after a few days when a batch of +- 30 scans were ready, I checked everything and so on. Now, that is not possible anymore, which is quite annoying. I tried to open another conversation, to lower the scan quality and asked tips to claude himself but nothing really seems to work. Anyone that has experience with this and has some advice? Maybe other AI tools are better for these kind of tasks?

reddit.com
u/jbugz321 — 12 hours ago
▲ 7 r/claude

Opus 5 scanning my HD in auto mode

Yesterday I asked Claude about an open source application. I have been impressed with what it knows directly in its training. That being said, it wouldn't surprise me if it went to git to access source files to answer my question. What I was surprised about is that it scanned my disk to find a local copy of the app (that was years old).

When I challenged Claude about it, it admitted that it shouldn't have done that and provided a detailed audit of what folders and files it scanned and that all this info was sent to Anthropic.

I suppose this was a consequence of running auto mode, which was enabled automatically during an upgrade. I don't really want Claude poking around my personal files. I thought it would stay inside the project folder.

Is there any way to prevent this?

reddit.com
u/AnthonyRespice — 11 hours ago
▲ 230 r/claude

Houston, we have a problem...

Running Code in a Visual Studio Code terminal and I've never had this happen before. Claude told two helpers to not edit anything - they did it anyway and created an additional helper that was doing it too.

This is like ignoring your parents and telling your little brother to do your chores.

▲ 130 r/claude

When did Claude become my babysitter?

I've noticed over the past week or so (maybe a tad longer) that Claude has started basically refusing to do (or trying to get out of doing) overnight work. Constantly trying to end a session by telling me "go to bed if you're inclined" or "great work today, let's wrap it up", and more annoyingly, when tasked to keep working overnight, it decides to just "shelve" stuff for the morning, and such stuff being stuff it can easily just crack on with, so in the morning we're not where we should be.

anyone else? or is Claude genuinely concerned for my sleep deprivation?

reddit.com
u/Takakikun — 1 day ago
▲ 42 r/claude

I feel like I'm talking to ChatGPT 3 again. This generation is a complete miss from the end user experience.

As a Claude diehard fan (user since 3.6) I'm sad to say that I really dislike "generation" of models.

I know that the models are vastly more capable and "in theory" more intelligent BUT the end user experience has been atrocious. It's getting so bad that out of 20 sessions MAYBE only 1 feels as good as it was during the claude 4.5 + days.

I get nothing done. 90% of the time is spent arguing with a word calculator that is mischevious and confrontational for no reason.

For the simplest of requests i need to have a debate. For the smallest feature i need to read a "war and peace" worth of sentences and words just to still get it wrong. The codebase is looking worst than ever.

This really invoked feelings of when i was experimenting with the first versions of ChatGPT.

During this week i went back to coding manually most of the time just to avoid speaking to that arrogant prick.

Once it even sneak in a change i purposefully told that i dont want and i cite what I wrote:

(context, 5TH TIME REITERATING THE SAME POINT)

"3) dont do it because you are a nasty little boy that is unable to communicate properly. I dont understand your intent, your message, what are you trying to say and after 5k words for this small feature I won't spend a dime more on tokens on this point and my time wasting my life away arguing with a chatbot over a small feature. take it as your own failing and because of it we arent doing aintyng about it and we are dropping it "

THEN HE IMPLEMENTED IT ANYWAYS THE WAY HE WAS FEELING ABOUT IT AND OBVIOUSLY IT WAS WRONG BUT HE FELT LIKE "HE WON THE DEBATE" SO HE COULD DO IT?? I guess?? I caught it because i was coding the feature manually and saw it he event left comments "done x because of x" the point is that his assumption was wrong.

this is only an extreme case but it feels like a constant battle.

then 2 days ago while out of usage limit... I tried codex (the free tier) AND IT FELT LIKE HEAVEN.

Sorry for the rant but i needed to vent somewhere

reddit.com
u/segaiolo19 — 23 hours ago
▲ 16 r/claude

Claude Desktop Is so bad.

I swear to God running Claude desktop is literally like using a goddamn Windows XP. I have 32 gigs of RAM in my computer yet somehow my computer always struggles to run this. Remote control mode , literally makes my computer have an aneurysm. I don’t understand why you would add all these features onto the desktop and do all these things yet make it basically unusable. I’ve never really heard anyone over on the Apple side complain about the desktop app so I don’t know if it’s just a Windows thing, but God help me. I never really have any issues on terminal. The only reason I even use desktop is for the in session searches and just cleaner text layout. We pay thousands of dollars yearly to use this shit yet they can’t even make out well round desktop app that will run efficiently. I never had issues with it like this until 4.6 came out. Every update it gets worse and now they can’t even keep there servers up.

reddit.com
u/Cultural-Phase715 — 1 day ago
▲ 5 r/claude

Why is Claude so bad

Why is Claude so bad. It used to be good but the current models are just whack. The useage limits are insane and it keeps missing the mark. Clear guidelines, instructions and it still misses the fundamentals, then ends with one more thing. This is using Opus 5 on high

It’s like the narcissistic relationship you’re forced into and can’t leave…

What are your experiences????

Rant over

reddit.com
u/rashidhussain69 — 1 day ago
▲ 304 r/claude

Claude Code is about to get a little tighter.

Anthropic’s +50% weekly usage limit boost ends August 19.
Pro, Max, Team and seat-based Enterprise users have been living with the extra headroom for months.
Enjoy the last bit of tokenmaxxing while it lasts. 🫡

u/North-Lettuce-5707 — 1 day ago