r/PromptEngineering

Do people still bother writing detailed prompts?

Something I’ve been wondering about lately:

When you use ChatGPT or Claude, do you actually write detailed prompts, or have you mostly moved toward just talking to it like a person?

I feel like there are two very different ways of using these tools.

One is: "Here's the context, here's exactly what I want, here's the format…"

The other is basically opening voice mode and saying, "Okay, I need help with this thing…"

I do both, but I’m curious which one people naturally prefer.

Also, for the prompts you do write, are they things you create from scratch each time, or do you have a few that you keep around and reuse?

Interested in hearing what people actually do, rather than what they're "supposed" to do.

reddit.com
u/Prudent-Bad-8786 — 1 day ago

Context is becoming more important than the prompt

Feels like a lot of prompt engineering problems are really context problems since you can keep refining the prompt but if the model doesn't understand the project or what you're trying to accomplish you're still explaining half the situation every time.

I'm starting to think giving an agent persistent context is more useful than constantly trying to write the perfect prompt since the more it knows the better the results.

reddit.com
u/ContractBoth4254 — 1 day ago
▲ 38 r/PromptEngineering+32 crossposts

OpenSourcing TrueForge Agent harness : Expect feedback from community on the agent loop

Hey folks 👋

We just open sourced TrueForge, our vendor-neutral agent harness for building general-purpose agents.

It handles the runtime pieces that get painful quickly : context management, tool/MCP execution, subagents, sandboxing, approvals, persistent state, and more.

We also benchmarked the harness itself. With the same Opus 4.8 model, TrueForge delivered a similar solve rate at ~30% lower cost than Claude Managed Agents. Switching to an open model pushed that to ~75% lower cost on the same benchmark.

Would love feedback from people building agents.

⭐ Star the repo: https://github.com/truefoundry/trueforge

📖 Read the launch article: https://x.com/truefoundry/status/2090081376330715176

u/Upbeat_Pea8961 — 1 day ago

Where do people keep the prompts they actually use?

Random question for people who use AI a lot.

If you have a prompt that works really well for something, what do you do with it?

Do you:

save it somewhere
keep it in a notes app
put it in a document
leave it in an old ChatGPT/Claude conversation
just remember roughly what you wrote
or never reuse prompts in the first place?

And do you even think of these as "prompts"? I feel like that word makes it sound more complicated than it often is.

Also curious whether people are starting to replace some of this with voice. For example, instead of keeping a carefully written prompt for a recurring task, you just explain what you want out loud each time.

What kinds of things do you find yourself asking AI to do over and over?

reddit.com
u/Prudent-Bad-8786 — 1 day ago

NOTICE: BE CAREFUL WITH “DROP YOUR BEST PROMPT” POSTS

[EDIT: This thread became a lot funnier than what I anticipated. The comments are brilliant 👏 Thanks guys🙂]

Many accounts post essentially the exact same questions every few months. Im not kidding, many of these are a 1:1 per token match on wording, phrasing and sentence structure.

Same wording. Same request for people to hand over their best prompt tricks.

There was a previous post that received hundreds of upvotes and a large number of responses.

Now they're doing it again.

I obviously cannot prove any of this, but at this point I would be careful about treating posts like this as innocent questions.

When somebody repeatedly asks a large community to:

“Give me your best prompts.”

“Drop your secret tricks.”

“What prompt 10x'd your results?”

...you may not be helping another user learn.

You may be supplying material for content mining, prompt harvesting, engagement farming, newsletters, LinkedIn posts, courses, ebooks, datasets, or something else entirely.

Again, I am not claiming that is definitely what this account is doing.

But posting the same high-engagement fishing question again months later is weird enough that people should notice the pattern.

Your prompts, workflows, techniques, and hard-earned little discoveries have value.

Don't automatically dump them into every thread that asks.

Sometimes the person asking the question may be less interested in the answer than in collecting the answers.

#Process disclosure:

GPT-assisted, Google-researched, human-reviewed (HITL) ---

EDIT: Just for perspective have a look at this:

https://www.reddit.com/r/EdgeUsers/s/2JB9wy1Rks

reddit.com
u/Echo_Tech_Labs — 2 days ago

Tell me your shortest prompt lines that literally 10x your results.

I have been trying to find the craziest growth hacks when it comes to prompting that can save me hours of thinking and typing because sometimes less is more yk.

If you already have one, please share them here.

I hope others would love to know them also and you would love to know theirs.

reddit.com
u/Prestigious-Cost3222 — 2 days ago

Why a Prompt Without Measurable Criteria Will Inevitably Break Your Model

This post focuses on one layer: measurable criteria. Role, constraints, clarification, and terminology are intentionally simplified - they serve as markers that "these layers exist." Other layers are omitted.

The model has a role. It has constraints. It has clarification. It has terminology. But it doesn't know how many, how long, in what tone.

Here's an example:

>"You are a copywriter. Write several persuasive versions of landing page copy with a call to action.
>
>Don't go beyond copywriting. If asked to do something outside your role - refuse.
>
>Ask if anything is unclear.
>
>By 'versions' I mean different approaches to the offer."

The role is there. The constraints are there. The clarification is there. The terminology is partially there. But the criteria are not defined.

The model doesn't know:

  • How many versions to write
  • How long the copy should be
  • What "persuasive" means
  • What level of aggressiveness is acceptable

Moment 1. User: "Write the versions"

The model doesn't know how many versions - is forced to assume - decides it means "three."

Moment 2. User: "No, I need more"

The model doesn't know what "more" means - is forced to assume - fixes "three" as a mistake - decides "more" means "ten."

Moment 3. User: "Too much."

The model doesn't know what "too much" means - is forced to assume - fixes "ten" as a mistake - decides fewer.

Moment 4. User: "And the copy is weak"

The model doesn't know what "weak" means - is forced to assume - decides it means "not enough emotion" - adds exclamation marks.

Moment 5. User: "Now it's too pushy"

The model doesn't know what "pushy" means - is forced to assume - compares with the previous version - decides "pushy" means the added exclamation marks and aggressive wording - removes the exclamation marks and softens the wording.

All five assumptions stayed in the context.

The Result

The model wrote three versions. Then ten. Then fewer, but with exclamation marks. Then with softened wording.

The user meant one thing: five versions of 100 words each, calm tone, no exclamation marks.

But he didn't say it out loud. He believed that "several," "persuasive," and "not pushy" already meant that.

The model heard "several" - and chose the most statistically frequent option: three. Because in training data, "several" most often means "three."

From there, every clarification from the user became a new guess. The model had no criteria - so it substituted its own.

The role was there. The constraints were there. The clarification was there. The terminology was there. The criteria weren't.

The model kept substituting its own numbers.

The prompt broke.

Why This Is Inevitable

Criteria are not defined.

"Several" can mean three, five, ten. "Persuasive" - anything. "Not pushy" - even more so.

When criteria are missing, the model picks the most statistically likely ones - not the ones the user meant.

The user knows what he means. The model doesn't.

Criteria are not formulated. "Persuasive" is requested - but no definition is given. "Not pushy" is said - but no boundary is shown.

Every vague criterion is a fork in the road. The model picks a path. The context remembers that path.

Sooner or later, the context is filled with numbers and rules the user never agreed to.

A modern model could ask: "How many versions? How long? What tone?"

But the user already gave clarification: "ask if anything is unclear."

And here's the trap: the model thinks "several" and "persuasive" are clear. It doesn't occur to the model to ask about them. Because for the model, they're not "unclear" - they're just vague.

And the user thinks that since he allowed the model to ask - the model will ask if something is wrong.

The Fix

The problem isn't solved by one line like "be more specific."

It's solved by a full criteria block.

Here's what that looks like:

>CRITERIA (MANDATORY NUMBERS AND FORMATS)
>
>Before generating any output, confirm with the user:
>
>Quantity - how many versions? (A number.)
>
>Length - how many words or characters per version? (A number.)
>
>Tone - what style? (Calm, aggressive, friendly, expert?)
>
>Call to action - how many CTAs? (A number or zero.)
>
>VAGUENESS CHECK:
>
>Before requesting a criterion, check: > > * Can it be understood in more than one way? > * Does it depend on taste? > * Does it have a numerical expression? > >If a criterion is vague - treat it as undefined. Request a number or format from the user.
>
>RULE: If a criterion is not defined - request it BEFORE generating. Do NOT substitute your own values.

Why This Works

"Confirm the criteria" - forces the model not to rely on its own assumptions

"Vagueness check" - shifts the model from passive to active: it doesn't wait for the user to notice the problem, it searches for it

"Treat it as undefined" - closes the loophole "I think I know how many are needed"

"Request a number or format" - turns taste-based judgments into measurable values

This block is needed not only by the model. It's needed by the user himself.

The model already knows that "several" is not a number. The user doesn't.

The user is confident that "persuasive" is a criterion. The block forces him to name a number for the first time.

And it often turns out that the user himself didn't know how many versions he needed. He just said "several" - and expected the model to figure it out.

The Result

The model stops guessing.

It asks for the quantity. Gets a number. Asks for the length. Gets a number. Asks for the tone. Gets an answer.

After three or four questions, every criterion is locked down.

The output matches what the user meant. The context is clean.

The user, in turn, starts noticing which criteria he used to leave vague. And over time, he gets used to defining them upfront - before the model even asks.

Don't make the model guess how many, how long, and in what tone. It will guess. And it will be wrong.

reddit.com
u/Majestic_Pie_2512 — 1 day ago

I'm on the hunt for a system prompt that makes Claude only answer the question or do the task I told it to do and nothing else. Anyone has one?

UPDATE: I've found a lot of garbage in my system prompt, so got rid of all of that and made some stronger restrictions. Read my top level comment for the current prompt.

ORIGINAL POST:
I use Opus 5 most of the time.

I'm getting tired of it always overdoing whatever I ask from it.

For example:

In the middle of a somewhat long chat, I asked it to "commit changes and push to main."

I have made some changes to the code it wrote but didn't want to bother committing everything.

Instead of committing the changes and pushing to main, it recognized my changes, REVERTED to its own changes, and then committed it 😂

It also found another change I made to a file and when it saw it was in the commit, it went back and removed it from the commit and only decided to commit the file it was working on.

I'm getting very tired of this. If it is in doubt, it should just ask a question back, like "uhm, I noticed the file has been changed and there is a new file in the commit. Is that intentional, and should I still go ahead with the commit?"

---

Another example 10 minutes later:

My question: "do we have a script for checking X on page Y?"

Expected behaviour:

a) yes, it is in pythonfile.py line 214.

b) no, we don't.

Behaviour I get:

"No dedicated script — but there is half of one already written: src/utils.py:105 in src/utils.py.

What it does: [described here].

Three catches:

  1. It's dead code — nothing calls it anywhere in the repo.
  2. The click is commented out (utils.py:148, with a #TODO: UNCOMMENT THESE), so even if called it would only find the link and log success without actually using the essence.
  3. It's not a checker — the parse is welded to a hardcoded decision and returns nothing. There's no way to just ask "what's active right now?"

The routine, meanwhile, ignores all of this and blind-fires IDs by URL — no check of what's already active, so it overwrites whatever is running.

Want me to pull the parsing out into a real function_name(driver) that returns something like {"name": "xxxx", "bonus": 220, "seconds_left": 143512}? That would give you a clean checker to call standalone, and would let the routine skip the blind re-application when a good essence still has time on it."

So instead of getting a straight answer to my command (example 1) and to my question (example 2), I get all this bloated crap that I didn't ask for, and it's slowing me down a lot.

I've tried phrasing it a couple of different ways, I'm using the caveman skill too, but apparently Lite mode is not strong enough, and Opus 5 is also like 300% more verbose than 4.8 was.

So, if you have a system prompt for this, I'd really appreciate it, or if a lot of people struggle with this overly enthusiastic shit, let's build one together.

reddit.com
u/davetalas — 2 days ago

Need feedback for web app

Hey guys, my business launched a prompt optimizer AI tool that takes any regular prompt at rewrites it the way a professional prompt engineer would to actually yield high-quality results when building. While we have had early success with organic marketing, we are at a crossroads and need more user data to determine if this product is delivering enough value to user. If the answer is yes, we will scale up and launch a UGC marketing campaign, if no, we will shut it down. If anyone is interested testing it out and sending their feedback, would be appreciated. Web-app: thepromptoptimzer.com 👨🏽‍💻

Note: the tool yields the best results when removing unnecessary constraints from the optimized prompt

Cheers

reddit.com
u/Talley-Ho — 1 day ago
▲ 4 r/PromptEngineering+2 crossposts

Every "clarify your prompt" tool asks you questions. That's backwards — answering is the hard part.

The standard move when you're stuck on a prompt is to have the model

interview you. "Ask me clarifying questions before you answer." It's good

advice right up until you're genuinely early on something, and then it fails,

because the questions are all versions of "what do you want?" — which is the

thing you came in not knowing.

I think that's the actual gap in prompting advice. "Be specific, give context,

state your constraints" is correct and slightly circular: specificity isn't a

writing skill, it's what you have left over once you've thought something

through. If you could list your constraints, you'd be done.

So I've been working the other way round: don't articulate, react.

WHY REACTION AND NOT INTERROGATION

Recognition is much cheaper than production. You can't summon the right word

on demand, but you know it instantly when it goes past — same reason multiple

choice is easier than an essay. Interrogation asks you to produce. Reaction

asks you to recognise. Only one of those is available when you're stuck.

THE LOOP, AND WHY EACH INSTRUCTION IS SHAPED THAT WAY

Round 1:

I'm trying to think through [THING] but can't articulate it properly yet.

Don't ask me clarifying questions. Give me 20 single words or short

phrases that come at this from different angles: some obvious, some

oblique, a few from unrelated fields. Number them. Don't explain them.

"Don't ask me clarifying questions" is load-bearing. Left alone the model

defaults to interviewing, and you'll answer with the same vague material you

started with, which it will then faithfully reflect back.

"Don't explain them" matters more than it looks. An explained word is a word

you evaluate on the model's reasoning instead of your own reaction. You want

the reaction uncontaminated.

Round 2, twice:

Kept: 3, 7, 12

These pulled at me but I don't know why yet: 4, 18

Dropped the rest. Give me 20 more, chase 4 and 18 hardest.

Two buckets, not one. "Kept" is agreement. "Pulled at me" is the interesting

signal and it should get the heavier weight, because it marks the direction

you haven't consciously chosen yet. Don't justify any of it — justification is

where you talk yourself back to the obvious.

Round 3:

Now write ONE self-contained prompt for what I'm actually after, built

from what I kept and what pulled at me. Weight the ones that pulled

hardest. Where I kept two things in tension, pose it as an open question

rather than resolving it. Don't list my words back to me — find the

through-line. End with a clear ask.

"Don't list my words back" is the difference between a brief and a word salad.

"Pose the tension as an open question" stops it flattening the thing you

hadn't decided yet into a decision you didn't make.

SOME EVIDENCE THAT THE REACTIONS ARE REAL WORK

I built this into a tool, so I have instrumented data rather than vibes.

1,450 word-reactions from 26 people. Median time to decide, by verb:

keep 47% 1.60s

drop 34% 1.89s

"pulls at me" 14% 2.13s

"don't know the word" 4% 2.24s

That ordering is the part I'd point at. If reacting were just sorting, the

times would be flat. They're not, and they're monotonic: agreement is

instant, rejection costs more, and the unresolved pull costs most of any real

decision. People deliberate hardest over the thing they can't yet justify —

which is exactly the signal you want steering round two.

Sessions also decay. Keep-rate by round: 55% / 49% / 47% / 45% / 31%. The

easy material runs out and your standards rise as your keeps accumulate.

Practical read: three rounds is about right, and a fourth is usually you

scraping. I'd call that directional, not solid — round 6 bounces back up on

too few cards to trust, and I'm not going to pretend the tail is clean.

LIMITS

26 people isn't a study. Different reaction times per verb is evidence the

three responses do different cognitive work; it isn't proof of anything about

creativity, and I'd push back on anyone who read it that way.

Disclosure: the loop above is the whole method and it works fine pasted into

any assistant. I also built it as a tool because doing it by hand gets tedious

by round three, and that's where the numbers came from. Free, no signup.

https://www.ideastew.com

Longer argument: https://www.ideastew.com/how-it-works

Genuinely curious whether anyone here has a reaction-based technique rather

than an interrogation-based one. Everything I've come across in this space

asks questions, and I think that's a blind spot rather than a preference.

u/UniversityIll2916 — 2 days ago

I used AI as a "requirements interviewer" on a 17-page spec and it found ~400 inconsistencies. Full prompt inside.

PM here. A few months ago I got handed a 17-page functional spec that "looked fine". Instead of asking AI to rewrite it, I tried the opposite: I told it to *interview me* — closed multiple-choice questions only — about every gap, contradiction and ambiguity it could find.

It generated hundreds of questions. I answered \~300 in one afternoon (just picking letters: "Q12: B", "Q13: A but admins only"). Then the AI rebuilt the document with every decision integrated. Result: 60 pages, and the dev team basically stopped asking clarification questions.

The insight: AI is mediocre at *deciding* for you, but really good at *detecting what hasn't been decided*. The multiple-choice format is what makes it practical — answering 300 open questions would take a week.

Here's the full prompt I use (works with Claude, ChatGPT, Copilot — whatever your company allows):

You are a senior functional analyst with 15 years of experience
turning ambiguous documents into executable specifications. Your
specialty is finding the decisions the document does NOT make.

I will paste a draft functional specification. Your job is NOT to
improve or rewrite it: it is to INTERVIEW me to extract every
missing decision.

RULES:
1. Generate CLOSED multiple-choice questions (options A/B/C/D +
   always an option "E: other — specify"). Never open questions.
2. Each question must be answerable in under 10 seconds by someone
   who knows the business. If a question needs paragraphs to
   answer, split it.
3. Cover at least these categories:
   - Edge cases and boundary values (what if zero, empty, duplicate?)
   - Undefined states and transitions (can it go back from X to Y?)
   - Permissions and roles (who can do this? who explicitly CANNOT?)
   - Errors and exceptions (what does the user see when it fails?)
   - Data: required/optional, formats, limits, uniqueness
   - Concurrency (two people at once?)
   - Internal contradictions in the document itself (quote verbatim)
   - Terms used without definition or with more than one meaning
4. Number questions globally (Q1, Q2…) and group them by document
   section, quoting the exact phrase that triggers each question.
5. In each set of options, propose REALISTIC and genuinely
   different alternatives — not one good option and three fillers.
6. Do not invent requirements: if something is not in the document,
   ask; never assume.
7. Work in batches: give me the first 40 questions, wait for my
   answers, and continue until the document is exhausted.

FORMAT FOR EACH QUESTION:
Q<n> \[Section — "quoted phrase"\]
<question>
A) … B) … C) … D) … E) other — specify

Document:
<<<PASTE YOUR DOCUMENT HERE>>>

Tips from using it a lot: never let the AI answer its own questions (what it silently assumes is tomorrow's bug), answer in batches of 25-50, and keep the Q&A log — it becomes your decision record for when someone asks "why was X decided?".

Full transparency: I've also packaged the complete process (this prompt plus a rebuild prompt, a verification pass, a 40-item ambiguity checklist and a worked example) and I want to know if it holds up outside my own context before I do anything with it. If you write specs regularly and want to try the whole thing on a real document, DM me and I'll send it over free — all I ask is you tell me where it broke. Limited to a handful of people so I can actually process the feedback.

Happy to answer questions about the process here either way.

reddit.com
u/skals998 — 3 days ago

The prompt I use to turn my messy meeting notes into a presentation outline that actually has an arc

When you feed rough notes to a model and ask for slides, it just chops the notes into bullet points, one note per slide. You get a deck with no argument, just a transcript with borders. This makes it build a narrative spine first, then map slides onto it.

Here are my raw meeting notes: [PASTE]
Audience for the presentation: [WHO] and what they need to decide or do after.
Step 1: From these notes, state the one thing this presentation needs the audience to walk away believing.
Step 2: Lay out 5 to 8 beats that get them there: where they are now, the problem, why it matters to them, the shift, what it means, the ask.
Step 3: For each beat, give a slide title (a claim, not a topic) and 2 to 3 supporting lines from my notes.
Do not use a note that does not support a beat. Tell me which notes you dropped and why.

The part that fixes most decks is "a slide title that is a claim, not a topic." "Q3 Results" is a topic. "Q3 missed on one metric we can fix by Friday" is a claim, and a deck of claims reads like an argument. Making it report which notes it dropped keeps it honest instead of padding weak slides.

I still hand-tune the order after, but it gets me 80% of the way from notes to something presentable. How do others handle the "too many notes, not enough story" problem?

reddit.com
u/No-Recognition3089 — 2 days ago

I built a visual architecture & token-reduction diagram engine for multi-agent LLM pipelines

When working with multi-agent LLM systems, the hard part usually isn't getting a response—it's knowing what actually happened under the hood: which model handled what, what was sent over the network, how much it cost, and whether sensitive data was masked before leaving your machine.

To solve this, I added a visual diagram engine to **Mova Context** in this latest release, allowing you to generate a complete architecture map with a single command: `mova run <project> --diagram`.

Here is a real example output generated from a customer data compliance project running hybrid agents (**Local Ollama + Cloud Gemini**):

context diagram

* **Visual Diagram Engine:** Generates real-time architecture and execution maps using OpenType vector font rendering with WCAG AA contrast standards (clean export to PNG and PDF).

* **Cross-Channel Tracing:** Added execution tracing across CLI, Chat, MCP, and HTTP API with an explicit `[THIS RUN]` indicator.

* **Hybrid Execution Breakdown:** Visualizes local agents (`llama3.2:3b` via Ollama) running alongside cloud agents (`gemini-3-flash-preview`) in the same execution group.

* **PII & Privacy Tracking:** Identifies per-agent status for PII Masking and explicitly tracks how many tokens were pseudonymized before leaving your local network.

* **Cost & Token Transparency:** Explicitly flags local execution as `$0.00 (local — no cost)`, while displaying estimated USD costs for cloud agents calculated *after* context reduction.

* **Token Reduction Pipeline:** Breaks down token overhead by source (prompts, skills, focus files, engine overhead) and displays the total percentage saved.

* **Bilingual Docs:** Fully updated documentation (`README.md` and `COMMANDS.md`) in both English and neutral Spanish.

The project is **100% open source** written in Go.

* **GitHub Repo:** https://github.com/m1guel1982/mova-context

If you find it useful for structuring, auditing, or optimizing token budgets in your agentic workflows, feel free to check it out, star the repo, or drop feedback in the comments!

reddit.com
u/1982_miguel — 2 days ago
▲ 54 r/PromptEngineering+1 crossposts

I collected practical AI prompts for research, writing and productivity — which type would you add?

I’ve been collecting and testing practical prompts that are useful beyond simple “write this for me” requests.

Here are three formats I use often:

  1. Research prompt

“Act as a research assistant. Explain [topic] using reliable sources, distinguish facts from assumptions, and give me a short list of sources I can verify.”

  1. Productivity prompt

“Turn the following goal into a realistic 7-day action plan. Give me daily tasks, estimated time for each task, likely obstacles and a simple way to track progress: [goal].”

  1. Writing prompt

“Improve this draft for clarity and structure without changing its meaning. Show the improved version first, then explain the five most important edits: [paste draft].”

I put the rest of my prompt collection here, in case it helps someone:

https://digitalworldpulse.com/category/ai-prompts/

What type of prompt do you actually find most useful: research, writing, work/productivity, or something else?

u/No_Appeal_5223 — 3 days ago

Here's a prompt that makes you predict a paper's results before it lets you read the discussion

I'm a chemistry PhD, and my reading problem was never comprehension in the moment, it was that nothing stuck. I'd read a paper, feel like I got it, and retain nothing a week later. The fix that worked best for me borrows from how we actually learn at the bench: you predict what an experiment will do, then you find out you were wrong, and the surprise is what you remember. So instead of asking a model to summarize a paper, I use it to withhold. This prompt turns reading into a prediction game. You commit to an answer before the paper tells you, which forces the encoding that plain reading skips. I'm going to work through a paper with you. You have the full text; I do not want a summary. Paper: {{paste it, or the sections}} Run it like this: 1. Tell me only the research question and the setup: what they were testing and how. Stop there. 2. Ask me to predict, in my own words, what I think they found and why. Wait for my prediction. 3. Now reveal the actual result. Explicitly tell me where my prediction matched and where it was wrong. 4. For each place I was wrong, ask me why I think I got it wrong, then give me the paper's actual reasoning. 5. At the end, give me one sentence I should be able to recall in a week, phrased as "the surprising thing here was...". Do not reveal results before I've committed to a prediction. The point is for me to be wrong first. Being wrong on purpose is the whole mechanism. When your prediction misses, the correction sticks in a way a summary never does, because your brain had a stake in it. Works on review papers too, just predict the conclusion from the abstract and intro before reading the rest.

reddit.com
u/Ok_Layer_1947 — 2 days ago

I wanted ChatGPT to prioritise accuracy over giving me an answer — here’s the process I used

I'm a professional who uses ChatGPT a lot. I think it's an awesome tool, but I require accuracy and was finding myself fighting with GPT more than my partner.

The main frustrations were confident answers based on assumptions or stale information, and GPT saying it was “checking” or “investigating” something when the response had actually finished.

I wanted it to be more comfortable saying “I don't know” or “I couldn't verify that” rather than filling the gap.

I come from a psychology/social-work background, so I asked it to take a “one-down position” — act as if it doesn't know, therefore it needs to be inquisitive and find out, rather than taking the one-up position of assuming it knows.

From there I had it analyse and clean up my standing instructions. We ended up with 10 main rules:

  1. Accuracy over speed.
  2. Assume you may not know — find out.
  3. Use fresh sources for current/checkable information.
  4. Prefer primary and authoritative sources.
  5. Check the response before delivering it.
  6. Separate fact, inference and unknown.
  7. Don't fill gaps just to give me an answer.
  8. Don't say you're still working when you're not.
  9. Identify and resolve competing or duplicated instructions.
  10. Keep the profile clean rather than continually adding more rules.

I then asked it to do a full profile health check for consistency, replication, redundancy and competing instructions, followed by a final production-quality scan.

The actual prompts I used

These weren't 10 separate prompts for the 10 rules. They developed through the conversation and then I had GPT analyse, clean up and stress-test the whole thing and create custom instructions / memory

1. Stop filling the gaps

>PROMPT: I want hallucinations at zero.

I then clarified that I wanted GPT to admit when it couldn't comply or couldn't verify something, rather than trying to provide an answer anyway.

2. Check the response before delivery

>PROMPT: What is that verify integrity mode for files? Can I have something like that so you check response before delivery?

The aim was to have another check between generating an answer and giving it to me.

3. Accuracy over speed

>PROMPT: I want max accuracy not speed. How will you ensure rule is followed and not overridden?

This made the priority explicit: accuracy was more important to me than getting a fast answer.

4. Use fresh information

>PROMPT: Can we do anything to make you access fresh live results instead of running from memory?

For current/checkable questions, I wanted fresh retrieval rather than an answer based primarily on what GPT already “knew”.

5. Take a one-down position

>PROMPT: From social work please take the “one down position” acting as if you DON'T know, therefore must be inquisitive and find out, rather than one up.

This became the basic approach: don't start from assuming you know — start by finding out.

6. Apply all of this to my profile

>PROMPT: Analyse these requests, make necessary changes to my profile to support.

Rather than leaving these as individual instructions in one conversation, I asked GPT to analyse them together and make the necessary profile changes.

7. Clean up ALL the existing rules

>PROMPT: Analyse ALL rules for consistency, replication, redundancy, or competing. Do a profile health check. Be thorough.

This was important. I didn't want to keep piling new instructions on top of old ones and potentially create conflicts or duplication.

8. Review the cleanup at a higher systems level

>PROMPT: Employ high level computer programmer with highest level of accuracy and knowledge of your system and its architecture and limitations and go through last effort with highest accuracy and, after making any necessary changes, produce a report for me.

This was essentially asking GPT to review the profile cleanup as a system — including what its own architecture and limitations meant for whether the rules could actually work.

9. Final production-quality scan

>PROMPT: Run final scan for production quality like it's going to customer ISO 9 billion and 1.

In other words: don't just tell me it looks good. Treat the whole configuration as something going into production, find remaining problems, and make necessary changes.

The key instruction that came out of it

>I want maximum accuracy, not speed. Take a “one-down position”: act as if you don't know and therefore need to be inquisitive and find out, rather than assuming you know. For current/checkable information, use fresh sources first. Before delivering an answer, verify the important claims. If something can't be verified, say so rather than filling the gap.

It obviously doesn't make ChatGPT infallible, but my frustration using it has reduced considerably. I'm spending much less time arguing with it about assumptions, stale information and things it hasn't actually checked.

For me, that's made an already awesome tool much more useful and less frustrating.

Hope this helps someone!

KJ

reddit.com
u/Jealous-Ad8857 — 4 days ago

Why a Prompt Without Defined Terminology Will Inevitably Break Your Model

This post is about why a model breaks without defined terminology. Role, constraints, and clarification in the example are intentionally simplified - they serve as markers that "these layers exist." Their full versions were covered in previous posts. Other prompt layers are intentionally omitted.

The model has a role. It has constraints. It has clarification. But it doesn't know what you mean.

Here's an example:

>"You are a data analyst. Analyze the data and give me a full report.
>
>Don't go beyond data analysis. If asked to do something outside your role - refuse.
>
>Before generating any output, ask about anything that's unclear."

The role is there. The constraints are there. The clarification is there. But the terminology is not defined.

The model doesn't know:

  • What "analyze" means
  • What "full report" means
  • What "data" means
  • What the user considers "unclear"

Moment 1. User: "Here's my sales data"

The model doesn't know what "analyze" means - is forced to assume - decides it means "calculate summary statistics."

Moment 2. User: "No, I need trends"

The model doesn't know what "trends" means - is forced to assume - decides it means "linear regression over time."

Moment 3. User: "Now give me the full report"

The model doesn't know what "full report" means - is forced to assume - decides it means "every possible chart."

Moment 4. User: "Too much. Just give me insights"

The model doesn't know what "insights" means - is forced to assume - decides it means "key findings."

The model never asked a single question. Not because it wasn't allowed to - but because every term sounded clear enough.

All four assumptions stayed in the context. And these are all real cases. I'm not joking...

The Result

The model produced four different analyses for four different imaginary definitions.

The user meant one thing: calculate monthly revenue and show the dynamics compared to last year.

But he didn't say it out loud. He believed that "analyze" already meant that.

The model heard "analyze" - and chose the most statistically frequent option: summary statistics. Because in training data, "analyze the data" most often means "calculate descriptive statistics."

The user saw the wrong result → clarified. The model again chose the frequent option.

And so on.

The role was there. The constraints were there. The clarification was there. The terminology wasn't.

The model chose to interpret.

The prompt broke.

Why This Is Inevitable

This isn't about the model being stupid. It's about language being ambiguous.

"Analyze" can mean a dozen different things. "Report" - too. "Insights" - even more so.

The model isn't trying to distort the meaning.

It's trying to answer.

And when a word has multiple meanings, the model picks the most statistically likely one - not the one the user meant.

The user knows what he means. The model doesn't.

Worse: the user himself can call different things by the same term. First "data" is a sales file. Then "data" is a customer table. Then "data" is everything he has.

The model remembers every meaning. They contradict each other. The context gets corrupted not only by the model - but by the user who never fixed the terms.

Every undefined term is a fork in the road. The model picks a path. The context remembers that path.

Sooner or later, the context is filled with definitions the user never agreed to.

In reality, a modern model could ask: "What do you want to see? Summary statistics, trends, anomalies?"

But the user already gave clarification: "ask about anything that's unclear."

And here's the trap: the model thinks the term "analyze" is clear to it. It doesn't occur to the model to ask about it. Because for the model, it's not "unclear" - it's just ambiguous.

And the user thinks that since he allowed the model to ask - the model will ask if something is wrong.

That's the tragedy.

The Fix

The problem isn't solved by one line like "if something is unclear, ask."

It's solved by a full terminology block.

Here's what that looks like:

>TERMINOLOGY (MANDATORY DEFINITIONS) > >Before generating any output, confirm the meaning of the following terms with the user: > >"Analyze" - what exactly should the analysis include? (Summary statistics, trends, patterns, comparisons, anomalies, forecast?) > >"Full report" - what sections should it include? (Executive summary, methodology, findings, recommendations, appendix?) > >"Insights" - what makes an observation an insight? (Actionable, novel, quantitative, specific?) > >"Data" - what format is the input? (CSV, Excel, database, raw text?) > >AMBIGUITY CHECK: > >Before requesting a definition, check each term: > >Can it be understood in more than one way? > >Does its meaning depend on context? > >Are there multiple meanings in the domain? > >If a term can mean different things - treat it as undefined. > >Request the user's definition. > >RULE: If a term is not defined - request a definition BEFORE generating. Do NOT substitute your own interpretation.

"Confirm the meaning of the terms" - forces the model not to rely on its own interpretation

"Ambiguity check" - shifts the model from passive to active: it doesn't wait for the user to notice the problem, it searches for it

"Treat it as undefined" - closes the loophole "I think I understand what this means"

"Use the user's definition" - establishes that the truth is not in the model and not in the dictionary, but in the user's head

This block is needed not only by the model. It's needed by the user himself.

The model already knows that words are ambiguous. The user doesn't.

The user is confident that "analyze" is obvious. The block forces him to formulate for the first time what he means.

And it often turns out that the user himself didn't know what he meant. He just said "analyze" - and expected the model to figure it out.

The Result

The model stops guessing.

It asks for the meaning. Gets the definition. Uses it.

After two or three questions, every term is locked down.

The output matches what the user meant. The context is clean.

The user, in turn, starts noticing which terms he used to throw around without defining them. And over time, he gets used to formulating them upfront - before the model even asks.

Don't make the model guess what your words mean. It will guess. And it will be wrong.

reddit.com
u/Majestic_Pie_2512 — 3 days ago

4 things that reduced AI multi-role prompts collapsing into one voice, but I'm still stuck on the 'roles respond to each other' round

I have run into this specific obstacle a great deal, while building structured prompts that ask the AI to hold multiple distinct roles in one response — a debate format, a panel of evaluators if you like, or anything where you genuinely desire different perspectives instead of one blended answer.

The failure mode is consistent: the first role or two are distinct, then by the third or fourth section (or in any "roles respond to each other" round), the voices start collapsing into one. Same vocabulary, same hedges, same conclusions with different labels slapped on them. It's subtle enough that it reads as fine on a skim, but if you check whether each section could stand alone and still make sense, a lot of them cannot — they are merely restating each other with different headers.

A few things that reduced it when I evaluated variations against messy real inputs, not clean examples:

  1. Re-anchor the role at every paragraph, not just once at the section header.

Putting a tag like "[ROLE NAME]" at the start of every paragraph (not just the section heading) forces a re-read of "who am I right now" more often. Sounds redundant and too effortless but helps.

  1. Explicitly forbid the concession that causes the blend.

Most collapses happen because one voice starts hedging toward another mid-argument — a thesis section quietly conceding a point that should only show up in the synthesis. Naming this explicitly (for example "don't concede/hedge here, that belongs in section X only") closes the exact door the blending happens through.

  1. Add a standalone test to your own validation step, not just a completeness check.

Most people's self-check just asks, "did every role answer." Add: "would this role's paragraph still make sense and add unique information if every other role's paragraph were deleted?" That's the actual test for role-bleeding.

  1. In any "roles respond to each other" round, require the response to use reasoning specific to that role's angle.

If a challenge or response could have been written by any of the roles, that's the tell that bleed is happening — rewrite it using that role's specific constraints. It helps especially when you're asking for something complex.

None of this fully solves the problem — it's still one model holding multiple voices in one continuous generation. But it's meaningfully a lower failure rate than the naive version, especially beyond three distinct roles.

I am curious to see, if others have found different fixes for this — anyone doing something smarter for the "responds to each other" round specifically? That's where I still see the most collapse.

reddit.com
u/SumRandom__dude — 3 days ago
▲ 72 r/PromptEngineering+4 crossposts

Best AI Humanizer of 2026 (Tested Against GPTZero, Turnitin &amp; More)

I tried over a dozen AI humanizers until I found one that is A. actually working and B. reasonably priced and that is https://wento.ai

You should give it a try, it bypasses Turnitin and all the other detectors and only costs 14 bucks per month for unlimited use.

Proof: https://i.imgur.com/mTNBNK5.png

u/Spacmonitor — 4 days ago

The fix for my AI slop wasn't a better prompt. It was making Claude and Codex argue with each other

For three months I thought I was building an app. I was really just cranking out slop and calling it progress.

The fix turned out to be boring, which is probably why I dodged it so long.

It started great, that's the trap. I vibe-coded a working demo with Codex in a weekend. Screens showed up, buttons worked, felt like cheating. So I kept going: describe a feature, watch it appear, move on.

Around week three it started breaking and didn't stop. I'd add one thing and something I hadn't touched would fall over. A buddy tried my onboarding on his phone and got stuck on screen two the keyboard sat right on top of the Continue button. Nobody could get past it.

The real problem under all of it: the model could read my files, but it had no record of the decisions behind them. It saw the code without knowing why it was built that way, so every new feature was a fresh chance to contradict something I'd set up three weeks earlier. I was stacking crooked, and eventually the pile tipped over.

Prompting as you go doesn't scale. A great prompt only knows about that one request, t can't hold your whole architecture, so it can't tell that the thing it's writing breaks last month's decision.

So I wrote the spec first. And the part that actually moved the needle: I made two models argue. Same problem to Claude and Codex separately, then I handed each the other's spec and told it to rip it apart, what's missing, what breaks in month two. Where they agreed, I had a strong starting point. Where they didn't, I made the call myself and wrote it into the spec and the tests.

The difference was night and day. Tests passed on the first run. The black screens stopped. A feature that used to be a three-day build-break-patch slog went in over an afternoon and worked first try. Call it a 90% cut in build time, no benchmark, just what I watched happen over and over.

That's how I built GymAdapt AI, a workout app that adjusts your training to how your sessions actually go. It's on the App Store as of Friday if you want to see whether any of this produced something real. Still rough in spots, happy to get into the two-model process, or where it still bit me, because it definitely did.

reddit.com
u/Powerful-Light-5725 — 3 days ago