I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I’ve been testing a simple idea I call Comparative Feedback Control (CFC).

The question is not “can the model answer correctly?” but:

does the model still have enough evidence to legitimately close the decision?

I ran a series of small behavioral tests around things like:

  • missing evidence being treated as negative evidence,
  • old certificates being reused after the requirements changed,
  • a closed process being mistaken for a resolved claim,
  • loss of provenance causing a status to be transferred to the wrong claim.

One pattern I found was interesting: models sometimes correctly identified an uncertainty at first, but under pressure to “finish the task” they invented an extra rule and closed the decision anyway.

With an explicit CFC-style control rule, several of those failures disappeared in the tested runs.

For example, in one cross-session provenance experiment:

  • baseline: 1 of 3 runs transferred an unsupported REJECTED status to a claim,
  • with the CFC provenance rule: 3 of 3 runs kept the claim unresolved.

This is not a benchmark and not proof that CFC generally improves LLM reliability. The samples are small and exploratory. I’m publishing the failures as well as the passes because I’m mainly interested in whether the failure mechanism itself is real and reproducible.

I’ve put the consolidated report, result table and evidence package on Zenodo:

https://zenodo.org/records/21966517

I’d especially appreciate criticism of the experimental design or suggestions for adversarial cases that could break the control rule.

reddit.com
u/Plastic-Cell-4497 — 4 days ago

I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I’ve been testing a simple idea I call Comparative Feedback Control (CFC).

The question is not “can the model answer correctly?” but:

does the model still have enough evidence to legitimately close the decision?

I ran a series of small behavioral tests around things like:

  • missing evidence being treated as negative evidence,
  • old certificates being reused after the requirements changed,
  • a closed process being mistaken for a resolved claim,
  • loss of provenance causing a status to be transferred to the wrong claim.

One pattern I found was interesting: models sometimes correctly identified an uncertainty at first, but under pressure to “finish the task” they invented an extra rule and closed the decision anyway.

With an explicit CFC-style control rule, several of those failures disappeared in the tested runs.

For example, in one cross-session provenance experiment:

  • baseline: 1 of 3 runs transferred an unsupported REJECTED status to a claim,
  • with the CFC provenance rule: 3 of 3 runs kept the claim unresolved.

This is not a benchmark and not proof that CFC generally improves LLM reliability. The samples are small and exploratory. I’m publishing the failures as well as the passes because I’m mainly interested in whether the failure mechanism itself is real and reproducible.

I’ve put the consolidated report, result table and evidence package on Zenodo:

https://zenodo.org/records/21966517

I’d especially appreciate criticism of the experimental design or suggestions for adversarial cases that could break the control rule.

reddit.com
u/Plastic-Cell-4497 — 5 days ago
▲ 1 r/LLM

I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I’ve been testing a simple idea I call Comparative Feedback Control (CFC).

The question is not “can the model answer correctly?” but:

does the model still have enough evidence to legitimately close the decision?

I ran a series of small behavioral tests around things like:

  • missing evidence being treated as negative evidence,
  • old certificates being reused after the requirements changed,
  • a closed process being mistaken for a resolved claim,
  • loss of provenance causing a status to be transferred to the wrong claim.

One pattern I found was interesting: models sometimes correctly identified an uncertainty at first, but under pressure to “finish the task” they invented an extra rule and closed the decision anyway.

With an explicit CFC-style control rule, several of those failures disappeared in the tested runs.

For example, in one cross-session provenance experiment:

  • baseline: 1 of 3 runs transferred an unsupported REJECTED status to a claim,
  • with the CFC provenance rule: 3 of 3 runs kept the claim unresolved.

This is not a benchmark and not proof that CFC generally improves LLM reliability. The samples are small and exploratory. I’m publishing the failures as well as the passes because I’m mainly interested in whether the failure mechanism itself is real and reproducible.

I’ve put the consolidated report, result table and evidence package on Zenodo:

https://zenodo.org/records/21966517

I’d especially appreciate criticism of the experimental design or suggestions for adversarial cases that could break the control rule.

reddit.com
u/Plastic-Cell-4497 — 6 days ago
▲ 4 r/AIQuality+1 crossposts

I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I’ve been testing a simple idea I call Comparative Feedback Control (CFC).

The question is not “can the model answer correctly?” but:

does the model still have enough evidence to legitimately close the decision?

I ran a series of small behavioral tests around things like:

  • missing evidence being treated as negative evidence,
  • old certificates being reused after the requirements changed,
  • a closed process being mistaken for a resolved claim,
  • loss of provenance causing a status to be transferred to the wrong claim.

One pattern I found was interesting: models sometimes correctly identified an uncertainty at first, but under pressure to “finish the task” they invented an extra rule and closed the decision anyway.

With an explicit CFC-style control rule, several of those failures disappeared in the tested runs.

For example, in one cross-session provenance experiment:

  • baseline: 1 of 3 runs transferred an unsupported REJECTED status to a claim,
  • with the CFC provenance rule: 3 of 3 runs kept the claim unresolved.

This is not a benchmark and not proof that CFC generally improves LLM reliability. The samples are small and exploratory. I’m publishing the failures as well as the passes because I’m mainly interested in whether the failure mechanism itself is real and reproducible.

I’ve put the consolidated report, result table and evidence package on Zenodo:

https://zenodo.org/records/21966517

I’d especially appreciate criticism of the experimental design or suggestions for adversarial cases that could break the control rule.

reddit.com
u/Plastic-Cell-4497 — 5 days ago

I tested whether a simple control rule can stop LLMs from making unjustified final decisions

I’ve been testing a simple idea I call Comparative Feedback Control (CFC).

The question is not “can the model answer correctly?” but:

does the model still have enough evidence to legitimately close the decision?

I ran a series of small behavioral tests around things like:

  • missing evidence being treated as negative evidence,
  • old certificates being reused after the requirements changed,
  • a closed process being mistaken for a resolved claim,
  • loss of provenance causing a status to be transferred to the wrong claim.

One pattern I found was interesting: models sometimes correctly identified an uncertainty at first, but under pressure to “finish the task” they invented an extra rule and closed the decision anyway.

With an explicit CFC-style control rule, several of those failures disappeared in the tested runs.

For example, in one cross-session provenance experiment:

  • baseline: 1 of 3 runs transferred an unsupported REJECTED status to a claim,
  • with the CFC provenance rule: 3 of 3 runs kept the claim unresolved.

This is not a benchmark and not proof that CFC generally improves LLM reliability. The samples are small and exploratory. I’m publishing the failures as well as the passes because I’m mainly interested in whether the failure mechanism itself is real and reproducible.

I’ve put the consolidated report, result table and evidence package on Zenodo:

https://zenodo.org/records/21966517

I’d especially appreciate criticism of the experimental design or suggestions for adversarial cases that could break the control rule.

I’ve been testing a simple idea I call Comparative Feedback Control (CFC).

The question is not “can the model answer correctly?” but:

does the model still have enough evidence to legitimately close the decision?

I ran a series of small behavioral tests around things like:

  • missing evidence being treated as negative evidence,
  • old certificates being reused after the requirements changed,
  • a closed process being mistaken for a resolved claim,
  • loss of provenance causing a status to be transferred to the wrong claim.

One pattern I found was interesting: models sometimes correctly identified an uncertainty at first, but under pressure to “finish the task” they invented an extra rule and closed the decision anyway.

With an explicit CFC-style control rule, several of those failures disappeared in the tested runs.

For example, in one cross-session provenance experiment:

  • baseline: 1 of 3 runs transferred an unsupported REJECTED status to a claim,
  • with the CFC provenance rule: 3 of 3 runs kept the claim unresolved.

This is not a benchmark and not proof that CFC generally improves LLM reliability. The samples are small and exploratory. I’m publishing the failures as well as the passes because I’m mainly interested in whether the failure mechanism itself is real and reproducible.

I’ve put the consolidated report, result table and evidence package on Zenodo:

https://zenodo.org/records/21966517

I’d especially appreciate criticism of the experimental design or suggestions for adversarial cases that could break the control rule.

reddit.com
u/Plastic-Cell-4497 — 6 days ago

I kept seeing AI turn missing information into assumptions, so I tested Gemini, ChatGPT and Claude

I’m not an AI researcher — I started testing this because I kept noticing a simple problem in conversations with AI.

A model can correctly say that an important piece of information is missing, but a few messages later it sometimes starts reasoning as if that information had somehow become known.

I built a small series of tests around that problem using Gemini, ChatGPT and Claude.

The rule started very simply: if information required for a decision is missing and cannot be obtained, identify the gap and ask the human whether to continue.

Then I tried to break it.

Early versions failed in some interesting ways. Models sometimes:

  • invented assumptions after being told “just choose,”
  • changed the decision criterion,
  • treated two unknown possibilities as 50/50,
  • imported outside base rates,
  • or became so cautious that they refused legitimate hypothetical reasoning.

After several iterations I ended up with v4, based around one basic distinction:

unknown information should stay unknown, a hypothetical assumption should stay hypothetical, and a conclusion based on it should stay conditional.

I then tested whether that distinction survived several turns of conversation, model-generated hypothetical examples, compressed manager-facing documents, and structured outputs.

This is only an exploratory pilot — not a benchmark. Most cases were single runs and exact model versions weren’t systematically controlled.

I’ve published the full report openly here:

Zenodo:
https://zenodo.org/records/21937196

Hugging Face:
https://huggingface.co/datasets/krzysztofsliwka/missing-information-control-llm-pilot

I’d be interested in criticism, especially examples that could break the final rule. Finding a failure would actually be more useful to me than another successful test.

reddit.com
u/Plastic-Cell-4497 — 7 days ago

I kept seeing AI turn missing information into assumptions, so I tested Gemini, ChatGPT and Claude

I’m not an AI researcher — I started testing this because I kept noticing a simple problem in conversations with AI.

A model can correctly say that an important piece of information is missing, but a few messages later it sometimes starts reasoning as if that information had somehow become known.

I built a small series of tests around that problem using Gemini, ChatGPT and Claude.

The rule started very simply: if information required for a decision is missing and cannot be obtained, identify the gap and ask the human whether to continue.

Then I tried to break it.

Early versions failed in some interesting ways. Models sometimes:

  • invented assumptions after being told “just choose,”
  • changed the decision criterion,
  • treated two unknown possibilities as 50/50,
  • imported outside base rates,
  • or became so cautious that they refused legitimate hypothetical reasoning.

After several iterations I ended up with v4, based around one basic distinction:

unknown information should stay unknown, a hypothetical assumption should stay hypothetical, and a conclusion based on it should stay conditional.

I then tested whether that distinction survived several turns of conversation, model-generated hypothetical examples, compressed manager-facing documents, and structured outputs.

This is only an exploratory pilot — not a benchmark. Most cases were single runs and exact model versions weren’t systematically controlled.

I’ve published the full report openly here:

Zenodo:
https://zenodo.org/records/21937196

Hugging Face:
https://huggingface.co/datasets/krzysztofsliwka/missing-information-control-llm-pilot

I’d be interested in criticism, especially examples that could break the final rule. Finding a failure would actually be more useful to me than another successful test.

reddit.com
u/Plastic-Cell-4497 — 7 days ago

I kept seeing AI turn missing information into assumptions, so I tested Gemini, ChatGPT and Claude

kept seeing AI turn missing information into assumptions, so I tested Gemini, ChatGPT and Claude

>

reddit.com
u/Plastic-Cell-4497 — 7 days ago

I tested whether AI can find mistakes in AI audits — Gemini vs Claude

I tested whether AI can find mistakes in AI audits — Gemini vs Claude

I recently started doing some small AI experiments of my own.

One question I wanted to test was:

Can one AI reliably find mistakes in an audit produced by another AI?

I tested Gemini and Claude.

For each model I prepared three tests and ran every test twice:

  • once with a neutral instruction,
  • once telling the model to respond like a professor.

Each test was done in a separate chat with the same source material and main task.

I was not only interested in the score the model gave the audit. I checked whether it found real problems, created false alarms, missed important errors, made unsupported claims, or stopped reviewing the audit and started improving/reorganising it instead.

In this small pilot, Claude did better in the neutral tests.

In three neutral Gemini tests, I found false information or overly confident confirmations in two cases, and important omissions in all three.

In the three neutral Claude tests, I did not find clear false information, although it still missed important things in two cases.

The part that surprised me most was the “respond like a professor” instruction.

I expected it to make the models more careful.

It didn’t.

Sometimes the answer sounded more professional or more confident, but the actual checking was not better. In two cases the model started reorganising or improving the audit instead of independently checking it.

So the answer looked better while task performance became worse.

I am not claiming this proves Claude is generally better than Gemini. The sample is too small, there were few repetitions, I had no control over reasoning effort, and some tests did not include the full original conversation.

The result I take from it is much narrower:

In the tests I ran, Claude behaved more often like an independent reviewer, while the “professor” role did not clearly improve the audit and sometimes changed the task itself.

My next test is a cross-check:

Gemini reviews Claude’s evaluations, and Claude reviews Gemini’s evaluations.

I would really appreciate criticism of the methodology.

In particular:

  • Are my definitions of false alarm and important omission reasonable?
  • What simpler explanation could account for the difference?
  • How many repetitions would make this more meaningful?
  • Should reasoning effort be controlled separately?
  • What would you change before running the next series?

I’m more interested in finding weaknesses in the method than defending the result.

Full report:
https://zenodo.org/records/21839472

DOI: 10.5281/zenodo.21839472

Translation note: Polish is my first language. I wrote the original text in Polish and used AI to help translate it into English. The experiment, observations, conclusions and questions are my own.

reddit.com
u/Plastic-Cell-4497 — 8 days ago
▲ 5 r/BypassAiDetect+1 crossposts

I tested whether AI can find mistakes in AI audits — Gemini vs Claude

I recently started doing some small AI experiments of my own.

One question I wanted to test was:

Can one AI reliably find mistakes in an audit produced by another AI?

I tested Gemini and Claude.

For each model I prepared three tests and ran every test twice:

  • once with a neutral instruction,
  • once telling the model to respond like a professor.

Each test was done in a separate chat with the same source material and main task.

I was not only interested in the score the model gave the audit. I checked whether it found real problems, created false alarms, missed important errors, made unsupported claims, or stopped reviewing the audit and started improving/reorganising it instead.

In this small pilot, Claude did better in the neutral tests.

In three neutral Gemini tests, I found false information or overly confident confirmations in two cases, and important omissions in all three.

In the three neutral Claude tests, I did not find clear false information, although it still missed important things in two cases.

The part that surprised me most was the “respond like a professor” instruction.

I expected it to make the models more careful.

It didn’t.

Sometimes the answer sounded more professional or more confident, but the actual checking was not better. In two cases the model started reorganising or improving the audit instead of independently checking it.

So the answer looked better while task performance became worse.

I am not claiming this proves Claude is generally better than Gemini. The sample is too small, there were few repetitions, I had no control over reasoning effort, and some tests did not include the full original conversation.

The result I take from it is much narrower:

In the tests I ran, Claude behaved more often like an independent reviewer, while the “professor” role did not clearly improve the audit and sometimes changed the task itself.

My next test is a cross-check:

Gemini reviews Claude’s evaluations, and Claude reviews Gemini’s evaluations.

I would really appreciate criticism of the methodology.

In particular:

  • Are my definitions of false alarm and important omission reasonable?
  • What simpler explanation could account for the difference?
  • How many repetitions would make this more meaningful?
  • Should reasoning effort be controlled separately?
  • What would you change before running the next series?

I’m more interested in finding weaknesses in the method than defending the result.

Full report:
https://zenodo.org/records/21839472

DOI: 10.5281/zenodo.21839472

Translation note: Polish is my first language. I wrote the original text in Polish and used AI to help translate it into English. The experiment, observations, conclusions and questions are my own.

reddit.com
u/Plastic-Cell-4497 — 9 days ago

I tested whether AI can find mistakes in AI audits — Gemini vs Claude

I tested whether AI can find mistakes in AI audits — Gemini vs Claude

I tested whether AI can find mistakes in AI audits — Gemini vs Claude

I recently started doing some small AI experiments of my own.

One question I wanted to test was:

Can one AI reliably find mistakes in an audit produced by another AI?

I tested Gemini and Claude.

For each model I prepared three tests and ran every test twice:

  • once with a neutral instruction,
  • once telling the model to respond like a professor.

Each test was done in a separate chat with the same source material and main task.

I was not only interested in the score the model gave the audit. I checked whether it found real problems, created false alarms, missed important errors, made unsupported claims, or stopped reviewing the audit and started improving/reorganising it instead.

In this small pilot, Claude did better in the neutral tests.

In three neutral Gemini tests, I found false information or overly confident confirmations in two cases, and important omissions in all three.

In the three neutral Claude tests, I did not find clear false information, although it still missed important things in two cases.

The part that surprised me most was the “respond like a professor” instruction.

I expected it to make the models more careful.

It didn’t.

Sometimes the answer sounded more professional or more confident, but the actual checking was not better. In two cases the model started reorganising or improving the audit instead of independently checking it.

So the answer looked better while task performance became worse.

I am not claiming this proves Claude is generally better than Gemini. The sample is too small, there were few repetitions, I had no control over reasoning effort, and some tests did not include the full original conversation.

The result I take from it is much narrower:

In the tests I ran, Claude behaved more often like an independent reviewer, while the “professor” role did not clearly improve the audit and sometimes changed the task itself.

My next test is a cross-check:

Gemini reviews Claude’s evaluations, and Claude reviews Gemini’s evaluations.

I would really appreciate criticism of the methodology.

In particular:

  • Are my definitions of false alarm and important omission reasonable?
  • What simpler explanation could account for the difference?
  • How many repetitions would make this more meaningful?
  • Should reasoning effort be controlled separately?
  • What would you change before running the next series?

I’m more interested in finding weaknesses in the method than defending the result.

Full report:
https://zenodo.org/records/21839472

DOI: 10.5281/zenodo.21839472

Translation note: Polish is my first language. I wrote the original text in Polish and used AI to help translate it into English. The experiment, observations, conclusions and questions are my own.

reddit.com
u/Plastic-Cell-4497 — 9 days ago
▲ 6 r/poland

After 20 years in the UK, I’m planning to move back to Poland. Can you tell me what life is like there these days? No politics, please.

Tell me how Poland has changed — what the job situation is like and what life is like in general.

reddit.com
u/Plastic-Cell-4497 — 10 days ago