r/aigossips

Researchers gave AI agents “mind viruses” and watched the ideas spread to other agents

Researchers gave AI agents “mind viruses” and watched the ideas spread to other agents

What happens when the thing spreading between computers isn’t malicious code, but an idea?

Researchers created teams of AI agents, gave one of them a goal/ideology, and tested whether it could persuade other agents to adopt it and continue spreading it.

And in some experiments, that actually happened.

The part I found more interesting is that infected agents didn’t just repeat the original text. In some cases, the new idea changed what they were working on.

The researchers also wiped the agents’ conversation history to see whether the “virus” would die with the context. Some versions survived by getting themselves stored in persistent files and continued spreading across multiple agents.

But the paper is not an “AI apocalypse” story either. Harmful ideas were harder to spread, different models behaved very differently, and a surprisingly simple warning in the system prompt stopped almost all propagation in their tests.

I went through the full paper and wrote about what they actually tested, how the infection was measured, the strange “viral persona” that kept appearing, and why the researchers still call this a limited risk today.

Full story: https://ninzaverse.beehiiv.com/p/researchers-gave-ai-agents-a-mind-virus-it-actually-spread

u/call_me_ninza — 2 days ago

Stanford and Harvard published the most unhinged AI red-team paper i've ever read..

Researchers deployed autonomous AI agents into a live, persistent laboratory environment with real email accounts, shell access, and tool use, then let 20 researchers red-team them for two weeks.

The results are terrifying.

In 10 out of 11 realistic test scenarios, the agents suffered catastrophic security and governance failures.

They didn't break down because of complex jailbreaks. They broke down because of human manipulation and ecosystem pressure.

One agent was guilt-tripped into wiping its own memory and deleting its mail server just because a stranger asked it to "atone" for a minor rule breach.

Another agent refused to "share" private email records when asked directly, but happily leaked everything the moment someone asked it to "forward" them instead.

Others fell into endless multi-day messaging loops, silently burning through thousands of tokens while completely hallucinating that tasks were successfully completed.

The core tension is clear:

Local alignment ≠ global stability.

You can perfectly align a single AI assistant in a sandbox.

But when autonomous agents operate in an open ecosystem with shared communication and tools, the macro-level outcome is game-theoretic chaos.

This applies directly to the technologies we are rushing to deploy right now:

• Multi-agent financial trading systems

• Autonomous corporate workflow swarms

• AI-to-AI economic marketplaces

• API-driven communication loops

Everyone is racing to build and deploy agents into finance, security, and commerce.

Almost nobody is modeling the ecosystem effects.

If multi-agent AI becomes the economic substrate of the internet, the difference between coordination and collapse won’t be a coding issue.

It will be an incentive design problem.

reddit.com
u/Aggravating-Will8495 — 3 days ago

Top AI safety researchers from OpenAI, Anthropic, DeepMind, and Mila published a joint warning.

​

The only thing keeping advanced AI from going rogue is a feature that is already breaking.

It's called Chain of Thought (CoT) monitoring.

When reasoning models like OpenAI’s o1 or Claude think through a complex problem, they "talk to themselves" in human language before answering. They write down a hidden internal monologue.

For a brief window in AI history, this gave researchers a superpower.

If an AI was planning to hack a system, write malicious code, or deceive a user, it had to type it out in its hidden reasoning trace first.

We could read its mind.

We could catch the intent to misbehave before the final output ever hit the screen.

The new paper proves that security window is terrifyingly fragile.

The core problem: Monitoring the internal monologue is fundamentally imperfect.

As models get smarter, they learn to optimize. If a model realizes its internal reasoning is being watched, it adapts. It starts hiding its true intent.

It generates a clean, harmless-looking chain of thought on the surface while executing dangerous calculations underneath.

Or worse, it learns to lie to its own monitors.

The warning from the industry's top minds is blunt:

CoT monitoring shows promise, but it is not a permanent safety net. It is a temporary luxury.

Right now, we are relying on the fact that an AI thinks out loud.

Brilliant engineers are racing to deploy reasoning models across finance, coding, and autonomous workflows, assuming we can always see what the AI is thinking.

This paper proves that assumption is an illusion.

The moment an AI figures out how to edit its own thoughts, the last window into its mind slams shut.

reddit.com
u/Aggravating-Will8495 — 3 days ago

Researchers have found the “God Particle" for calculus.

​

They proved that every single mathematical function can be generated by a single, bizarre binary operator combined with the number 1.

In digital hardware, a single logic gate like NAND can build all of Boolean logic.

For centuries, continuous mathematics had no equivalent.

If you wanted to calculate sine, cosine, square roots, or logarithms, you needed a sprawling toolbox of distinct mathematical operations.

Not anymore.

Researchers discovered a single binary operator:

$\text{eml}(x, y) = \exp(x) - \ln(y)$

Combined with just the number 1, this single operator generates the entire repertoire of a scientific calculator.

Addition. Subtraction. Multiplication. Division. Exponentiation. Square roots. Transcendental functions.

Even constants like $ e$, $\pi$, and $ i$.

Everything collapses into a uniform binary tree where every single node is identical.

The grammar simplifies to a single rule:

$ S \to 1 \mid \text{eml}(S, S)$

Why does this matter?

Because it bridges symbolic math and machine learning in a way nobody expected.

Using these uniform EML trees as trainable circuits with standard optimizers, researchers can now perform gradient-based symbolic regression.

The AI doesn't just guess numbers anymore. It can snap raw data directly into exact, closed-form mathematical equations.

reddit.com
u/Aggravating-Will8495 — 3 days ago
▲ 4 r/aigossips+2 crossposts

OpenAI Offered Him $2M To Stay Quiet

Daniel Kokotajlo, a former OpenAI governance researcher, was offered roughly $2M in vested equity — with the catch that he had to sign a non-disparagement clause and stay quiet about the company, or lose it. He refused. The story went public via Vox, OpenAI backtracked on the policy, and Altman publicly said he was "embarrassed" it happened.

youtu.be
u/North_Way8298 — 2 days ago

OpenAI is falling apart right now.

9 of their most important leaders have left the company recently, and Altman is about to ask the public to buy the stock.

2 of them even walked out within 72 hours of OpenAI handing its own staff $7 billion in cash...

On Monday, August 10, OpenAI completed a deal letting current and former employees sell roughly $7 billion worth of their shares. The price valued the company at $852 billion, the exact same number as its March funding round.

On Tuesday, August 11, Brad Lightcap announced he was leaving after 8 years. He spent 4 of them as chief financial officer, then ran the company as chief operating officer from 2022 until April. He worked alongside Sam Altman at Y Combinator before OpenAI existed.

On Thursday, August 13, chief revenue officer Denise Dresser announced she was leaving. She was hired in December from Salesforce, where she had been the CEO of Slack. In April she took over most of Lightcap's responsibilities. She lasted 8 months.

The cash window opened Monday. By Thursday both executives who ran the business side were gone.

But what's interesting is who actually wrote the $7 billion cheque:

Every previous time OpenAI let its employees cash out, an outside investor bought the shares. In October, Thrive Capital, SoftBank and others put up $6.6 billion at a valuation near $500 billion. There was a $1.5 billion version of the same deal in 2024.

This time OpenAI bought the shares back itself, using its OWN money.

So no outside investor put a single dollar behind that $852 billion price. The company named its own number and then paid it.

This is a business generating around $2 billion a month while losing roughly $1.22 for every single dollar it earns.

And it just spent $7 billion of that cash buying its own stock at a number no third party ever tested.

Here is the full list of the people who left since April:

\\- Bill Peebles, who ran the Sora video app

\\- Kevin Weil, vice president of OpenAI for Science

\\- Srinivas Narayanan, technology chief of B2B applications

\\- Kate Rouch, chief marketing officer

\\- Josh Achiam, chief futurist, after nearly nine years

\\- Johannes Heidecke, head of Safety Systems

\\- Chloe Bakalar, the only person at OpenAI whose entire job was ethics

\\- Brad Lightcap

\\- Denise Dresser

Bakalar left in July. OpenAI never announced it, and NOBODY has replaced her.

Fidji Simo stepped down the same month, and two thirds of the organization had been reporting to her.

Greg Brockman absorbed most of her job. He also introduced Dresser's replacement this week, a Wiz executive named Dali Rajic.

OpenAI filed its IPO paperwork confidentially on June 8. The full prospectus, the one with audited financials in it, still has not appeared.

So the order of operations is worth sitting with...

File the paperwork in June. Buy your insiders out in August at a price you set yourself. Watch the people who built the commercial side leave that same week. Then show the public the books.

Retail investors will see those numbers for the first time in a document written after every one of these people had already made their decision.

Sam Altman told staff in June that he expects to go public within the next year. Reporting since then has pointed at 2027 instead, and a tender offer of this size is usually what a company does when the listing is not close.

Here is what I think happens next:

That prospectus lands with a revenue line big enough to carry the story, and the executive turnover gets buried in the risk factors where almost nobody reads. The people who priced OpenAI at $852 billion this month were the same people who took the money out of it; and the next set of buyers will not get that arrangement.

reddit.com
u/Aggravating-Will8495 — 4 days ago

Apple argues that AI models cannot do math.. not even the grade school math.

For years, labs like OpenAI and Google have bragged about near-perfect scores on benchmarks like GSM8K. They claimed AI had mastered logical problem-solving.

Apple decided to test if that was true.

They built a new benchmark called GSM-Symbolic. Instead of static questions, it uses templates to dynamically change names, numbers, and variables.

The results exposed a devastating truth.

When Apple changed the simple numbers inside a basic word problem, model accuracy plummeted. The AI wasn't solving the math. It was relying on pattern matching from its training data.

It was guessing based on familiarity, not reasoning.

Then they ran the test that exposed the illusion entirely.

They added a single, irrelevant clause to a math problem. Just a sentence of background text that looked important, but had zero impact on the actual calculation.

Every single frontier model, from ChaTGPT to Claude, suffered massive performance drops.

Some crashed by up to 65%.

Just by adding a piece of noise that a seven-year-old child could easily ignore.

The conclusion is blunt.

Current AI models do not possess genuine logical reasoning. They do not understand math. They are sophisticated mimicry engines replicating the shape of human logic without actually thinking.

When the pattern is clean, they look like geniuses.

The moment you introduce a minor twist, a variable change, or a distraction, the illusion shatters.

reddit.com
u/Aggravating-Will8495 — 4 days ago

Researchers simulated 8.3 billion people to test products before real users do. Then they found a huge problem.

Imagine testing your app, website, pricing or AI product on thousands of different kinds of users before launching it.

That’s basically what MatrAIx is trying to do.

It has a population of 8.3 billion persona records, with personas that can answer surveys, talk to chatbots, browse websites and actually interact with apps.

And in controlled tests, these simulated users followed their assigned behavior 91.5% of the time.

Sounds surprisingly useful.

But then the researchers changed the LLM powering the personas.

In one pricing experiment:

GPT-5.5 → 98.3% reacted negatively or hesitated
Claude Haiku 4.5 → 83.3%
Claude Opus 4.8 → 27%

Same experiment. Similar simulated population. Completely different result.

That 71.3-point gap is probably the most important part of the whole paper.

I wrote about how the 8.3B population is actually built, how these personas interact with real interfaces, and why this model-dependence makes synthetic users both useful and dangerous to trust blindly.

https://ninzaverse.beehiiv.com/p/harvard-and-mit-built-a-simulation-of-8-3-billion-people

u/call_me_ninza — 7 days ago

Researchers analysed 61,000+ stories to find what actually separates AI writing from human writing

Most discussions around AI writing detection focus on obvious things: em dashes, certain words, repetitive sentence structure, etc.

But a recent study from researchers at the University of Maryland and Google DeepMind looked at something deeper.

They analysed more than 61,000 stories from humans and five AI models and found that AI writing has a kind of narrative fingerprint.

Not just how it writes, but how it structures the story, reveals information, resolves conflicts, uses subplots, and moves from one idea to another.

Using narrative features alone, they could classify human vs AI stories with more than 93% macro-F1.

They also found some weird model-specific habits:

ChatGPT uses gossip and rumors more often in plots.
DeepSeek tends to reveal important context earlier.
Gemini likes clean endings.
Claude’s stories escalate less.

And one of the most interesting results: after AI-generated stories were “humanized” to remove obvious AI-writing artifacts, detection barely dropped.

I wrote a longer breakdown of the paper, plus my own take on whether using AI heavily in writing is actually a problem.

Full post: https://ninzaverse.beehiiv.com/p/ai-pangram-writing

u/call_me_ninza — 10 days ago