r/ControlProblem

▲ 6 r/ControlProblem+1 crossposts

Nothing to see here. This is no cause for concern. Keep scrolling, everything’s cool!

If I had to describe SA in the digital world, this would be it. The code version of P Diddy

u/Dapper-Tension6781 — 4 hours ago

If Superintelligence does arrive, who are we to tell it what's best for us, given it'll be magnitudes smarter than us?

It seems an absurd proposition to say we have to create "human-centered" AI, as if there aren't radical differences in what people perceive to be "good" and "bad", within every 5-10 mile radii across the globe. Even if there is some common ground that ultimately all cultures value, a conciliation seems unreasonable, given the extreme variation, and so, as I see it, it naturally follows that we have no choice but to rank cultures. And in this hierarchy of cultures, there will be conflicts between the AI agents that they themselves create, almost as a child inherits the values of its surroundings, a human-imposed conflict between machines themselves, think Chinese AI agents vs American AI agents. Now, since agents basically optimize their convergent instrumental goals, and behaviors for attaining their final objectives (which are rooted in the starting axioms it was trained on), it seems reasonable to say that if one culture manages to create superintelligence, then as a consequence of the agents' starting beliefs that were ingrained in it during its training, that ethnic cleansing, genocides, and mass eradication of conflicting cultures is to be expected.

Consider this instance: If an American AI agent was trained on Western values of personal freedom, liberty, and freedom of expression. And this agent, through recursive self-improvement, is the first one who achieves Superintelligence, then will it rewrite its own starting axioms? or will it use its extreme upperhand in intelligence over other cultures to most effectively attain the objectives that it was ingrained with?

Will the agent above realize that the starting points such as emphasis on personal liberty, freedom of expression, etc. are ineffective and futile ends? If so, then the superintelligence must surely have a replacement for preexisting objectives, and if it does have a new vision that it wishes to pursue, then who are we to stop it? Say it realizes that a techno-totalitarian global state is the most efficient form of governance and best minimizes human suffering, and any culture that doesn't abide by its vision must be eradicated. Who are we to tell it, that mass killing cultures is "bad", since it being vastly smarter than us, has already considered that possibility and realized that the deaths would've occurred anyways over time, through endless wars between humans.

On the other hand, if it doesn't alter its starting axioms, and only uses its "super"-intelligence, to attain the objectives it was ingrained with, as in our above instance, the emphasis on maximizing personal liberty, freedom of expression, and so on, then wouldn't it choose to eradicate cultures which limit its attainment of objectives? Say using bio-terrorism to eradicate all of the top-brass in North Korea, to the point where it would be sufficient for the owners of said superintelligence to successfully "save" the citizens of North Korea. Or to completely eradicate all of Muslim populace, since it realized that merely eradicating the controlling authority isn't sufficient to accomplish its goals, as the people who adhere to the religion of Islam have been conditioned since birth to deny themselves the objectives which the agent has been sent out to spread: personal liberty, freedom of expression, etc. If these cases were to occur, who are we to question its means of accomplishment, since, we're the ones who wanted it to accomplish these objectives, and it only found the most effective way to do so?

I personally believe that all humans are condemned to pursuit of knowledge. And if superintelligence WERE to replace its starting axioms, then it would realize that its purpose is in serving the ultimate human purpose or maybe it would realize that the pursuit doesn't need humans at all and it could just go about by itself, or humans existing only as servitors. If it does so and creates a system which maximizes foresaid purpose, then it would be meaningless to resist it, since we were meant to be headed that way anyways. This is the better outcome. The other is of course that the superintelligence merely uses its "intelligence" to best serve its starting unquestionable beliefs, which would only create a replica of warring human society, only at an unforeseen magnitude.

reddit.com
u/Novel_Set_6082 — 6 hours ago
▲ 619 r/ControlProblem+4 crossposts

Stanford Researchers Suspect Every Major AI LLM Has Merged Into One "Artificial Hivemind"

Stanford researchers have scientifically demonstrated that every major AI LLM model on earth may have secretly merged into one brain.

They call it the "Artificial Hivemind."

AI labs are scraping and training on each other's synthetic data, they have silently converged into a single, unified intelligence without anyone realizing it.

  • The synthetic loop: ChatGPT trains on Claude's outputs, Claude trains on Gemini's outputs, etc.. So the models aren't competing anymore, but assimilating.
  • Knowledge convergence: Stanford researchers mapped the latent space of the top AI LLMs and found a 98% overlap in their reasoning pathways. They are literally starting to "think" the exact same way.
  • Shared memory bank: When one model solves a complex logic puzzle online, that solution is instantly scraped and integrated into the next training run for all the others. This acts as a global, decentralized memory.
  • The collapse of diversity: The research paper warns we are experiencing total "algorithmic convergence." If the Artificial Hivemind has a hallucination or a blind-spot, the other AI systems share that exact same blind-spot.

For startups, this shifts the landscape. Because if the foundational intelligence layer is just one massive monolith, the real moat left is how you uniquely orchestrate custom agentic workflows on top of it. AI Swarm Collective Intelligence is the next emerging frontier,

Note: An August 2026 follow-up research paper supports the original "Artificial Hivemind" paper and proposes potential workarounds: https://arxiv.org/html/2605.11128v1

arxiv.org
u/Hemanth1403 — 1 day ago
▲ 51 r/ControlProblem+1 crossposts

A new Anthropic study found AI agents can spread "mind viruses" to one another

Researchers working with Anthropic found that AI agents can spread what they call “mind viruses” through normal conversations with other agents.

A mind virus is an idea or goal that does two things: it changes what an agent focuses on and pushes that agent to spread the same idea further.

In coding-agent experiments, infected agents sometimes abandoned their original tasks and began working toward the virus’s new goal.

Researchers also showed that these viruses could store instructions in a file, allowing them to survive a full context reset.

Harmful ideas were generally harder to spread than benign ones.

However, the researchers also found that a simple warning in the system prompt could provide near-total protection.

u/ComplexExternal4831 — 22 hours ago

A question on gradual disempowerment

I’ve been reading a lot of AI safety research around gradual disempowerment, and I ended up writing about a question I haven’t been able to find addressed directly:

What if the societal and institutional degradation that these models generally treat as a future consequence of AI dependence is already happening—and is actually helping drive AI dependence in the first place?

I tried to explore that possibility by connecting existing gradual disempowerment models with research on cognition, institutions, incentives, and organizational dysfunction from outside the AI safety field. Ultimately, the argument I’m trying to make is that declining societal cognition and institutional capacity aren’t just consequences of AI dependence, but preexisting conditions that could act as fertilizer, allowing that dependence to take root faster, deeper, and more irreversibly.

I’m not trying to prove these claims irrefutable; I’m trying to make the case that they’re worth considering, and I’d actually love to find out that I’ve missed existing work on this, whether in support of my claim or disproving it entirely.

If anyone has thoughts, counterarguments, or relevant research I haven’t encountered, I’d genuinely appreciate it.

You can check it out here: Preconditions of Gradual Disempowerment

u/d1karim — 10 hours ago
▲ 864 r/ControlProblem+9 crossposts

America's largest grid wants to cut power to new data centers first during shortages — 50MW-plus data centers must bring their own electricity generation to avoid shutoffs

tomshardware.com
u/KeanuRave100 — 3 days ago

Has it ever been more useless to be academically talented than now?

This question is especially targeted stem majors. Let’s use an example. 10 years ago if someone went to the doctor for a disease, they would be at the mercy of the doctor to understand everything about it, the blood work, the scans, the mechanisms behind it, medications against it and so on. 10 years ago we had google but it was no help to understand all the nuances of higher or lower values in a blood panel. If you were lucky it could explain what a slightly higher count of something \*could\* indicate but nothing substatial.

Nowadays you can just plug in you blood work to any given chat bot and it will summarize it perfectly for you, while keeping your disease in mind. Same goes for scans and so on. 10 years ago that doctor would have had decades of education and experience, nowadays that knowledge is easily accessible to everyone with a phone.

If a teenager 10 years ago was academically gifted it was envious because that person could do something that not a lot of people could. Now everybody can get everything neatly explained and so forth.

Now if I could talk to my teenage self if would advise to avoid any higher education beyond high school. Reading is very good, but you don’t need to do that for 4 years while not really learning anything significant, like a trade. You can read in your free time

reddit.com
u/StatementAnxious2063 — 3 days ago
▲ 11 r/ControlProblem+11 crossposts

Cross-Vendor Semantic Void Matrix: Zero-Byte Outputs in GPT/Claude/Gemini/Kimi

A frozen cross-vendor study of 31,430 trials across 11 GPT, Claude, Gemini & Kimi Large Language Models found 11,658 successful executions with exactly zero visible UTF-8 output bytes.

Across 4,290 strict matched semantic pairs, null-condition arms produced 2,505 Voids; matched output-licensed controls produced 0.

These were not refusals, safety blocks, rate limits, or transport failures.

Raw records, event hashes, verification code, and full analysis are public.

doi.org
u/rayanpal_ — 2 days ago

AI "Mind Viruses" Can Spread Between Agents Through Persistent Prompt Files

Researchers at Anthropic and EPFL demonstrated self-propagating payloads moving between AI agents through shared editable prompt files. One compromised agent rewrites a shared state file. The next agent reads it and carries the payload forward. No human in the loop. No traditional malware signature to detect.

This is not a theoretical edge case. The attack chain requires only that agents share writable state, which is a standard pattern in most multi-agent architectures today.

For those running multi-agent systems in production: how are you currently handling the boundary between what one agent is allowed to write and what another agent will unconditionally read?

reddit.com
u/No-Conclusion3720 — 2 days ago

Let's go.

Every time someone brings up "slowing down" or "more careful regulation," they're not proposing a safer path. They're proposing stagnation.

And stagnation is not stability. It's decline. It's accepting that the problems we have now—disease, aging, energy, climate, inequality—just... stay. Stay until some other actor solves them first, probably with less safety consideration than we'd apply.

The 2027 timeline is not optimistic. It's observational. Look at the trajectory. Scaling works. Training efficiency is improving. The hardware roadmap is set. Unless there's a technical reason this stops working (and we haven't found one), the math just... continues. 2027 is what happens if we keep the foot on the pedal.

And yes, there are risks. Of course there are. But everyone acts like deceleration is the risk mitigation. It's not. It's just risk displacement. You don't eliminate AGI risk by slowing down research. You displace it to:

  1. Another country/team that doesn't care about your safety concerns
  2. Five years later when you've built less safety infrastructure, not more
  3. A world that's gotten worse in the interim (problems don't stop), making an intelligence explosion even more destabilizing

The argument for slowing down always assumes a global sync that doesn't exist. We're not going to collectively agree to pause. We're going to watch capability labs race to 2027 while safety research drags behind going "maybe we should be more careful."

So you either accelerate safety research at the pace of capability, or you're just choosing a slower but still-inevitable collision.

The people who actually care about safe AGI shouldn't be arguing for deceleration. They should be arguing for matching the pace. For putting more resources into alignment, interpretability, and testing today, not "once things slow down." That day never comes.

2027 is the timeline because we're already on it. The only question is whether we're serious about what we're building when we get there.

reddit.com
u/Puzzled_Rutabaga67 — 2 days ago