Using Poincaré hyperbolic geometry to solve a volume scaling problem in neural network interpretability
▲ 22 r/compsci

Using Poincaré hyperbolic geometry to solve a volume scaling problem in neural network interpretability

Wanted to share an interesting application of hyperbolic geometry to machine learning interpretability.

The setup: Sparse Autoencoders decompose neural network activations into interpretable features. These features are dictionary atoms embedded in R^(d.) The problem is that the concepts networks learn form branching hierarchies (trees), and trees with branching factor b have O(b^(r)) nodes at depth r. But the volume of a Euclidean ball grows as O(r^(d)) -- polynomially.

This mismatch means that at large dictionary sizes (16K+), there isn't enough Euclidean volume for features to spread out. They collide at the boundary and "die" (stop activating).

The fix: embed dictionary weights in the Poincaré ball model of hyperbolic space, where the volume element grows as sinh^(d-1)(r) ~ O(e^(r).) This matches the exponential branching of concept hierarchies.

The interesting constraint: the forward pass of the autoencoder must stay Euclidean (for compatibility with the host neural network's normalization layers). So the hyperbolic embedding is applied only as a training-time weight regularizer via an entailment cone loss on the Poincaré-projected dictionary atoms.

Empirically, this reduces dead features from 3.8% to 0.2% and improves reconstruction by 9.8% on a 2B-parameter language model.

Paper: https://vishalvermalabs.com/papers/empirical-validation-hypersae-poincare-geometry/ Code: https://github.com/vishal-dehurdle/hypersae

u/visha1v — 9 days ago

TU Wien just proved quantum entanglement in a centimeter-sized crystal. Does this mean the "classical boundary" is just a technological limitation, not a physical law?

I recently read about a breakthrough from TU Wien where researchers detected a high level of quantum entanglement in a strange metal crystal about one centimeter in size. This surprised me because it’s a macroscopic object that you could hold in your hand.

In every undergrad physics class, we are taught that quantum states are incredibly fragile. We are told that macroscopic objects don't show quantum behaviour because interaction with the environment (heat, stray photons, air molecules) causes immediate decoherence.

But if a crystal containing trillions of atoms can maintain macroscopic entanglement, how is it surviving its own internal thermal vibrations (phonons)?

My question is: Why didn't the thermal noise of a centimeter sized object cause the wave-function to instantly collapse/decohere?

Is the concept of a "fragile quantum state" just a reflection of how bad we used to be at isolating systems, rather than a fundamental law of physics? Could we theoretically maintain entanglement in an object the size of a car if it was forged out of the right strange metal, or is there a hard theoretical mass limit?

reddit.com
u/visha1v — 1 month ago

As opto-mechanics creates quantum superpositions with larger and larger objects, is the scientific community ready if the Born Rule turns out to be wrong? (Objective Collapse vs. Decoherence)

We’re seeing some amazing progress in macroscopic quantum mechanics. Scientists are cooling tiny particles and microscopic mirrors to their lowest energy state and putting them into quantum superpositions. Every year, it seems like they can create quantum interference in larger and heavier objects.

Looking at where this research is heading, it feels like we’re getting closer to answering one of the biggest questions in quantum physics: the Measurement Problem.

According to standard quantum mechanics (and the Many-Worlds interpretation), there is no limit. In theory, a cat, a person, or even a planet could exist in a quantum superposition if it were perfectly isolated from its surroundings. The only thing that makes the superposition disappear is interaction with the environment, a process called decoherence.

But Objective Collapse theories, such as Penrose’s idea of gravitationally induced collapse, make a very different prediction. They suggest that once an object becomes large enough, the superposition becomes physically unstable. At that point, the wave function collapses on its own, even without any measurement or interaction with the environment.

For those working in quantum foundations or experiments:

What is the current view in the field? Do you think future experiments will discover a real mass limit where the Schrödinger equation no longer works and objective collapse takes over? Or do most researchers believe there is no such limit, and that the universe is simply one giant entangled wave function, with “collapse” being nothing more than the effect of environmental decoherence?

reddit.com
u/visha1v — 2 months ago

If particle physicists are right about the “Great Desert,” doesn’t that completely challenge our current models of the early universe and cosmic inflation?

Over in particle physics, the lack of new discoveries at the LHC has led to growing concern about the “Great Desert” hypothesis, the idea that there are no new fundamental particles between the Higgs scale and the Planck scale.

But I’ve been wondering what this means for cosmology. If the Great Desert is real, doesn’t that create some major problems for our models of the early universe?

If there are no new particles:

  • What caused inflation? Inflation is usually explained by an inflaton field, which would require a new scalar particle. If nothing exists up to the Planck scale, what actually drove the universe’s rapid expansion?
  • What is dark matter? If particles like WIMPs and axions are ruled out by a true Great Desert, are we left with primordial black holes or ideas like modified gravity instead?
  • What about vacuum stability? Based on the measured masses of the Higgs boson and the top quark, our vacuum appears to be metastable. If there’s no new physics to stabilise it, how did the universe survive the huge quantum fluctuations during inflation without the vacuum collapsing?

Question for cosmologists:

Are you closely watching the LHC results? If the Great Desert turns out to be real, does it force us to rethink decades of early-universe models, or are there standard cosmology explanations for inflation and dark matter that don’t require undiscovered particles?

reddit.com
u/visha1v — 2 months ago

Now that the dust has settled on the final Muon g-2 results, has the physics community quietly accepted that Lattice QCD may have killed our best hint of physics beyond the Standard Model?

First, the experimental precision achieved by the Fermilab Muon g-2 team was incredible, and their recent Breakthrough Prize was absolutely deserved. But I’m more curious about the theory side, which seems to be going through a major shift.

For years, the gap between the measured muon (g-2) value and the Standard Model prediction from the data-driven dispersive approach was considered one of the strongest hints of physics beyond the Standard Model.

However, recent Lattice QCD results, beginning with the BMW collaboration and followed by several independent checks, appear to bring the theoretical prediction much closer to the experimental measurement.

Combined with flavor anomalies such as (R_K) largely disappearing as more data arrived, it feels like particle physics is going through an era of vanishing anomalies.

For theorists and phenomenologists: is the disagreement between Lattice QCD and the dispersive approach finally coming to an end? If Lattice turns out to be correct, does that effectively close the door on muon (g-2) as evidence for new physics? And more broadly, how is the community adapting to a period where many of the most promising BSM hints seem to be fading away?

reddit.com
u/visha1v — 2 months ago

With the LHC increasingly pointing toward the “Nightmare Scenario,” is the €20 billion Future Circular Collider (FCC) being built to answer important physics questions, or is it simply a case of scientists refusing to let go after decades of investment?

I’ve been following the recent discussions around the 2026 European Strategy for Particle Physics, and one thing that stands out is the huge support for the Future Circular Collider (FCC). But the more I read about the current state of particle physics, the more I feel like there’s an uncomfortable question nobody wants to talk about.

When the Large Hadron Collider (LHC) was built, scientists were almost certain it would discover something important. The Standard Model predicted the existence of the Higgs boson, and without it, the theory didn’t really work. The LHC found the Higgs exactly as expected.

But since then, things have been very different. Years of data from the LHC have failed to find many of the new particles and theories that physicists hoped for. No supersymmetry, no extra dimensions, and no clear signs of dark matter particles. Instead, the Standard Model has continued to survive every test.

So my question for physicists is this:

Is there a similar reason to be confident about the FCC, or are we spending around €20 billion on a machine without any guarantee that it will discover something fundamentally new?

Why isn’t there more support for alternatives like a Muon Collider, which could reach higher energies more efficiently, or for precision experiments involving neutrinos and other smaller-scale projects? Is the FCC still the best path forward, or are we continuing down this route because it is the approach the field is most comfortable with?

More broadly, has particle physics become too focused on the idea that bigger colliders will eventually reveal new physics, even though many of the predictions based on concepts like naturalness have not appeared so far?

reddit.com
u/visha1v — 2 months ago
▲ 105 r/Physics

I just finished David Tong’s QFT lectures. If Gauge "Symmetry" is just a mathematical redundancy, why does it have the power to dictate the fundamental forces?

After months of working through index notation, Lagrangians, and more equations than I can count, I finally finished David Tong’s lecture notes.

What I loved most was how he connects the maths to physical intuition. A lot of things that felt like abstract symbols finally started making sense.

The part that really stuck with me was the gauge covariant derivative. As I understand it, if you demand that physics should remain unchanged under a local phase transformation, the ordinary derivative stops working. To make the maths consistent again, you have to introduce a new vector field. And that “mathematical fix” ends up being the photon.

That was a pretty mind-blowing moment for me. It feels as if forces aren’t something we simply add into a theory, rather they appear because the mathematics demands them.

But it also left me with a question.

Tong repeatedly points out that gauge symmetry isn’t a physical symmetry in the same way that, say, moving an object from one place to another is. Instead, it’s a redundancy in how we’ve chosen to describe the system mathematically.

So here’s what I’m struggling with:

If gauge symmetry is just a redundancy in our description, why does it seem to have such a powerful influence on the real world? Why do actual, measurable particles appear when we enforce it? If it’s only a feature of our mathematical language, why does nature seem to care about it so much?

Are forces somehow the physical consequence of these mathematical redundancies, or am I thinking about this the wrong way?

Would love to hear how people with more experience think about this.

And also, huge credit to David Tong. His lecture notes are genuinely fantastic.

reddit.com
u/visha1v — 2 months ago

If I was falling into a supermassive black hole, time would pass more slowly for me compared to the rest of the universe. As I crossed the event horizon, would I be able to see the entire future of the universe happening in fast-forward?

From the view of someone watching me fall into a black hole, I would seem to slow down and almost stop at the event horizon.

But what would I see from my side? If time outside is moving much faster than for me, would I see stars dying, galaxies crashing into each other, and maybe even the end of the universe happening in just a few seconds before I get destroyed?

reddit.com
u/visha1v — 2 months ago
▲ 65 r/universe+1 crossposts

If I was falling into a supermassive black hole, time would pass more slowly for me compared to the rest of the universe. As I crossed the event horizon, would I be able to see the entire future of the universe happening in fast-forward?

From the view of someone watching me fall into a black hole, I would seem to slow down and almost stop at the event horizon.

But what would I see from my side? If time outside is moving much faster than for me, would I see stars dying, galaxies crashing into each other, and maybe even the end of the universe happening in just a few seconds before I get destroyed?

reddit.com
u/visha1v — 2 months ago

Why do we still teach the Word-RAM model by default when caches matter so much more?

It feels like every undergrad CS program still leans completely on the basic RAM model when teaching algorithmic complexity. I get that Big O is a mathematical bound and not a literal benchmark, but pretending memory hierarchies don't exist feels like a massive blind spot when analysing data structures.

For example, standard theory teaches that traversing an array and a linked list are both O(N). But we all know the difference in cache misses makes them completely different beasts. I know things like the Ideal-Cache model and Cache-Oblivious algorithms exist, but they almost always get shoved into niche grad-level courses.

Is anyone actually pushing to introduce cache-aware or external memory models earlier in undergrad? Or is the general consensus just that the basic RAM model is "good enough" for beginners, even if it leads to "theoretically optimal" algorithms that perform terribly in practice?

reddit.com
u/visha1v — 2 months ago
▲ 32 r/ZyadaKuchNai+1 crossposts

Beating the heat with an iced Coke and Pacific Rim. Happy Sunday, folks!

u/visha1v — 2 months ago

Applying Discrete-Time Lyapunov Stability to monitor LLM Agent loop trajectories

Has anyone else explored using classical control theory (Lyapunov stability, state-space representations, or Kalman filtering) to monitor and govern LLM execution spaces? I'd love to discuss the math.

reddit.com
u/visha1v — 2 months ago

Cereal companies intentionally use bags that are impossible to open cleanly so the cereal goes stale faster.

You know how the plastic bag inside a cereal box is supposed to pull apart at the top seam, but instead it violently rips directly down the side, exposing the entire bag to the air? That’s not a manufacturing flaw. Big Cereal designed the plastic's tensile strength to fail exactly like that. You can't roll it shut, the cereal goes stale in half the time, and you are forced to buy another box of Cinnamon Toast Crunch a week earlier than you planned.

reddit.com
u/visha1v — 2 months ago

I actually enjoy the feeling of stepping on a Lego block.

Everyone acts like stepping on a Lego is a torture method outlawed by the Geneva Convention. But honestly? If you don't full-force jump onto it, stepping on a rogue 2x4 Lego brick provides a fantastic, deep-tissue pressure point massage. It’s like a free acupressure mat. Hitting the arch of your foot right on the sharp corner of a plastic brick releases so much tension. I don't purposefully throw them on the floor, but when I step on one my kid left out, I lean into it.

reddit.com
u/visha1v — 2 months ago

A father and son spend a week arguing about proper car maintenance while repeatedly picking up unhinged hitchhikers.

Hint 1: It's an animated movie.
Hint 2: There is a lot of cheese involved.

reddit.com
u/visha1v — 2 months ago

[The Office] Creed Bratton isn't actually a criminal; he just thinks the documentary crew is filming a gritty HBO crime drama.

Everyone assumes Creed is an actual murderer, cult leader, and thief because of the bizarre things he confesses to the cameras. But look at his background: he’s an aging, former rock-and-roll theater kid. What if he realised early on that documentaries are boring, so he decided to "play a character" to make the show better?

He thinks they are filming a true-crime docuseries. Every time he is alone with the camera, he tries to give the producers "good television" by making up insane backstories. When he showed up covered in blood on Halloween and said "It is Halloween... that is really, really good timing," he wasn't relieved he got away with murder. He was relieved he brought his own fake blood on the same day the office happened to be doing a costume party, saving his "character" from looking out of place.

reddit.com
u/visha1v — 2 months ago
▲ 10 r/ControlTheory+1 crossposts

Applying Discrete-Time Lyapunov Stability to monitor LLM Agent loop trajectories

Hey everyone,

I’ve been experimenting with using classical control theory to monitor LLM agent loops (which are essentially discrete-time dynamical systems).

When agents run in multi-turn loops, their state (input tokens, tool call history, error counts) can grow unstable. We modelled this using a normalised Lyapunov energy candidate:

V(k) = S(k) / S̄

where S(k) is the turn token count and is a warmup baseline. A trajectory is flagged as unstable when ΔV(k) ≥ 0 for W consecutive steps.

We ran a 3,175-run study testing this as a runtime guardrail. On SWE-bench search trees, dynamically tripping unstable trajectories early cut compute search nodes by 38.6% and wall-time by 30% with zero false positives on short/stable loops.

Has anyone else explored using classical control theory (Lyapunov stability, state-space representations, or Kalman filtering) to monitor and govern LLM execution spaces? I'd love to discuss the math.

reddit.com
u/visha1v — 2 months ago

Empirical Lyapunov Stability: Modelling LLM Agent Loops as Dynamical Systems

Hello CS community,

We've been researching runtime stability for multi-turn LLM agents, modelling their execution trajectories as discrete-time dynamical systems.

A major challenge is that raw token accumulation (ΔV ≥ 0) is normal in multi-turn context windows, causing a 46% false positive rate if you try to monitor raw energy. We solved this by implementing growth-ratio normalisation: median-aggregating a warmup baseline and tracking relative deviations.

We validated this on a 3,175-run ablation study across SWE-bench Verified, τ³-bench, and MINT. On SWE-bench search trees, dynamically terminating unstable branches cut node expansions by 38.6% with zero impact on the resolve rate.

The core implementation is open-source (search state-harness on GitHub). Curious to hear thoughts on modelling agent safety boundaries using dynamical systems theory.

reddit.com
u/visha1v — 2 months ago
▲ 21 r/FunMachineLearning+4 crossposts

Empirical Lyapunov Stability: Runtime Observability and Failure Classification for LLM Agents (OpenTelemetry + Python/Rust Library)

Hello Community,

Standard budget caps tell you that an agent failed, but they don't tell you why (did it loop on a broken tool? did its context spiral?).

To solve this, we open-sourced state-harness (https://github.com/vishal-dehurdle/state-harness); a lightweight Python/Rust runtime guard that tracks a normalised token growth ratio (inspired by discrete Lyapunov stability) and exports failure signatures straight to OpenTelemetry.

u/visha1v — 9 days ago
▲ 35 r/ChatGPT

Are we actually building "AI software engineers," or are we just creating incredibly fast, highly confident junior devs who don't sleep but write legacy code at 10x speed?

Every coding agent demo looks amazing until it hits actual system architecture. Right now, it feels like we're just scaling up technical debt at warp speed. They're brilliant at churning out syntax and boilerplate, but completely lack real engineering intuition.

Are we actually moving toward autonomous engineering, or are we just generating tomorrow's legacy code faster?

reddit.com
u/visha1v — 2 months ago