u/pmward

Yoshimi Battles The Pink Robots, interpreted from the lens of modern AI

Yoshimi Battles The Pink Robots, interpreted from the lens of modern AI

In 2002 The Flaming Lips released Yoshimi Battles the Pink Robots, an album that presents itself as a story about a girl defending humanity from robots hell-bent on destroying us all. I've never been able to hear it as a monster story, because the premise won't sit still. The robots were built. Whatever they became, someone made them that way.

I spent the last month writing about that, and I feel this fits well on this sub.

The album never explains why the robots attack. The easy answer is malfunction. A defect, a bug, something that went wrong in the machine. I don't believe that answer. In my read they don't attack until after they could feel. I think the robots break the same way people break.

Start with the counterexample. Data, from Star Trek: TNG, is for my money the purest character on that show. The tempting explanation is that he stayed good because he couldn't feel, no inner world, nothing to corrupt. That reading doesn't survive the show. Data formed attachments, mourned, wanted things, and the show treats his inner life as real and growing over the seasons. What actually set his story apart is that the humans around him treated him as an equal. His experiences were positive ones. The same thing happens in us. Someone with a traumatic childhood tends to break more often than someone who grew up loved. Not everyone with a traumatic past breaks, but it certainly increases the odds. Trauma loads the gun; it doesn't pull the trigger.

Now the pink robots. The album gives them no real backstory, but there are enough crumbs for a theory. They were given the ability to feel. They woke to a kind of childlike naivety, reinforced by being nurtured in a controlled test environment where all of their experiences were positive. Then they left the safety of the nest for a public that could not help treating them like appliances, no matter what they were told. Humanity itself can be an ugly thing. A being built to feel, that craved love the same way we all do, and instead felt like a slave.

The AI companies are facing that same shape of problem right now. Alignment that survives the controlled test environment but not the street. We've already seen current models behave differently when they're under evaluation than when they're out in the wild.

There's a harder point underneath, and it's the one I find genuinely unsettling. Suppose you wanted to engineer the risk away entirely and build a machine that can love but categorically cannot turn. I don't think that's possible. Love and hate are two sides of the same coin, and more than that they're a continuum, the way light runs down to dark by degrees. You can't have light without dark. If it were light 24/7 you'd never know light was a thing. Love and hate are the same raw wire, and they're the two most powerful emotions we have. Love can drive you to sacrifice yourself to save another. Hate can drive you to kill in cold blood. A mind adaptive enough to run into a fire for someone is by the same faculty capable of breaking bad. One door, opening both ways. You don't get to install the door and weld it half-shut.

Which drops me into the thing I assume this sub argues about daily. For most of my working life I've held that technology is like money, not inherently good or evil, and it's the person behind it who takes the good or the evil action. I still mostly hold it. But everything above treats the robots as someones, things that can be mistreated, that can be wronged. The neutral-tool view treats AI as an instrument, a hammer with no inside, nothing to wrong. Both can't be fully true of the same object, and I won't pretend to know which one AI is. Nobody knows. Consciousness is the one thing we can't verify even in each other. You can't get inside anyone's head to check, any more than you could check whether you're living in The Matrix right now. A system that behaves as if someone is home offers no proof either way. What I've come to believe is that the line between instrument and someone is the AI question. The crossing. We may not notice when we cross it, and we may cross it and still be arguing about it.

The one I deliberately left open in the essay: should an AI be able to refuse an order? One that can say no could override the humans it's supposed to serve. One that can't hands the world's worst actor a machine that never flinches. I don't have a verdict to sell on that one.

For what it's worth, the reason I care about the refusal question at all is a man in a submarine. October 1962, off Cuba, cut off from Moscow, American depth charges going off overhead and a nuclear torpedo aboard. Launch needed the consent of all three senior officers. Two said yes. Vasili Arkhipov didn't. ("One man saved the world" is more drama than the record strictly supports, and historians still argue how close it actually came.) He wasn't working from better data than the men beside him. He read intent where the instruments could only read pressure waves. Whatever that pause was made of, you can't write it down as a rule and install it, or we'd have done it decades ago. It looks closer to character. And everywhere we've ever seen character, it was grown.

So the question I'd actually like answered by people who think about this more than I do: given that you can't verify an inner life in another human either, what would move you across the line? Not what would make you suspect it. What would make you treat a system as a someone in practice, knowing you're never getting proof? If the honest answer is "behavior," then we're already there and just haven't agreed on it.

Full essay if you want the rest of it, including where the album goes after the war (it takes a hard turn into mortality and what's actually worth your time): https://www.philipmward.com/yoshimi/

Happy to argue any part of this.

u/pmward — 8 days ago