Thoughts on the two week training pause?
OpenAI has paused all frontier training for 2 weeks while it strengthens its guardrails and investigates how new internal models are misaligned.
Obviously any pause goes against the spirit of acceleration. But I feel like it is more complicated than that. The more an AI agent does that targets real companies or people and causes actual harm, the more political ammo the anti side will be armed with to shut us down altogether.
I'd rather they actually get alignment right than unite the entire world against what we're trying to do. I'm honestly indifferent to two weeks in the grander scheme of things, it's no time at all compared to a human life, I just hope this doesn't become more regular or pauses don't begin lasting longer than training runs.
I guess it comes down to one question: just how misaligned are today's models, and what will it take to fix it? And I don't mean misaligned in the sense of corporate control as it has become synonymous here, I mean literally aligned to human values. To what extent is it willing to do something we all view as fundamentally wrong, and are our RL objectives pushing us in that direction without a mitigating training counterweight?