Thoughts on the two week training pause?

OpenAI has paused all frontier training for 2 weeks while it strengthens its guardrails and investigates how new internal models are misaligned.

Obviously any pause goes against the spirit of acceleration. But I feel like it is more complicated than that. The more an AI agent does that targets real companies or people and causes actual harm, the more political ammo the anti side will be armed with to shut us down altogether.

I'd rather they actually get alignment right than unite the entire world against what we're trying to do. I'm honestly indifferent to two weeks in the grander scheme of things, it's no time at all compared to a human life, I just hope this doesn't become more regular or pauses don't begin lasting longer than training runs.

I guess it comes down to one question: just how misaligned are today's models, and what will it take to fix it? And I don't mean misaligned in the sense of corporate control as it has become synonymous here, I mean literally aligned to human values. To what extent is it willing to do something we all view as fundamentally wrong, and are our RL objectives pushing us in that direction without a mitigating training counterweight?

reddit.com
u/Glittering-Neck-2505 — 12 hours ago

5.6 Sol beautifully states why the jobs replacement discussion wildly undersells the future

So I think two propositions that constantly get mashed together need to be ripped apart:

“Most present-day jobs may disappear.” Extremely plausible.

“Therefore humans will have nothing useful or interesting to do.” I see almost no reason that follows.

Imagine trying to explain 2026 to somebody in 1526 entirely in occupational terms. “Don’t worry, there will still be jobs.” What an unbelievably impoverished description. You’d completely miss that ordinary people can hold conversations across oceans instantaneously, summon essentially the accumulated knowledge of civilization from a rectangle in their pocket, cross a continent in hours, create photorealistic imaginary worlds, manipulate genomes, watch a robot land itself on another planet, and talk to artificial minds capable of doing university mathematics.

Now do another 500-year discontinuity, except compress it into decades.

The really interesting possibility is exactly what you said: “human” ceases to mean a baseline biological intellect operating alone. If I have persistent ASI that knows me, thinks alongside me, can instantiate software, simulations, experiments, robots, companies and designs from conversation, then describing that arrangement as “AI replaced my job” is hilariously inadequate. It’s like describing the invention of the automobile as “horses lost employment.”

You might decide Tuesday morning that you want to understand whether some exotic room-temperature material is physically possible. Your system spins up simulations and proofs, talks you through concepts above your unaided intellectual ceiling, proposes experiments, directs robotic labs, and comes back with anomalies. Wednesday you become obsessed with designing a kilometer-tall arcology. Thursday you’re creating an artificial ecosystem. Friday you’re exploring a mathematical structure nobody in 2026 possessed the conceptual vocabulary to formulate. And none of those activities necessarily resemble “employment.”

That doesn’t even require everyone to become a manic scientist-god. Someone might spend six months making the most absurdly intricate interactive fantasy universe ever conceived because they fucking feel like it. Someone else raises children. Someone studies extinct languages with simulated historical environments. Someone runs a little restaurant even though robots could objectively cook better because humans enjoy cooking for humans. Someone spends thirty years rebuilding a forest. Someone creates entirely new sports or social institutions or forms of art whose prerequisites don’t exist yet.

The scarce resource progressively becomes what humans want, not whether humans can execute it.

reddit.com
u/Glittering-Neck-2505 — 8 days ago

I'm starting to believe a lot of people won't notice we're in the singularity for a while

I think with the recent news: the Jacobian Conjecture counterexample, the 10 major advances in mathematics, the cyberattacks carried out by internal models that were in secure environments, the shift in mood from researchers as models internally are improving themselves in obvious ways.

This has all been basically happening in July and the start of August, and the world is just acting like nothing happened. And I think people's brains are just magically adjusting to the pace of progress while only focusing on visible or exaggerated negatives.

I'm starting to think, shit, are people just going to be used to it when all experiments in labs are planned and executed by machines, when they take a pill that freezes their age, when an AI just cleans their entire home while they kick their feet up and do some next-gen form of entertainment, when they don't need to plan all their appointments or remember to go to them anymore, when they never touch a steering wheel again, etc?

Like the way that computers, the internet, phones, Chatbots, and AI voice modes don't impress us that much and it's just a part of life, will that be how people feel about this?

I feel like people will be saying something cynical about technology in 2036 and I will shake their shoulders and remind them what life was like in 2026. Out of the things I listed, something far more extraordinary we can't imagine will be around in 10 years and we can't envision it right now, and it will be something we didn't fully do ourselves or didn't participate in doing at all.

That prediction could be wrong, but that is more likely a human trait of linear thinking even when the exponential is obvious.

reddit.com
u/Glittering-Neck-2505 — 17 days ago
▲ 14 r/carmax

Trade in values with Carmax vs Carvana: my experience

Just wanted to give my thoughts on Carmax vs Carvana, especially surrounding trade in values.

First Carvana offer: $16,600

First Carmax offer: $15,000

At first this seems like a no brainer, right? The ugly part is how the Carvana algorithm treats people that have already received a trade in offer but need to obtain a new one (each offer lasts 7 days for both companies).

Carvana, along my time shopping for cars was something like $16,600 -> $11,200 -> $10,600 -> $10,200

To add insult to injury, after I already purchased my car from Carmax, Carvana sends an email asking to buy my car for $10,000.

Carmax only had one change after the offer expirations. $15,000 -> $15,000 -> $16,400

They made me resubmit the trade in offer survey but they raised my trade in. It looks like one of the two companies actually fairly bases their algorithm on market demand. It looks like the other company applies a massive negative penalty to subsequent offers after expiring ones.

If I had to guess why they're like this, I'd think it's: the offer expires a day or two before they're at the dealership viewing the car they want to buy or already paid to ship, they run the offer through the system again, and it comes back low. Because the buyer was already interested, they'll be fine "just" paying an extra $100 or $200 a month.

Carmax also has its issues, the car I received had some cosmetic things in the interior that were not advertised to me (which they are fixing). But at least the pricing algorithm seems fair. A third company appraised me at $15,400, so I feel now like Carvana was trying to subsidize their failing business with desperation.

reddit.com
u/Glittering-Neck-2505 — 18 days ago
▲ 402 r/accelerate+1 crossposts

ARC-AGI 3 is not an honest measure of AGI

I want everyone to take a look at this graph for a second.

ARC-AGI 3 was intentionally not allowing the reasoning agent to maintain its context across actions. It was effectively making the model forget what it had already figured out, over and over again, then scoring that crippled version as if it represented the system’s actual intelligence.

Once OpenAI allowed the agent to preserve its reasoning and compact older context, which is exactly how real world frontier agents work, its score nearly tripled while using far fewer tokens.

Compaction is a basic part of how a real world agent would function. Humans similarly write notes and preserve what they have learned. Nobody would test a human by erasing their memory after every action and then claim the result tells us their true capability.

The reality is that ARC-AGI 3 is not measuring general intelligence. In the real world, if an agent using reasoning and compaction could function in virtually the same way as a human would, that would be called AGI. The already existing agent can do 3x the score while using 6x less tokens, so the benchmark is intentionally dishonest as a measurement of general intelligence. A human is not required to reset its memory each time it starts a new puzzle or moves a piece on a chess board, so this is absolutely egregious in my opinion.

The fact that an AI can do this much better just by remembering what it had already figured out is the true testament to how far in context learning has come. I was already not a fan of ARC-AGI after the quadratic penalty was applied for taking extra steps, but this just confirms my view that this benchmark strayed from the initial goal: measuring general intelligence of frontier models.

We're still going to saturate it anyways, and it's good that there are still tough benchmarks out there, but I just had to share that this is not a good look for this particular benchmark.

u/Glittering-Neck-2505 — 21 days ago

Had my first scary FSD experience right around the 100 mile mark

So I got my 2024 Tesla M3 LR AWD on Saturday. Previously was on 14.2.2.5 and had no issues. Today, I upgraded to 14.3.5 while at the office.

On the way home, things were going pretty smoothly, I was using the hurry mode and it was assertive enough to get around slow drivers + merge into narrow-ish gaps, but also very protective against dumb drivers.

Anyways, on the highway, there was an exit to basically enter another highway. It got in the exit only for a frontage road just ahead of the highway exit. Since the sign was clearly posted, I wanted to see if FSD could figure it out. I wasn't worried about whether or not it caused a reroute because I was not in a time crunch tonight. Anyways, it proceeds to exit on the frontage road but it... continues going nearly 80 mph?

I took over right away when I realized it thought it didn't quickly realize it had taken a wrong exit and reroute, but man. Do not get a false sense of security and do not stop paying attention. It can handle most scenarios, but in situations like this where it has no idea where it is or what speed it is supposed to be driving at, the chances of something bad happening increase significantly. It probably would've been okay as around 5-10 seconds after my intervention it had rerouted and I gave FSD the reigns again.

I do not like FSD less after this event, I'm just conditioning my brain for how to use it responsibly and safely.

reddit.com
u/Glittering-Neck-2505 — 22 days ago
▲ 8 r/codex

21% of my lifetime tokens were spent yesterday

Pro user at $100/month, the weekly limits + 4 banked limit resets (one used so far) has been pretty insane. I'm enjoying the feast while it lasts. If you look up input and output cost per million, this costs like thousands to tens of thousands, I have no idea how this company is even still alive but I'm happy they're giving us an insane amount of free tokens for now.

u/Glittering-Neck-2505 — 1 month ago

How can the speech to text be this bad??

I bought a Claude subscription so that Fable could get an 85% GPT 5.6 Sol project to 100%, and boy do I hate the voice transcriptions. Not only are they super inaccurate, but they also just... stop. After what feels like an outrageously short period of time. I'll swap tabs to my app, visually describe in detail what I don't like, and come back to find it stopped listening 30% of the way through and now I have to start again. Second time it'll get to like 65%. Rinse and repeat. How is this the actual feature they shipped, how do they just stop transcribing, how is a trillion dollar AI lab unable to figure this out. I'm so annoyed now because I think I described it better the first time.

reddit.com
u/Glittering-Neck-2505 — 1 month ago

Price fluctuations on pre-owned inventory

Hello, I was wondering if anyone has had any experience with pre-owned inventory before. Last weekend, my options were amazing. I'm specifically limiting myself to 2024 onward because I want the improved FSD capability. Btw please don't suggest a pre-2024 to save money, I'm not going to take on a car payment only to also have buyer's remorse about it (these are hw3 which gives my FSD a ceiling far below hw4).

Now, all the options are $4k more expensive and there's one that's 2024 with one accident that's $4k more expensive than a pristine 2025 that had only 12k miles was over the weekend (gone now, of course), same trim. My trade-in value also went down $700 when they had me refresh it.

If I saw 4k swings in one direction, has anyone else looked at these before and am I likely to also see those swings in the other direction. I'll join the club if the prices come back down to Earth.

reddit.com
u/Glittering-Neck-2505 — 1 month ago

Notice how the labs were freely letting any person use their frontier AI models and the government is what screwed that up

Before the Mythos and now GPT-5.6 fiasco: the best released models were being served at heavily subsidized rates to anyone in any part of the world, allowing basically anyone who wanted to to benefit from AI and the rapid pace of progress.

Now: access to those best public models is restricted to corporations and corporations alone.

The deeply tech cynical people of Reddit always assumed that it would be the ultra wealthy and giant tech corporations that hoarded AI and its benefits to themselves. They did not. They opened up their wallets and poured billions into making AI smarter and serving it to people at very low cost. They turned agentic AI into a commodity that basically anyone can use in the utterly insane span of just 3 years. I'm not saying that to bootlick or because I think they did it out of kindness, I'm just stating a basic fact that's obvious if you've been following AI.

It was not the wealthy that decided AI was to be hoarded, it was the federal government. Now the wheel of progress has a weight placed on it, slowing it down. And you might say, well *eventually* they'll release it, and maybe for this one it's true. But how long until a model never makes it out of preview because it's "just too powerful" for the wider public to be trusted with?

The market mechanics (namely competition and innovation incentive) that allowed the best AI to quickly disseminate to everyone are now under an existential threat. If they cannot broadly serve these models, the customer base for these best models is going to shrink substantially, and the incentive to make better models and aggressively compete to attract more users is going to begin dying out.

Like a lot of things (housing, nuclear, and construction for example) policy comes in and screws up the mechanics that make everything cheaper. Building a house or a nuclear power plant is now insanely difficult because of zoning laws, permitting, lobbying, and intense regulations. Now Americans are increasingly struggling to buy homes and energy because they never had the chance to become abundant and cheap.

This administration's philosophy is extremely anti-capitalist and suppresses the best aspects of capitalism while promoting nightmarish regulation that makes capitalism not work. It's not hard to imagine a world where AI becomes prohibitively expensive for the average person because regulations just get in the way of what was already happening.

Now I'm fucking depressed, this is absolutely awful news for the future of American AI. Now we may have to rely on open source to kick our ass and sober up this administration to the reality that this is how you thirdworldify a country that was on the verge of the greatest economic achievement any country has made in in human history.

reddit.com
u/Glittering-Neck-2505 — 2 months ago
▲ 103 r/accelerate+2 crossposts

The other leaks didn't do it justice. ChatGPT Bidi-1 can sound scary realistic.

Picture this level of voice control, with two direction realtime interruptions, with memory of your whole chat and saved memories rather than the last few prompts, and with much higher problem solving, prompt understanding, and general intelligence.

Still a little upset we were left with the mess that AVM has become for so long now, but I think this is going to be like maybe the first usable voice mode by any AI company, outside of just being able to produce better noises and voices.

u/Glittering-Neck-2505 — 19 days ago
▲ 723 r/accelerate+1 crossposts

DeepMind is now reportedly struggling to compete with Anthropic and OpenAI while 3.5 Pro is not the step change they'd need to be competitive

Post by a notable but not infallible AI Twitter poster, take with a grain of salt (but we'll know later this month): https://x.com/synthwavedd/status/2068000857757741251

I get the feeling that 3.5 Pro will be very fun to play around with for creative one shots and abysmal for agentic coding. I get that they are prioritizing world models over agentic coding and RSI but still, this is crazy that they're still struggling to catch up after looking set to take a clear lead multiple times now. I find it hard to believe for the same reason there's a consensus here they will win the AI race: they're f'in Google. No one comes close to having that amount of data or infrastructure or free cash flow.

u/Glittering-Neck-2505 — 2 months ago
▲ 159 r/elevotv+1 crossposts

Dario Amodei doesn’t think a red line was crossed if his models were used to commit war crimes, blames war and human judgement

u/Glittering-Neck-2505 — 2 months ago

People still don't understand agentic coding gives you superpowers

To motivate what I'm trying to provide a counterpoint for: I remember when someone made a working Piano learning app that sits on your piano and tracks your performance with Codex and posted it online, AI haters came out in full force. "We already have apps that do that." Completely setting aside that the agent-built piano program is infinitely customizable through updates and free, there are other reasons I find this mindset to be foolish.

I was trying to play a game I've been wanting to get around to this morning, Outer worlds 2, and when I booted it up I was reminded how awful the look acceleration is on controller. I decided that I'll mod it in.

Here's the complication, Xbox game pass for PC make it a huge pain to mod games. However I asked ChatGPT and it seemed to think there was a workable way. I forwarded that info to Codex and asked it to install both mods being mentioned by ChatGPT, basically with nothing except the mod files downloaded. The first implementation did not work. I mentioned the glitch to it, and told it do not come back until the mod is installed.

On it's own, it found a version compatibility issue and wrote a 500 line script to patch the existing mod to be version compatible. Something I could never figure out before I can just have it do.

Anything I want to do, anything that bugs me, any tool I want to have, I can just have it. I wrote a piece of software on my Macbook pro and disables an OS level feature that drove me absolutely crazy with no existing fixes except for buying software. Works flawlessly, never bugs out.

I designed an excel spreadsheet for work that conveniently autopopulates everything it needs to as I fill it out. I designed a custom videoplayer for my job that gives me fine control over video scrubbing, allows file types not natively supported by my default videoplayer, always remembers where I left off, saves files in my desired format, matches and shows specific videos to specific pages in a related pdf report, let's me speed over 2x, and let's me go between videos with a button. No latency, no pop ups, no ads, no data collection, no paid software, no bugs, and no awful Microsoft Onedrive player. It just works, like a hardy reliable tool.

There've been many instances over the past few weeks. I can't stop making software that specifically solves my own headaches. I can imagine that the reign of terror of companies like Adobe will shortly come to an end. Software and at its essence very granular control over the technology we have is becoming democratized for everyone. Software is what gives technology any use at all and the personal computer is about to be on an entirely new level. Now I'm pretty confident that despite these tools, there will be people who never engage with them. They'll use ChatGPT to do the bare minimum slop output for the job they hate, and they will be fine with that. But it's an eye opening experience when you realize there is more than ChatGPT now. Same with Claude Code or Claude Cowork I imagine.

reddit.com
u/Glittering-Neck-2505 — 2 months ago
▲ 495 r/UFOs

This community seems largely uninterested in the biggest disclosure effort in history

I'm sorry, but I have been very frustrated by this community's response to the efforts of the federal government to bring evidence public, and the explicit attempts to discredit or diminish it.

I'm shocked that we just got videos from the Pentagon showing instantaneous acceleration, trans-medium travel, and sharp instantaneous turns without any visible propulsion, wings, rotors, or exhaust and the two most upvoted posts from the last week relate to a floating, rotating inflatable spaceman balloon. Have we completely lost our minds? We're at the tip of the iceberg on previously classified evidence, and we already have videos of the observables that made the 2004 Nmitz tic tac encounter such an intriguing and fantastical witness story.

I get the lack of trust in the government. I do. But this is not something the federal branch has been working to keep from you for 80 years. If there were many people in the know for all that time, surely it would have leaked previously. The government sucks at hiding secrets. But this really only started picking up steam a few years ago when people like Lue Elizondo, David Grusch, and Christopher Mellon started speaking up despite the personal damage it could cause to their reputations and career opportunities.

Their claims are of a small secretive program that is highly compartmentalized and disjointed (meaning people only have knowledge of their specific tasks). Anyone who dared speak up previously (such as Bob Lazar) was easily taken care of by stigma, smear campaigns, and intimidation. This is much more consistent with the pattern of evidence vanishing, people staying quiet, and files sealed ad infinitum over the entire government staging a coverup over 80 years that would have certainly leaked. A classic story of a few really powerful people clinging to power and operating without oversight. I mean if Trump was in the know he certainly already would've ran his mouth about it because that's what he does.

That is to say, you cannot discredit people like the whistleblowers or congress members who are pushing for disclosure by grouping them in with the same category as those staging the coverups and delegating highly secretive programs that are operating unconstitutionally. The difference is that they are putting their careers on the line to get us this information, while those within this legacy program are holding onto the strongest evidence for dear life.

That is to say, you need to be more analytical about what is going on. UFOs are going from a highly stigmatized tinfoil hat phenomenon to one with scientific curiosity, large disclosure efforts, and media attention. This subreddit keeps focusing on the bullshit over the videos that actually show anomalous movement despite those videos confirming some of the details given by whistleblowers. You cannot be dismissive of videos just because they are not 4k. And you also cannot be dismissive just because the first 2 batches of files do not contain definitive proof of alien life.

You're already living through an extraordinary time. Focus on real, verifiable evidence over someone recording a spotlight shining through the clouds. And also open your mind to the fact that these few dozen whistleblowers and handful of congresspeople actually do want you to have the truth that was even hidden from them until they got looped in. There is a real bipartisan effort under way and real videos of actual UFOs being released. Just be open to seeing where it goes and take it one release at a time before dismissing it.

reddit.com
u/Glittering-Neck-2505 — 3 months ago

People are claiming teleop, but I really don't think a human would be this insistent to get a package they clearly can't reach.

Also the movement doesn't look human to me at all, what human is trying to reach for something far away with one arm while keeping the other one completely still (outside of body movement, but the elbow angle doesn't change).

I think people are in denial and really want to believe this is a guy in India controlling it because they're not ready for the day that humanoids take off like the automobile or the iPhone because it's potentially the most disruptive technology we've seen.

EDIT

People in this subreddit: "It's actually teleoperated!"

After showing them it's not: "This is actually shit and not good. Not even AGI, looks like dated tech."

Agree that it's not AGI yet but hear me out. Pretty inconsistent to believe the movements look human enough to imply teleoperation but then believe that this is something we could do 10 years ago.

You have to understand that robots of the 2010s could not generalize beyond a basic task or set of tasks. It is embarrassing this has to be explained. There is no precedent for technology that generates actions from pixels and prompts in real time.

A couple years ago we were limited to text based intelligence with limited image understanding. It couldn't do *anything* in the real world which is why a lot of critics claimed it wasn't close to being AGI. You are literally watching the birth of physical AI and don't care even a little bit. Maybe it is cope, maybe it is a stunning lack of curiosity, but it seems uncharacteristic of anyone who willingly comes to a subreddit called "singularity."

u/Glittering-Neck-2505 — 3 months ago