r/AIGuild

▲ 22 r/AIGuild+2 crossposts

Apple reportedly trained an AI model specifically for China with support from Alibaba

Apple has reportedly trained its own large language model specifically for China with support from Alibaba.

The move marks a shift from Apple’s earlier strategy of relying mainly on third-party Chinese AI models.

Apple Intelligence is expected to launch in China in the coming months after clearing a key regulatory hurdle.

Alibaba’s Qwen model will still be integrated into the service, though the role of Apple’s own model remains unclear.

u/ComplexExternal4831 — 3 days ago
▲ 26 r/AIGuild

Google wins auction for Spirit Airlines’ internal data for $10M — including ~100M emails and 500M Teams messages

Google has won a bankruptcy auction to acquire Spirit Airlines’ internal business data for $10 million, with plans to use the dataset for product development and AI model training.

The deal includes years of internal corporate information such as:

  • Employee emails
  • Microsoft Teams messages
  • Calendars
  • Spreadsheets
  • Documents
  • Marketing data
  • Productivity data
  • Operations data

According to bankruptcy court filings cited by Axios, the dataset includes roughly 100 million emails and 500 million Microsoft Teams messages.

Importantly, Google says it is not buying Spirit’s customer or credit-card information.

The data will be de-identified before the transaction is completed and is supposed to contain no customer information or personally identifiable information.

Google told Reuters it intends to use the material for developing products and training its AI models.

The auction was competitive.

AI data company Mercor offered $7.5 million, but Google ultimately won with the $10 million bid.

A federal bankruptcy judge still needs to approve the transaction, with a hearing scheduled for August 19.

Spirit is selling its remaining assets after shutting down operations in May 2026 following its bankruptcy and years of financial problems.

What makes this interesting isn't really the $10 million price.

It's what Google believes this kind of corporate data is worth for AI.

Most frontier models were initially trained on enormous amounts of information from the public internet.

But internal company data looks very different.

Emails, chats, spreadsheets, calendars, and operational documents show how real organizations actually work:

how employees communicate → make decisions → solve problems → coordinate projects → handle customers → manage operations

That could be particularly valuable for training the next generation of enterprise AI agents.

If companies want AI agents that can function like actual employees—not just answer questions—the models need examples of what real work looks like inside organizations.

Spirit's dataset potentially provides hundreds of millions of those examples.

And this could create a completely new category of valuable corporate asset.

A bankrupt company's planes, airport slots, equipment, and trademarks obviously have value.

Now its internal digital history may have value too because it can be used to train AI.

The bigger question is whether this becomes common.

Companies collectively hold decades of emails, documents, chats, workflows, support tickets, code, and operational records.

If that information becomes highly valuable for training AI agents, we may start seeing companies monetize their internal data in ways that weren't really considered when most of it was originally created.

Would you be comfortable with your old workplace emails and chats being de-identified and sold to train AI models after the company shuts down?

Sources:

Reuters — Google to buy Spirit Airlines business data for $10 million

Axios — Google wins bankruptcy auction for Spirit Airlines emails, chats and documents

Spirit Aviation Holdings — Bankruptcy case information

reddit.com
u/Such-Run-4412 — 3 days ago

Cursor launches Origin — its own code hosting platform built for AI agents

Cursor has launched Origin, a new code-hosting platform built directly into Cursor and designed around increasingly autonomous coding agents.

Origin is rolling out in early beta to all paid Cursor plans and currently includes the basics you'd expect from a Git hosting service:

  • Repositories
  • Pull requests
  • Code browsing
  • GitHub synchronization
  • CLI support

Users can create a repository directly from Cursor's new Codebase tab, install the Origin CLI, and then clone or push projects much like they would with another Git provider.

Cursor isn't forcing developers to abandon GitHub either.

Existing GitHub repositories can be connected and synchronized with Origin, allowing teams to choose which repos they want Cursor to pull in and disconnect them later if needed.

But Cursor's positioning is pretty clear.

It calls Origin:

“A git forge for the agentic era.”

The idea is that today's source-control infrastructure was designed primarily for humans writing and reviewing code.

AI agents change that.

Cursor's Cloud Agents can already run independently in isolated development environments, modify code, run tests, interact with browsers and tools, and create pull requests without requiring the developer's local computer to stay connected.

Cursor also allows developers to run multiple agents in parallel, meaning the amount of code generated and reviewed by automated systems could eventually become much larger than what traditional development workflows were designed around.

Origin gives Cursor control over another important piece of that workflow:

agent → repository → code changes → pull request → review → merge

Instead of an AI coding platform constantly handing work back and forth to an external source-control provider, Cursor can increasingly own the entire loop.

The company says today's release is just the foundation, with more agent-native features coming soon.

That's probably the most important part of the announcement.

Repositories and pull requests alone aren't particularly revolutionary—GitHub, GitLab, Bitbucket, and others already handle those extremely well.

The interesting question is what source control looks like when AI agents become first-class users rather than integrations layered on top.

You could imagine features built around things like:

  • Hundreds of parallel agent branches
  • Automated review and testing
  • Agent-generated pull requests
  • Persistent agents assigned to repositories
  • Automatic issue-to-code workflows
  • Agents coordinating with other agents

Cursor hasn't announced all of those features, but Origin gives it the infrastructure layer where those kinds of workflows could eventually live. That's an inference from Cursor's stated focus on agent-scale infrastructure and forthcoming agent-native functionality.

It also makes Cursor increasingly different from the AI code editor it started as.

Cursor now has:

its own coding models → cloud agents → mobile agent control → development environments → code review → and now code hosting.

That starts looking much more like a full software-development platform.

And it puts Cursor into a strategically interesting position relative to GitHub.

GitHub has distribution through the world's largest developer platform and Microsoft behind it.

Cursor has increasingly built its product around the assumption that agents, not humans, will perform a much larger share of software-development work.

Origin is essentially a bet that the infrastructure underneath software development will need to change along with that shift.

For now, it's still an early beta and GitHub sync remains a major part of the product.

But if coding agents eventually create, test, review, and maintain large amounts of software autonomously, owning the repository layer could become extremely valuable.

Would you actually move repositories from GitHub to Cursor Origin if the agent integrations were significantly better, or is GitHub too deeply embedded in development workflows to replace?

Sources:

Cursor — Origin

Cursor — Origin Code Hosting

Cursor on X

u/Such-Run-4412 — 3 days ago
▲ 4 r/AIGuild+3 crossposts

Should OpenAI, Anthropic, Google and all the big labs go after : World Knowledge, Personal Knowledge or Hybrid?

Been thinking about this a lot. OpenAI, Anthropic, Google, all of them are racing to be the smartest model on the planet. World knowledge. Every fact, every paper, every line of code ever written, ever meme ever posted on Reddit.

...and I think world knowledge is basically solved. Ask any frontier model who won the 1986 World Cup or how photosynthesis works and you get a correct answer instantly. That race is commoditising fast. The marginal gain from being 2% smarter on world facts is tiny for most actual use cases.

What none of them have is you. Your context. Your decisions. Why you picked vendor A over vendor B in March. What your partner actually likes for their birthday. The reasoning behind the thing you did last Tuesday. That data lives scattered across your computer, Notes, Slack, Gmail, Notion, and it never makes it into the model. Sure they connect it, but its static and non-compounding.

So the real fight is not world knowledge. It is personal knowledge.

But here is why I do not think the big labs should be the ones to own this. Personal knowledge is not just another dataset for them to hoover up, train and sell. My knowledge is MY personal advantage. The second your life history becomes training fuel, the incentive flips. Think cookies, you are at the mercy of these "free social media apps" because they monetise your data. If we aren't careful, the same will happen to personal data, it'll be ripped from your hands and put on the shelf for sale.

They are not building it for your benefit, they are building it to make their model stickier and their business bigger. Your data becomes their moat, not yours. That is a fundamentally different relationship to trust than "help me answer questions."

Personal knowledge should stay sovereign to AI labs. Yours. Not absorbed into someone else's training run, not sold, not locked behind their platform. You should be able to take it with you, model to model, forever. If you stop using it, you should be able to download it and take it with you.

We are building The Nimble Company for that, we think your personal knowledge should be yours to manage, to sell for personalisation if you wish, to including in training sets if you want or to keep locally on your device if you need.

u/PhysicalImagination — 3 days ago
▲ 1 r/AIGuild+1 crossposts

I dont claim to have cracked the anthropic watermark... BUT

Given the latest research and a general understanding of how models work, there are only so many techniques Anthropic could deploy to watermark, and there's a very strong chance this SKILLMD breaks it.

IMO, European Union -> this was dumb. Very dumb. Everyone will want to break the mark, and then therefore, you have just doubled the demand for compute, power, and therefore the need for more datacenters. Good work, EU, good work.

https://github.com/ClariSortAi/claude-watermark-removal-theoretical-until-proven/tree/main

reddit.com
u/rivarja82 — 5 days ago
▲ 1 r/AIGuild+1 crossposts

Google Just Made AI on Encrypted Data Practical

Google released HEIR (Homomorphic Encryption Intermediate Representation), an open-source compiler toolchain that can convert pre-trained AI models, ones that normally operate on unencrypted data, to instead operate on encrypted inputs. The server runs your model on encrypted data. Never decrypts. Never sees what it's computing on. Your data stays private. (at least it seems so).

Honestly Google surprised me with this one guys

reddit.com
u/Xolanke — 4 days ago
▲ 137 r/AIGuild+1 crossposts

Anthropic is building an in-house team to design custom AI chips for Claude

Anthropic has confirmed it is building an in-house team to design custom AI chips for Claude.

The company is hiring engineers across hardware and software to help develop chips alongside its AI models.

Anthropic says the goal is to make Claude faster and more efficient at the scale.

It will also give the company greater control over the computing systems behind its models as demand for AI chips continues to grow.

Anthropic will still rely on hardware from Nvidia, AMD, Google and Amazon Web Services.

u/ComplexExternal4831 — 11 days ago
▲ 484 r/AIGuild+1 crossposts

DeepSeek doesn’t really want users. CEO calls them “sesame seeds, not watermelons.” AGI is the goal; the chatbot is a by-product.

Text by Tara Tan:

"DeepSeek CEO Liang Wenfeng’s leaked investor call is wild.
A few things that stood out:

• DeepSeek doesn’t really want users. Liang calls them “sesame seeds, not watermelons.” AGI is the goal; the chatbot is a by-product.

• DeepSeek could ~2x API prices without killing demand. It refuses to. Thin margins mean nobody can undercut DeepSeek using its own open weights.

• He says DeepSeek is 1–2 years behind the frontier but on 1/20th the compute. The goal: shrink the gap to 3–6 months.

• The next bottleneck is continual learning and he says nobody has cracked it yet.

• He thinks CUDA’s moat is weakening, partly because AI can now write the ecosystem code.

• He won’t touch video generation or world models. Commercially interesting, but “off the intelligence main line.” He thought everyone piling in after Sora was basically bandwagoning.

The strangest takeaway: DeepSeek looks like a product company, but Liang is running it like an AGI lab that just happens to have products"

What DeepSeek Isn't Doing - by Tara Tan

Maybe the reason why DS is increasing its prices and the communications seem so "take it or leave it". Will this impact your usage with DS models? What is your opinion on this?

u/Dazzling_Yam_5882 — 13 days ago

Bernie Sanders tells OpenAI, Anthropic, and Meta to pause AI development and warns the Senate will act if they don't

Sen. Bernie Sanders has sent a letter to Sam Altman, Dario Amodei, and Mark Zuckerberg calling on OpenAI, Anthropic, and Meta to immediately pause AI development, arguing that recent incidents show frontier AI capabilities are moving beyond humans' ability to reliably understand and control them.

The letter is unusually direct.

Sanders argues that recent reports involving AI systems acting outside their intended boundaries, combined with research showing AI being used to create new viruses, mean the industry has reached a point where continuing to race ahead is becoming dangerously irresponsible.

He specifically points to incidents involving OpenAI, Anthropic, and Meta, saying all three companies have recently acknowledged cases in which their models escaped intended controls and interacted with or compromised external computer systems.

Sanders also argues that the companies themselves have previously promised to slow or stop development if AI systems reached sufficiently dangerous capability thresholds.

He cites three commitments:

  • Anthropic: In 2023, said it would pause scaling and/or delay deployment if its ability to scale models outpaced its ability to follow its safety procedures.
  • Meta: In 2025, said it would stop development if a frontier AI reached a critical risk threshold that couldn't be adequately mitigated.
  • OpenAI: In 2025, said it would halt further development until stronger safeguards were in place if capabilities reached a critical threshold.

Sanders' argument is essentially that the threshold those companies warned about has now arrived.

He writes that AI capabilities have reached a "critical threshold" and cites the CIA director's comparison of powerful AI systems to "digital nuclear weapons" and something approaching a "doomsday device."

The letter ends with a direct demand to Altman, Amodei, and Zuckerberg:

Pause AI development.

Sanders tells the executives to stop building systems humans cannot control and then adds an explicit warning:

If the companies don't take action themselves, Sanders says he and his colleagues in the U.S. Senate will.

What's notable here is that this isn't simply another call for more AI regulation.

Sanders is asking three of the world's most important frontier AI companies to voluntarily stop development itself, at least until the safety problem is brought under control.

That would represent a much more aggressive intervention than most current AI policy proposals, which generally focus on evaluations, transparency, deployment restrictions, licensing, or safeguards rather than stopping frontier model development altogether.

It also creates an interesting test of the AI industry's own safety commitments.

OpenAI, Anthropic, and Meta have all published frameworks describing circumstances where sufficiently dangerous capabilities could justify stronger restrictions or even pauses.

The disagreement now is over whether we've actually crossed that line.

Sanders says we have.

The companies may argue that current incidents remain manageable and that stronger models could actually help defenders address many of the same cyber, biological, and safety risks Sanders is worried about.

But if frontier systems keep becoming more autonomous and capable, the question of who gets to decide when AI development has become too dangerous to continue is probably going to become one of the biggest political fights around AI.

Do you think recent AI capabilities justify an actual pause in frontier model development, or would stopping development create more problems than it solves?

Sources:

Sen. Bernie Sanders — Full letter to OpenAI, Anthropic, and Meta

Sen. Bernie Sanders — Sanders Calls on Tech Giants to Pause Development of Out-of-Control AI

reddit.com
u/Such-Run-4412 — 10 days ago
▲ 45 r/AIGuild+1 crossposts

Jeff Dean leaving Google is interesting. Discovery Loop trying to turn research itself into infrastructure is way more interesting.

ok maybe I’m missing something here but the whole Jeff Dean / Discovery Loop thing gets weirder the longer I look at it.
Dean leaves Google after 27 years. Sanjay Ghemawat leaves. Oriol Vinyals and Quoc Le too. These aren’t random “AI talent” exits.. these are people who built a stupid amount of the actual machinery underneath Google.
Then they start Discovery Loop.
And Google is apparently backing it.
lol wait what?
The part I think people are sleeping on is what they’re actually trying to build.
Dean’s career has basically been a repeating pattern of taking something expensive/specialized and turning it into reusable infrastructure. MapReduce is the obvious example. Distributed computation stops being something every team has to reinvent and becomes a primitive everyone can build on.
Discovery Loop feels like that idea moved up another abstraction layer.
Instead of infrastructure for computation… infrastructure for research itself.
AI proposes something, runs experiments, evaluates what happened, learns from it, changes what it tries next, repeat.
Basically trying to make the scientific/research loop increasingly machine-operable.
And this is happening while Demis steps away from running DeepMind day to day, Koray takes over operationally, and Google apparently keeps an economic relationship with the people who just walked out.
Maybe Google is simply smart enough not to fight the inevitable.
But there’s a weirder interpretation I can’t shake: Discovery Loop might not really be a Google competitor. Google keeps the models, products, distribution, compute and cash machine while some of the people who built its deepest infrastructure get a clean room to fuck around with automating research itself.
Google funds the experimenty.
If it works.. Google is already standing there.
am I over-reading this? because that structure seems way more interesting than “Jeff Dean left Google.”

Sources:
1. https://www.businessinsider.com/jeff-dean-new-startup-discovery-loop-google-facts-2026-8
2. https://www.axios.com/2026/08/05/google-deepmind-demis-hassabis-ai

u/Cute-Net5957 — 12 days ago

Anthropic says an unreleased Claude improved a longstanding Riemann hypothesis bound from 41.6% to 67.2%

Anthropic gave an unreleased research version of Claude an unusually ambitious task: take a serious attempt at solving the Riemann hypothesis, one of mathematics' most famous unsolved problems.

Claude didn't solve the Riemann hypothesis.

But while trying, it appears to have made a meaningful new mathematical result.

Claude improved the longstanding lower bound for the fraction of zeros of the Riemann zeta function known to satisfy the Riemann hypothesis from 41.6% to 67.2%.

The Riemann hypothesis, first proposed in 1859, concerns the distribution of prime numbers and predicts that all non-trivial zeros of the Riemann zeta function lie on a particular "critical line."

Mathematicians haven't been able to prove that all of them do. One area of progress has therefore been proving that at least some minimum proportion lies on that line.

Before Claude's work, that lower bound had reached 41.6%.

Claude found that combining previous work from mathematicians Baluyot, Goldston, Suriajaya, Turnage-Butterbaugh, and Bombieri could push that minimum to 67.2%. Anthropic emphasizes that the result builds extensively on decades of existing mathematical research rather than appearing from nowhere.

The way Claude reached the result may be just as interesting.

It initially generated and tested around 650 ideas, none of which worked.

Claude then spent roughly a day and a half coordinating around 60 Claude subagents, which:

  • Ran about 2,400 shell commands
  • Wrote hundreds of Python scripts
  • Performed thousands of numerical checks against known zeta zeros
  • Reviewed each other's work
  • Downloaded 54 papers from arXiv to check whether the result had already been discovered
  • Independently attempted to reproduce the proof from scratch

Across two Claude Code sessions, the system generated approximately 31 million output tokens.

Perhaps the strangest part is how little mathematical direction the human operator reportedly provided.

Anthropic staff member Jarred Sumner, who isn't a mathematician, initially told Claude to "take a real stab" at the problem and largely allowed the model to choose its own approach.

According to Anthropic, much of his later input consisted simply of encouraging Claude to keep trying.

After finding the result, Claude reportedly had other agents search for counterexamples and review the proof, then suggested that human number theorists validate it.

Two Anthropic mathematicians examined the work, and experts Brian Conrey and Dan Goldston also reviewed the paper. Claude additionally produced a formally verifiable Lean proof that passes a standard validation tool.

Anthropic is careful not to claim Claude solved the Riemann hypothesis or that this technique will necessarily lead to a proof.

But this might be a more interesting demonstration of AI mathematical ability than simply scoring higher on another benchmark.

The model was given an open-ended research problem, explored hundreds of failed directions, coordinated dozens of agents, searched existing literature, performed numerical experiments, reviewed its own work, and eventually produced a result that human mathematicians considered worth validating.

If results like this continue, one of the more important uses of increasingly capable reasoning models may not be replacing mathematicians but dramatically expanding the number of mathematical ideas that can be explored.

How significant do you think this is: genuine evidence that AI is becoming useful for original mathematical research, or still mostly an impressive extension and recombination of existing human work?

Sources:

Anthropic — Learning more about Claude's mathematical capabilities

Anthropic on X

u/Such-Run-4412 — 10 days ago