A quick thought experiment: Can we view the taking of AI training data as a modern land grab?
Most of the current debates around artificial intelligence seem focused on copyright issues, job losses, or the disagreements between open-source and closed-source software. But lately, I have been wondering if we might gain some interesting insights by looking at generative artificial intelligence through the classic framework of land, labor, and capital.
I have been sketching out a loose idea about digital land grabbing, and I would love to run it by this group to see if it makes sense, where the weak points are or if I am stretching the metaphor too far.
Here is a simple way we might map Henry George’s factors of production onto the modern artificial intelligence world.
- Labor: The millions of human writers, artists, forum posters, and programmers who spent decades building the open web.
- Capital: The physical and digital tools, large computer systems, data centers, and specialized software programs, built and run by tech companies.
- Land (The Digital Commons): The collective knowledge base created by humans. Much like physical land, real human data is limited in supply (studies suggest we might run out of high-quality public text in the coming years) and hard to recreate (as seen when artificial intelligence models break down after training only on computer-generated data). Its value comes from the public that created it, rather than the company that pays to collect it.
The Parallel to Closing Off Public Land
Historically closing off public land involved putting fences around shared fields and turning them into private property that earns money for an owner.
It feels like we might be seeing a digital version of this process today. Tech companies collect information from the open web for free, hide that collective human work inside private computer models and then charge monthly fees or subscriptions to sell that same knowledge back to the public.
From a certain angle, it feels like a confusion between the value of the soil (the public information created by people) and the value of the tractor (the software and computers). Capital clearly deserves a fair return but I wonder if private companies are also taking unearned land rent in the process.
A Thought Experiment on Policy: A Data Tax and Open Source
If we look at it this way just for the sake of discussion, what kind of policy response might make sense, setting aside standard copyright battles? Here are a few simple ideas to think about.
- A Tax on Private Data Use: A tax on private artificial intelligence models based roughly on how much public data they gathered without paying for it.
- A Public Data Dividend: Directing that tax money into a public dividend fund, recognizing that the underlying information belongs to everyone.
- An Open-Source Exception: If a developer shares their models and training data openly, they pay no tax at all, essentially returning the digital land back to the public for anyone to use. You would only tax companies that build a fence around the data.
I would love to hear your thoughts, critiques or alternative ideas! :)