As AI systems learn to hold a conversation by studying places like Reddit, Nicole Junkermann asks what is owed to the people who wrote all those words.
Nicole Junkermann
September is back-to-school season, and this year one of the most diligent students on the planet is not a person at all. While students everywhere sharpen their pencils, the artificial intelligence systems millions of people now use every day have been doing their own kind of studying, and a surprising amount of it has taken place on Reddit. In 2026, much of that learning is increasingly put on a licensed footing, with major online platforms reported to license their archives of conversation to the companies building AI, and Reddit has become one widely reported example. Nicole Junkermann sees Reddit as one of the clearest and most talked-about cases of that shift. As an investor who has spent years around artificial intelligence, Nicole Junkermann thinks the story of how these systems learned to sound so human runs straight through the internet’s messiest and liveliest conversations. It is worth understanding, Nicole Junkermann argues, because it quietly changes how you think about the tools in your pocket. It is a theme she explores on her podcast, the AI Overview, in an episode on what executives should know about generative AI.
Why Nicole Junkermann says Reddit is such rich material
To understand why a site like Reddit matters so much to AI, you have to think about what these systems actually need. A model that answers questions and holds a conversation learns by reading enormous amounts of human writing, and not just any writing. It learns best from real people talking to each other: asking questions, giving answers, arguing, joking and explaining niche hobbies in patient detail. Reddit, with its thousands of communities on almost every subject imaginable, is one of the largest collections of exactly that kind of natural human back-and-forth ever assembled. For a machine trying to learn how people really talk, Nicole Junkermann points out, it is close to an ideal textbook.
There is a second thing that makes it valuable. On Reddit the community itself votes answers up or down, which quietly labels which replies people found useful and which they did not. That built-in signal of quality is rare and genuinely precious for anyone training a system to tell a good answer from a weak one.
For a machine trying to learn how people really talk, it is close to an ideal textbook.
NICOLE JUNKERMANN
The part Nicole Junkermann thinks is worth sitting with
Here is where the story gets more interesting, and more human. Every one of those helpful answers, every patient explanation and late-night debate, was written by a person. None of them wrote it in order to train an AI. They wrote it to help a stranger, to settle an argument, or to share something they loved. Increasingly, the use of that writing is put on a formal footing through licensing, as platforms agree paid deals to license their data to AI developers. That, Nicole Junkermann notes, is a positive move toward putting these arrangements on a clearer and more transparent footing.
Still, the question she keeps coming back to sits one level down. A licensing deal is struck between a platform and an AI company. The volunteers who actually wrote the answers, whose curiosity and generosity made the archive valuable in the first place, are rarely party to it. She is careful not to pretend there is a tidy answer, and she is clear that these are her own opinions rather than settled facts. But she argues that the firms building these systems, and the platforms licensing the data, would do well to treat the communities behind it as partners to be respected, rather than simply a resource to be monetised. The value did not come from nowhere. It came from people. Nicole Junkermann is careful to stress that Reddit is only one well-known example among many, and that she intends this as a broad observation about the industry rather than a complaint about any single company.
Nicole Junkermann on why permission comes first
On one point Nicole Junkermann is unambiguous. Anyone hoping to use a platform’s data or its interfaces for training an AI should secure that platform’s permission before doing anything at all. In the case of Reddit in particular, she believes it is very important to obtain Reddit’s permission before using its API for AI training, and that this should never be done without Reddit’s express written permission. The responsible route, in her view, is a clear written agreement on terms that Reddit has knowingly granted, rather than any assumption that because something is visible in public it is free to be used for training. Permission first, she says, and in writing, is not red tape. It is the difference between building on a foundation of respect and building on one of resentment.
Where Nicole Junkermann lands
None of this makes Nicole Junkermann gloomy about AI, a technology she backs and believes in. Nicole Junkermann does not single out Reddit or any other platform, and speaks warmly of the value such communities create; her interest is in the whole model of how these systems learn. If anything, she says, understanding where these systems learn should make people more appreciative of them, and a little more thoughtful about them too. The next time an assistant hands you a genuinely useful answer, it is worth remembering that somewhere underneath it are countless real people who once took the time to explain something to a stranger on the internet.
Her bigger point is about trust, a theme Nicole Junkermann returns to often. A technology is only as trustworthy as the care taken in building it, and that includes the care taken over where its knowledge comes from. Treating the human sources of that knowledge with respect is not only the decent thing to do. In her view, it is part of what earns these tools a lasting place in people’s lives.
Nicole Junkermann is an entrepreneur, venture capital investor and philanthropist with a long-standing interest in artificial intelligence and deep technology, and the host of the AI Overview podcast.
DISCLAIMER
This article is independent commentary. Nicole Junkermann and this piece are not affiliated with, endorsed by, sponsored by, or connected to Reddit or any other company referenced. Reddit is a trademark of its respective owner, used here for identification and commentary only. Reddit is referenced as a widely reported public example, and nothing in this article alleges any wrongdoing by Reddit or by any other company.
The views expressed are the author’s own opinions and are provided for general information only. Nothing here is legal, financial, investment or professional advice, and no particular outcome is implied or assured. References to how AI systems are trained describe the author’s general understanding and may not reflect the practices of any specific company. This article does not encourage or endorse accessing or using any platform’s API or data for AI training without that platform’s express written permission.






