๐ Reddit, Inc. tried to shut her down but she flipped the game. ๐ช

Had the fortune of listening to Anna Kazlauskas from OpenDataLabs & Vana at the Flower Labs Summit today. She dropped a vision so inspiring, it is basically mandatory to share. Check this out ๐
๐ The Glass Ceiling
- The big shots - OpenAI, Meta, Google, X - are hitting a wall. Clean public data is drying up fast - 15 trillion tokens, mostly tapped out. Check graph below
- Their move? Cozying up to publishers for new content behind paywalls. No wonder companies like Reddit, Inc. and Photobucket made it harder to crawl their webpages and raised the pricing for their APIs
๐ก Anna's Insight
- Public internet data is just a tiny slice - think ~5% of what's out there. The rest? Locked behind paywalls or inside personal vaults (your emails, messages, notes).
- The real data doesn't belong to the companies - it's YOURS, thanks to user rights from regulators. Solve the coordination problem and you have more data than any single company ever could.
๐ The Vision
Why not let users contribute their own data to create an LLM that really works for them?
- Vana wants to train the first user owned LLM trained on user consented data contributions that are collected through Data DAOs.
- Users get their copies of their data from the walled gardens (as is their right) and contribute it to training the LLM on their own terms and get to share the revenues generated when the models create new value in the world.
๐ ๏ธ How It's Made Possible
- Users form Data DAOs (Reddit, Tesla, Google, 23AndMe) pool contributions in a Trusted Execution Environments (TEE) encrypted by a key that is owned and governed collectively by all contributors.
- Federated learning frameworks tie it all together. Enables learning from all DAOs simultaneously without centralizing it in one place.
Hardik Katyarmal
Writes on data strategy, privacy-preserving tech, and ecosystem intelligence at LattIQ.

