Love it. Yeah I keep going back and forth between fully community driven (everybody throws in any embeddings) vs community organized (make set of categories, work on embedding all of it together methodically).
@yoheinakajima
-
Third-party system for platform data import export
By
–
Ohh I like the idea of a third party system that other platforms can import/export out of
-
Automating Tax Processes With Secure Personal Data Storage
By
–
The next step is not needing find, copy, and paste the tax code. Add a safe way to store and share personal information and all of this gets automated!
-
Pinecone Hybrid Search for Few-Shot Examples Database
By
–
Ooh both of these are great. Pinecone hybrid seems like a good approach for managing scale (unverified comment). Yours is interesting, is the purpose here for having a public database of few-shot examples?
-
Open Standards for AI Models: Navigation and Implementation Challenges
By
–
Hm, yeah an open standard seems to make sense. Though seems complicated to navigate which models to use etc, though worth it IMO
-
Decentralized AI infrastructure with working group governance
By
–
Oh awesome. Random next Q: should this be built with redundancy and no reliance on a single corporate entity? Maybe a working group makes sense.
-
Public Data Embedding Efficiency and Database Solutions
By
–
Hm good point to consider. I’ve seen at least 5 @paulg bots where they embedded his essays, and @gpt_index will chunk up and embed websites on the fly. I don’t have a clear answer, but this seems inefficient, at least for publicly available data – I’d love a public database of
-
Automatic multi-layer clustering with continuous user feedback optimization
By
–
Hm good point. How about automatic clusters embedded regularly, multi-layer as clusters grow, with constant testing and iteration through user feedback and observed speed.
-
Public Embedding Database with Monetization Model
By
–
A lot of people are embedding the same publicly available content. Would be fascinating to have a public database where you could contribute your embeddings that anyone could query. Include source URL and text as standard metadata. Maybe it’s one cent per 1000 calls to this