Interesting. You think it’s better to hide the training data then? Currently in NLP the only chatbots widely deployed (and replacing search) are closed source so the situation is maybe a bit different. I’ll try to take some time write a blog post on this topic one day. It’s a
@thom_wolf
-
Credit vs Monetary Reward in Open Source Communities
By
–
I make a distinction between crediting & monetary reward. I know many people happy to share without monetary reward (the OSS community is all about that) but I know very few people who don't care at all about being credited (even anonymously) for something they've spent time on
-
Need for More Research Beyond Bing Chat
By
–
Huge need for more research here imo. Bing chat is not nearly enough.
-
AI Content Reposting Threatens Creator Attribution and Artistic Value
By
–
There is no excitation in seeing your blog post ideas reposted by a chat AI as his own personal thoughts No excitation in sharing art if the main way to consume it becomes through a white page anonymous art drawing chatbot
-
Knowledge Creator Attribution Crisis in Generative AI Models
By
–
The disconnection between knowledge creators & knowledge consumers in Gen AI models is a pressing problem that's not yet discussed enough Nobody will be excited to share knowledge on the internet if the default way to consume it becomes through anonymizing/unreferencing chatbots
-
Search Engines vs Anonymous AI Training: Creator Economics Shift
By
–
Search engines had the huge advantage that they would point to your personal page where you could get credit/revenue for your ideas, build your own personal brand Writing anonymous tokens for a held-out training dataset is the dream of no stackoverflow or wikipedia contributor
-

StarCoder Fine-tuned as Chat Assistant Using Community Datasets
By
–
So by popular demand, we've starting exploring the finetuning of BigCode's recently released code-completion model StarCoder (
https://
x.com/BigCodeProject
/status/1654174941976068119
…) to be a…. Chat Assistant We finetuned it on two high-quality datasets created by the community:
– OpenAssistant’s -
Data licensing and commercial AI use ethics
By
–
"having all the data" doesn't mean that the people who wrote the data agreed for you to use it commercially and without crediting, e.g. GPL licences
-
Waiting for GPT-4 detailed specifications and capabilities
By
–
Waiting to read the same details about your GPT4