Use an appropriate license compatible with Fair Use. LAION's creation of the dataset required one or more Copyright Exceptions with strict terms. Picking a license that respects the terms of these exceptions would reduce legal risk significantly, and show good faith.
OPEN SOURCE
-
LAION criticized for partnering with questionable data hosting site
By
–
Don't partner with dubious websites to host files! LAION shouldn't have requested TheEye, a questionable "archival" website, to host their data in order for them to link to it! Maybe instead get approval from platforms like HF so they share legal responsibility for hosting.
-
Respecting Manual Opt-Out Requests After Fact
By
–
Provide opt-out after the fact. To operate in good faith, manual opt-out requests should also be respected after the fact too. Because the dataset format allows users to bypass and override opt-out client side, not honoring request makes LAION look complicit in infringement.
-
Copyright Compliance: LAION Must Prove File Deletion Records
By
–
Publish code & logs upfront show files were deleted. Complying with Copyright exceptions means the burden of proof is on LAION to show they qualify. They're required to delete the files so keeping records of this (or publishing them upfront) could help avoid court.
-
Software Tool Recommendations for Robots.txt Compliance Checking
By
–
Recommend software that checks robots.txt. One of LAION's members built img2dataset, which does not check robots.txt. It would be a stronger legal argument if this recommended tool did check at least robots.txt now, even if better forms of opt-out are added later.
-

LAION-5B Legal Compliance: Dataset Challenges and Solutions
By
–
Legal uncertainty is all around the AI industry, in particular the LAION-5B dataset. Despite the RELAION update, there are still many problems to be addressed… (My peer review follows.) Here's what LAION could have done to make their dataset more legal and more certain!
-
FineVideo Dataset Released for AI Model Training
By
–
The dataset is here: https://
huggingface.co/datasets/Huggi
ngFaceFV/finevideo
…. Great job @micuelll and team! -
Hugging Face Releases FineVideo Dataset for Advanced Multimodal Analysis
By
–
Hugging Face releases FineVideo 🚀🚀🚀
— abhishek (@abhi1thakur) 12 septembre 2024
Check out this groundbreaking dataset, built for advanced video analysis in multimodal settings:
• 43,751 videos across 122 categories with over 3,425 hours of content 🤯
• Focuses on mood analysis, storytelling, and media editing
•… pic.twitter.com/dp62jHPhYFHugging Face releases FineVideo Check out this groundbreaking dataset, built for advanced video analysis in multimodal settings:
• 43,751 videos across 122 categories with over 3,425 hours of content • Focuses on mood analysis, storytelling, and media editing • -

DataGemma: Open Models Grounding LLMs in Real-World Data
By
–
Announcing DataGemma, a set of open models that utilize Data Commons through Retrieval Interleaved Generation (RIG) & Retrieval Augmented Generation (RAG) to ground LLMs in real-world data for fact-checking, responsible AI development & more. Read more at https://
goo.gle/datagemma -

Hugging Face Releases FineVideo Open-Source Dataset
By
–
Did you notice how most video AIs are commercial, closed-source & secret? It sucks and a big reason for the lack of open-source & transparency is that there are very few high-quality open video datasets. The @huggingface team is changing this today by releasing FineVideo, a