-LAION & the copyright crisis: @Lawgeek -iNaturalist & how communities make data: @JerThorp -Making See:Set and the critical study of data:
@christo_buschek
DATA
-
Data, Copyright, and Community: Critical Perspectives on AI Training
By
–
-
The Exponential Growth of AI Training Datasets Over Two Decades
By
–
Datasets are now impossibly large. Back in 2003, Caltech 101 was 10,000 images. By 2010, ImageNet had 14m images. In 2023, LAION has 5 billion images and text captions – it'd take you 1000+ years to look at it all. The entire territory of the internet has become the map of AI.
-
See:Set: New Tools for Analyzing Dataset Patterns and Worldviews
By
–
That doesn't mean we stop analyzing them – we just use new tools. At Knowing Machines we've developed See:Set, a search engine to look into datasets and track patterns. These 9 essays come from direct experiences using See:Set to understand how datasets encode worldviews.
-
9 Ways to Study and Analyze Datasets: Research Findings
By
–
New release: '9 Ways To See A Dataset" https://
knowingmachines.org/publications/9
_ways_to_see_a_dataset
… This is a 9-part series about how to study datasets – and what we found. Dives into individual datasets like LAION, ImageNet, NYT Corpus, C4, NABirds & iNaturalist. But studying datasets is harder than ever… -
AI Ethics Experts Lack Technical Understanding Says Critic
By
–
I believe in ethical standards but @pierrepinna unfortunately there some of the “AI ethics experts” who don’t understand the underlying tech nor understand the role and nature of data.
— AI (@DeepLearn007) 12 juillet 2023
This has ranged over the years from some AI ethics experts making calls and demands on the… https://t.co/spWWuSh4YFI believe in ethical standards but @pierrepinna unfortunately there some of the “AI ethics experts” who don’t understand the underlying tech nor understand the role and nature of data. This has ranged over the years from some AI ethics experts making calls and demands on the
-
Neon Database Vector Analysis Compared to pg_vector
By
–
What I really like about the analysis the @neondatabase team did is that they did a pretty detailed comparison to `pg_vector`. s/o @raoufdevrel for this They both have their pros and cons! @andrewkane @kiwicopple would love your thoughts here as well
-
PostgreSQL HNSW Extension for Vector Embeddings by Neon
By
–
pg_embedding @PostgreSQL is popular database choice. Excited to share a new extension from @neondatabase to help you use it for embeddings as well (with HNSW)! Blog:
-
Area Chair Award: Machine Translation with Minimal Data for Linguistic Diversity
By
–
Area Chair Award on Linguistic Diversity
Small Data, Big Impact: Leveraging Minimal Data for Effective Machine Translation https://
bit.ly/3POTFbl -
PropSegmEnt: Expert-Annotated Dataset for Textual Attribution
By
–
PropSegmEnt is a large expert-annotated dataset for for realistic textual attribution, claim provenance, and more. Drop by the @aclmeeting Google booth at 12:30 today to learn about it from authors Alex Fabrikant & @soshsihao
. -

Future of Decision-Making: Machines and Human Collaboration
By
–
What will the future of decision-making look like? Only 4% think machines will make decisions alone. Explore and learn more in our report. Download now: https://
ow.ly/P3Y250P87Vr #Analytics #AnalyticsInsights #DataDriven #AI #AlteryxAiDIN #DecisionIntelligence