OpenAI says it will follow public opinion on whether GPT4 should include face, emotion and gender recognition. Which public? In which nations? Whose opinions will count?
@katecrawford
-
Data, Copyright, and Community: Critical Perspectives on AI Training
By
–
-LAION & the copyright crisis: @Lawgeek -iNaturalist & how communities make data: @JerThorp -Making See:Set and the critical study of data:
@christo_buschek -
Amazing team of scholars creating datasets and visualizing bird data
By
–
It's from an amazing team of scholars, coders & artists: -NYT & institutional data: @ananny -How to see birds: @JerThorp -C4 & the making of datasets: @_will_orr
-Unstable categories & biodiversity: @Hamsini_S
-The trouble with ImageNet: @SashaMTL -
See:Set: New Tools for Analyzing Dataset Patterns and Worldviews
By
–
That doesn't mean we stop analyzing them – we just use new tools. At Knowing Machines we've developed See:Set, a search engine to look into datasets and track patterns. These 9 essays come from direct experiences using See:Set to understand how datasets encode worldviews.
-
The Exponential Growth of AI Training Datasets Over Two Decades
By
–
Datasets are now impossibly large. Back in 2003, Caltech 101 was 10,000 images. By 2010, ImageNet had 14m images. In 2023, LAION has 5 billion images and text captions – it'd take you 1000+ years to look at it all. The entire territory of the internet has become the map of AI.
-
9 Ways to Study and Analyze Datasets: Research Findings
By
–
New release: '9 Ways To See A Dataset" https://
knowingmachines.org/publications/9
_ways_to_see_a_dataset
… This is a 9-part series about how to study datasets – and what we found. Dives into individual datasets like LAION, ImageNet, NYT Corpus, C4, NABirds & iNaturalist. But studying datasets is harder than ever… -
China’s dominance in gallium production raises semiconductor concerns
By
–
Gallium, of course, is a critical material for semiconductors. And China produces 80% of it. If you listen closely, you can hear the emergency meetings in session.
-
Archiving Influential Datasets: A Critical Gap in AI
By
–
I couldn't agree more. And they could have a powerful role in the conversation about how/if to archive influential datasets which is a woeful gap at the moment
-
Adobe prioritizes corporate users over creatives in legal protections
By
–
And let's not forget that even as Adobe says it is working on a compensation model for creatives, the first people to get legal protections here are corporate users of Adobe products. cc @alondra
-
Adobe Offers Legal Indemnification for Firefly AI-Generated Images
By
–
Whoa. Adobe is offering *full* legal indemnification for copyright lawsuits over generated images that enterprise users produce in Firefly. Their model is trained on licensed & out of copyright images, which others don't do – so it's a big throw-down.