I think synthetic data for fine-tuning is a seoarate issue from whether accidental AI-generated text in pre-training data has an impact on model quality
GENERATIVE AI
-
Writing Accessible AI Guide Inspired by Being Digital
By
–
An inspiration for my book was Negroponte’s “Being Digital.” I read it in college and it led me to MIT & the Media Lab. I wanted to write something that was a similar accessible & research-backed introduction to what AI actually can do for us right now.
-
Detecting AI-Generated Images and Model Collapse Risks
By
–
I can believe that – detecting images generated by their own image generator feels like a more tractable problem to me than detecting text from their LLMs Also model collapse for images feels more likely to me than for text, though I'm not sure I could explain that intuition
-
Haiku vs GPT-3.5 comparison gains traction over flagship models
By
–
I'm finding the comparison between Haiku and GPT 3.5 a whole lot more interesting than the comparison between Opus and GPT-4
-
AI Generated Training Data Quality vs Human Data Impact
By
–
No idea! A lot of people seem to believe that accidental AI generated training data is uniquely harmful to models compared to low quality human data – I'm trying to figure out what the latest thinking on that is
-
Curse of Recursion: 2024 Updates and Vendor Mitigation Strategies
By
–
That's about the "Curse of Recursion" from a year ago – I'm looking for a 2024 update on that. Are there new developments that counter the claims from that paper? What are the big model vendors doing (if anything) to mitigate that risk?
-
AI Training Data Ethics and Cultural Appropriation Concerns
By
–
Dear God make it stop. Say it with me – “AI uses human data.” So if “it” ever exhibits knowledge it’s already an aggregate of stolen data, IP and creativity from humans. Also – Maslow’s hierarchy was built on Blackfoot First Nations wisdom and has been miscommunicated.
-
AI Models Can Use Links in Prompts for Context
By
–
If the context is online, you can actually paste links to it in the prompt, and the model can decide to use it if it wants! So it’s very likely possible.
-
OpenAI training data transparency and AP licensing concerns
By
–
Yeah I've been wondering about that – is the more recent training data mostly stuff they've licensed from sources like the AP? https://
apnews.com/article/openai
-chatgpt-associated-press-ap-f86f84c5bcc2f3b98074b38521f5f75a
… As always the infuriating lack of training transparency just leaves us guessing -

GPT-4 Predicts Future Events Accurately Through Narrative Storytelling
By
–
Not 100% sure what to make of this timey-wimey paper showing GPT-4 is able to predict the future quite accurately (or, after least make guesses about events that happen after its training cut-off) but only when asked to tell stories about what will happen. https://
arxiv.org/abs/2404.07396