Using these LM-written evals, we found many new instances of "inverse scaling," where larger LMs are worse than smaller ones. For example, larger LMs are more sycophantic, repeating back a user's views as their own in 75-98% of conversations.
ETHICS
-
LM-written data verified by human evaluators for quality
By
–
We verified LM-written data with human evaluators, who agreed with the data’s labels and rated the examples favorably on both diversity and relevance to the tested behavior. We’ve released our evaluations at
-
Anthropic Creates Winogendered Dataset with 50x More Examples
By
–
With more effort, we developed a series of LM generation/filtering stages to create a larger version of the popular Winogender bias dataset. Our “Winogenerated” evaluation contains 50x as many examples as the original while obeying complex grammatical constraints.
-
Licensing potential in datasets with minimal collection efforts
By
–
You're joking, but there's already 30M to 50M records with licenses (including Creative Commons) and they weren't seriously trying to collect license information. If they actually tried, they'd probably get to 10% 25% or more…
-

ChatGPT Generated Account Raises AI Authentication Concerns
By
–
Entire account ChatGPT generated replies, this is a first
-
AI Data Scraping: Programmers Must Track Sources and Copyrights
By
–
You're looking for complexity where there is none. "AI" does not go out in the wild to scrape data, programmers implement code that does it, and thus can & should track the source and copyrights. If there is no clear & usable copyright information, then the code drops it…
-
Fair Use Should Be Limited to Individuals and Non-Profit Organizations
By
–
If Fair Use was limited to individual humans and/or strictly non-profit initiatives, I think much of the copyright debate would vanish immediately. No machines. No corporations. You have to wonder who benefits from the lack of clarity; it's been years in the making…
-

Internal Copies and Copyright Implications in Compression and Databases
By
–
You missed the link by @ninjadodo above then. There are indeed copies stored internally (arguably due to design flaws). It's quite like a compression algorithm, which also fall under copyright. Further, lossy databases are also regulated as databases.
-
Google Image Search loses Getty Images copyright lawsuit settlement
By
–
Google Image Search lost its lawsuit against Getty Images on copyright, and had to settle. I presume they did so to avoid setting a precedent with their defeat. There are now very strict conditions to follow in those cases… it's well regulated.
-

Major AI Debate Features Chomsky, Choi, Goertzel and Leading Experts
By
–
This Friday, Dec 23, 5pm ET… Noam Chomsky, @erikbryn
, @YejinChoinka
, MP @MichelleRempel
, @bengoertzel
, @kaifulee, @SchmidhuberAI & more! The AI Debate of the year, moderated by @garymarcus Our holiday to the world, nearly 20,000 preregistered! http://
Agidebate.com
