This would also be a good thing for professional organizations to do. AMA benchmarks, New York Bar Association benchmarks, Modern Language Association benchmarks, American Psychological Association benchmarks…
@emollick
-
Leaderboards Insufficient Generic Testing Standards AI
By
–
Leaderboards are fine but not a good generic testing standard. And also likely game-able
-
AI Needs Independent Testing Standards Like Consumer Reports
By
–
We need a Consumer Reports or Underwriters Lab for AI testing. All the public benchmarks are game-able and mostly not useful measures of things LLMs do. We need secret test batteries for subject areas (coding, reasoning, human conversation, writing) and secret red team tests, too
-

Claude 3 Creates Original Optical Illusions and Visual Patterns
By
–
I asked Claude 3 to create original optical illusions. You are supposed to see:
— Ethan Mollick (@emollick) 24 avril 2024
"A perception of tilted lines from a pattern of alternating rows of tiles"
"A mesmerizing pattern"
"A square that appears to be breathing or pulsating, even though it's actually just changing size" pic.twitter.com/FrMOAx7isiI asked Claude 3 to create original optical illusions. You are supposed to see:
"A perception of tilted lines from a pattern of alternating rows of tiles"
"A mesmerizing pattern"
"A square that appears to be breathing or pulsating, even though it's actually just changing size" -

Claude 3 excels at creative puzzle-solving in teaching games
By
–
Claude 3 is pretty good at puzzles. One of our teaching games takes place on a doomed mission to Saturn, and a minor challenge is naming the ship where you are given an over-the-top manual and a set of criteria and have to come up with a name. People struggle, Claude nailed it.
-

Phi 3 Mini Model Issues: Looping and Autoregressive Problems
By
–
As I use Phi 3 mini more, it is clear that model is a bit weird. It is usually quite good to start, but tends to loop quite often, repeating things, and also sometimes descend into autoregressive madness.
-

Phi 3 Mini 3.8B: Powerful Reasoning in Tiny Model
By
–
For such a tiny model (it can run on your phone!) Phi 3 mini 3.8B seems to do a pretty nice job on real world reasoning and planning tasks.
-

Ethan Mollick Co-Intelligence Book Signing at Harvard Coop
By
–
I’ll be doing a signing & discussion for my book Co-Intelligence at the Harvard Coop Bookstore this Thursday at 6pm. If you are interested, here is the link: https://
eventbrite.com/e/author-event
-ethan-mollick-tickets-886462762987
… -
Chain of Thought in Foreign Languages for AI Reasoning
By
–
This idea of having the AI do Chain of Thought in foreign languages & tell you the results in English so as not to ruin the surprise of the game is fascinating.
-

AI Levels Up Bottom Performers, Offers Strategic Opportunity
By
–
A persistent result of studies of AI on work (like our BCG study) is that AIs improve the performance of bottom performers. Yet AI is frequently benchmarked to average or best performers. A winning strategy could be to focus on maximizing the benefits of the leveling effect.