Here is an example of how I plan to test next generation models: It is a GPT that writes academic papers, given a dataset, and it almost, but not quite, works well with GPT-4. I am curious how much closer GPT-5 gets. Idiosyncratic benchmarks are useful.
@emollick
-
Writer anticipated GPT-5 release for content creation
By
–
I wrote it with the expectation of GPT-5 coming out.
-

Co-Intelligence Book Launch with Interactive AI Tools
By
–
My book Co-Intelligence on living & working with AI comes out on April 2. If you pre-order it (and follow the directions on the website below), you will get access to some GPTs on March 31 that provide useful tools & new ways of interacting with the book. https://
moreusefulthings.com/book -
Model Preview Assessment: Tuning and Performance Evaluation
By
–
I will say it doesn’t feel as tuned yet as other models, it is very much a preview, so it is hard to judge directly.
-
Gemini 1.5 Pro Outperforms Gemini 1.0 Ultra with Massive Context
By
–
Gemini 1.5 Pro is a very good model, by the way. Appears to beat Gemini 1.0 Ultra which was itself GPT-4 class on the stats. Also impressive that Google can more widely release massive context windows. Google seems to be moving quickly.
-
Google Gemini 1.5 Now Available in AI Studio
By
–
Link: https://
aistudio.google.com Select Gemini 1.5 -
Gemini 1.5 Million Token Context Window Capabilities
By
–
If you want a hint about the future of AI, it is worth trying Gemini 1.5 with the 1M token context window, now available to everyone, apparently. Some of my experiments: giving it a video and having it figure out a recipe, execute instructions, watching my screen, summarize work
-

Claude 3 outperforms GPT-4 at Elvish translation
By
–
I am not sure this is useful, but Claude 3 is much better at Elvish (both Sindarian & Quenya) than ChatGPT, though they can sort of communicate with each other. When asked to translate "My hovercraft is full of eels" Claude 3 does an original translation, GPT-4 searches the web.
-
Proprietary Data Less Useful Than Expected for Pre-trained Models
By
–
After years of AI meaning building your own models trained on proprietary data, one of the hardest things for firms to grasp is that their own data is much less immediately useful for pre-trained models The value in adding data in a use case will require experimentation to learn
-
GPT-3.5 API Release Reception Compared to ChatGPT
By
–
GPT-3.5 was released in the API playground 2 days before ChatGPT-3.5 was released. The difference in reception was noticeable.