2. Building evaluations. Many benchmarks get quickly saturated, and we need more to evaluate the frontier of language models. In addition, it’s still an open question of how to evaluate language models generally. The new OpenAI evals library could be good: https://
github.com/openai/evals
AI
-
Building Evaluations for Frontier Language Models
By
–
-
Prompting Research: Just the Beginning for Language Model Guidance
By
–
1. Prompting research. Maybe hot take, but I think we’ve just reached the tip of the iceberg on the best ways to prompt language models. As language model capabilities increase, the degrees of freedom for guiding a particular generation via a good prompt will increase.
-
PhD Research Directions in the Era of Large Language Models
By
–
I’m hearing chatter of PhD students not knowing what to work on.
My take: as LLMs are deployed IRL, the importance of studying how to use them will increase.
Some good directions IMO (no training):
1. prompting
2. evals
3. LM interfaces
4. safety
5. understanding LMs
6. emergence -

ESM-2: Meta’s Large Language Model Advances Protein Structure Prediction
By
–
One of the most important advances for #AI in life science has been #AlphaFold. The prediction of proteins from amino acid sequences at atomic level has a new large language model, ESM-2, from @Meta https://
science.org/doi/10.1126/sc
ience.ade2574
… @ScienceMagazine -
Staying Updated on LLM Jailbreaks and Exploits
By
–
well, now that @gdb qt'd this tweet, I feel I have to share this… keep up w the current state of jailbreaks and LLM exploits by subscribing to my newsletter here: http://
thepromptreport.com -
GPT-4 fails to write without ‘e’ via naive prompting
By
–
Writing without the letter "e" is still beyond GPT-4 with naive prompting, but that's maybe a cheap shot.
-
Democratized Red Teaming: Building Robust AI Models Through Adversarial Testing
By
–
Democratized red teaming is one reason we deploy these models. Anticipating that over time the stakes will go up a *lot* over time, and having models that are robust to great adversarial pressure will be critical. Also considering starting a bounty program/network of red-teamers!
-

Customized AI Models Drive Enterprise Efficiency and Competitive Advantage
By
–
“Compared with one-size-fits-all-models where your data needs to be shipped off to a third-party API, customized owned models bring huge efficiency savings that give businesses a clear competitive advantage in their lines of work.” – @MarshallChoy https://
techmonitor.ai/technology/ai-
and-automation/microsoft-copilot-for-work-office-365-ai
… -

Microsoft 365 Copilot Launch: Understanding Usefully Wrong
By
–
From Microsoft 365 Copilot launch post one hour ago… What is “usefully wrong”?
-

Economy Impact on Manufacturing, Security and Women Leadership
By
–
Had an insightful conversation with #AndeHazard on #CXOSpice. We discussed the impact of economy on #manufacturing, network security & talent shortage, improving diversity & inclusion, and empowering women to lead. Her advice? Be proactive, embrace https://
youtube.com/watch?v=_TfyRw
aZ-w8&t=14s
…