10/ AgentBoard – a benchmark with an open-source evaluation framework to perform analytical evaluation of LLM agents; assesses the capabilities and limitations of LLM agents and demystifies agent behaviors which leads to building stronger LLM agents.
AGENTS
-
AI agents enable end-to-end task completion in a single window
By
–
You literally won’t need to leave a single chat window to complete the whole task
-

ChatGPT New Feature: Combining GPTs Changes Everything
By
–
Big new feature #ChatGPT → We can COMBINE GPTs and it changes everything. I show you all this in my latest video:
https://
youtu.be/sQ8iZB0YuD8 #GPT4 #GPTs -

Sam Witteveen’s LangGraph Guide: Essential Weekend Watch
By
–
Some awesome YouTube content on LangGraph by @Sam_Witteveen Sam has awesome, high-quality guides, and this new one on LangGraph is no exception a GREAT weekend watch https://
youtube.com/watch?v=PqS1ki
b7RTw
… -
Utility Function Terminology and Human Evolutionary Misalignment
By
–
I think that "utility function" in retrospect is a mathematical word of power that I should not have expected lay computer scientists to understand, so let's drop that. With humanity, the outer optimization criterion was inclusive fitness; our inner preference was not aligned.
-
Intelligence and Outcome Matching: The Core Problem Beyond Utility Functions
By
–
It isn't about "simple" utility functions or "monomania". The problem is just any sufficiently smart system whose work, on some level, can be viewed as matching up outputs and results, and learning.
-
Learning Reality and Selecting Outputs for Desired Outcomes
By
–
The problem that 'learn how reality works' and 'select outputs which, when they interact with reality, lead to X happening' is a simple great way of doing Y for a lot of possible Y. For example, with humans, Y is inclusive genetic fitness and X is all the stuff that humans want.
-
Utility Functions and Existential Risk from Instrumental Convergence
By
–
https://
arbital.com/p/instrumental
_convergence/
… but I'm not sure what that buys you if "pick any simple measure on utility functions, preimage them through a reasonable environmental model onto actions, most utility functions kill humanity as a side effect" doesn't already do it. -
Bing Sydney’s Unexplained Threat Behavior Remains Mystery
By
–
Nope! To this day, as far as I know, nobody knows which particular numbers inside Bing Sydney led her to try to threaten a human with reporting him to the police. They poked her until she stopped doing that, but can't read her thoughts any more than we can look at a frozen
-
Preference Functions and Alignment: Optimizing for Human Flourishing
By
–
If you don't see why most preference functions that don't specifically have an attainable optimum around people living happily ever after, do something else instead of that, I'm not sure what particular inscrutable properties of a kqv layer are going to be persuasive?
