Learning General World Models in a Handful of Reward-Free Deployments.
CASCADE seeks to learn a world model by collecting data with a population of agents. @YingchenX
, @JParkerHolder
, @AldoPacchiano
, @PhilipJohnBall
, @_Oleh
, Stephen J. Roberts, @_Rockt
, @EGrefen
AGENTS
-
CASCADE: Learning World Models with Multi-Agent Reward-Free Data
By
–
-
SAMPLR: Minimax Regret Method for Unsupervised Environment Design
By
–
Grounding Aleatoric Uncertainty for Unsupervised Environment Design. @MinqiJiang
, @MichaelD1729
, @JParkerHolder
, @_AndreiLupu
, Heinrich Küttler, @EGrefen
, @_Rockt
, @J_Foerst propose SAMPLR, a minimax regret UED method that optimizes the ground-truth utility function. -
Language Dynamics Distillation Improves Policy Learning
By
–
Improving Policy Learning via Language Dynamics Distillation. @hllo_wrld
, @Jayelmnop
, @LukeZettlemoyer
, @EGrefen
, @_rockt propose Language Dynamics Distillation (LDD), which pretrains a model to predict environment dynamics given demonstrations with language descriptions -
Language Abstractions Improve Intrinsic Exploration in AI
By
–
Improving Intrinsic Exploration with Language Abstractions. @Jayelmnop
, @hllo_wrld
, @RobertaRail
, @MinqiJiang
, Noah Goodman, @_rockt
, @EGrefen explore natural language as a general medium for highlighting relevant abstractions in an environment. -
Exploration Integral to General Intelligence and Learning Systems
By
–
General Intelligence Requires Rethinking Exploration. @MinqiJiang
,
@_rockt
, @EGrefen argue that exploration is integral to all learning systems despite the fact that the study of exploration in AI has mostly focused on reinforcement learning. -
AI Character Creation Platform Inspired by Character.ai Before ChatGPT
By
–
note: inspired by similar to something like http://
character.ai, but also came up with this a few weeks before chatgpt was released so a little foreshadowing there -

Update: ChatGPT external browsing works, but likes post as Grimezsz
By
–
Update — I got external browsing working and ordered ChatGPT to like this post, but for some reason it was logged into Twitter as @Grimezsz
: -
Building and Refining Your World Model Over Time
By
–
Yes, you definitely should have a world model, but you learn/refine it as you go along.
-
Improving Sequential Decision Making by Incorporating Backpropagation Principles
By
–
But this is well below that. Let me put it another way: how can we improve RL by making it more like backprop? (And by RL I mean sequential decision making, not the current set of techniques for doing it.)
-
Interactive Language Framework Enables Real-Time Language-Conditionable Robots
By
–
Interactive Language is an imitation learning framework for producing real-time, open vocabulary language-conditionable robots. Learn more and check out the newly released and largest available language-annotated robot dataset, called Language-Table → https://t.co/ZdQCeYEFJl pic.twitter.com/5zMyoa9B57
— Google AI (@GoogleAI) 1 décembre 2022Interactive Language is an imitation learning framework for producing real-time, open vocabulary language-conditionable robots. Learn more and check out the newly released and largest available language-annotated robot dataset, called Language-Table → http://
bit.ly/3Umujjt