For my own work, I view the arXiv version as the definitive version, no matter what conference or journal it's in. No weird page limits, templates, etc., everything is just the way I intend it. It may also be a carry-over from other fields, where proceedings are often closed.
RESEARCH
-
Intelligence lies in optimization methods, not goals
By
–
Intelligence is not in the objective function, it’s in how you optimize it.
-
Building and Refining Your World Model Over Time
By
–
Yes, you definitely should have a world model, but you learn/refine it as you go along.
-
Blurred Line Between AI Training and Copyright Infringement
By
–
There’ll be a fuzzy & heavily litigated line between training an AI and copyright infringement. Is it a fuzzy interpolated search restating data from Yelp, Quora, Stack Overflow? Or remixing a thousand copyrighted artists, coders, authors? Or is it learning, thinking, creating?
-
Actor-Critic Reinforcement Learning as Backpropagation Through Time
By
–
So actor-critic RL is backprop through time.
-
Improving Sequential Decision Making by Incorporating Backpropagation Principles
By
–
But this is well below that. Let me put it another way: how can we improve RL by making it more like backprop? (And by RL I mean sequential decision making, not the current set of techniques for doing it.)
-
Reinforcement Learning and Backpropagation: A Dynamic Programming Perspective
By
–
You’re taking it too literally. I’m trying to point to a way to improve RL by recognizing that at a certain level of abstraction it’s doing something similar to backprop (namely, efficiently propagating later results back to earlier choices via dynamic programming).
-

DeduceLogic Facts Based on Logic
By
–
We can also check how it can deduce facts based on logic. Here is an example
-

Stable Diffusion 2: Foundational Model Powering Next Generation
By
–
Perhaps @StableDiffusion 2 was misnamed. the “2” branding implies it’s same but better, but it’s not – out of the box. it’s a LOT better when you know how to wield it, or finetune it. SD2 is a *foundational* model now, that will power an entire generation of txt2img models.
-
Interactive Language Framework Enables Real-Time Language-Conditionable Robots
By
–
Interactive Language is an imitation learning framework for producing real-time, open vocabulary language-conditionable robots. Learn more and check out the newly released and largest available language-annotated robot dataset, called Language-Table → https://t.co/ZdQCeYEFJl pic.twitter.com/5zMyoa9B57
— Google AI (@GoogleAI) 1 décembre 2022Interactive Language is an imitation learning framework for producing real-time, open vocabulary language-conditionable robots. Learn more and check out the newly released and largest available language-annotated robot dataset, called Language-Table → http://
bit.ly/3Umujjt