Self-training: F(x,y) is a filtering/ranking function, eg., what we call a reward/return. The input x may be chosen by humans, but the model generates the y’s and F ranks and selects for further rounds of self-training. F can be explicit or implicit (human in the loop as in RLHF)
AI
-
Policy Gradients: Q-Functions and State-Action Value Learning
By
–
Policy gradients: F = Q(x,y) (the state-action value function), and x and y are generated by the model acting on an environment with policy p(y|x). The x’s are from the invariant state distribution as in the policy gradients theorem.
-
Supervised Learning: Function Identity and Human-Labeled Data
By
–
Supervised learning: F = I (identity), and x and y are produced by humans. E.g. x is images taken by humans and y are corresponding labels. E.g. 2, x is text and y is the next text token.
-
Learning Methods Unified Through Gradient Optimization Framework
By
–
Funny @sirbayes Learning methods — supervised, RLHF, policy gradients, Dagger, self-training — can be seen as optimisation with the following gradient: grad = Expectation_x,y [ F(x,y) grad log p(y|x) ] Choices of F and how x and y are produced determine the learning type 1/n
-
Mobile AI Tool Offers Better Interface Than ChatGPT
By
–
Just tried it on a few – really good! Definitely easier than ChatGPT interface on mobile
-
AI-Powered Long-Form Content and Podcast Summarization
By
–
I want to try it with a realllly long ramble. Would also love to have it listen to podcast convos and write the summary for me.
-
Oasis AI Portfolio Company Opens TestFlight Access
By
–
Oasis is a portco and @mattmireles let me share the TestFlight if you want to try it out: http://
oasis.so/Testflight -

Oasis AI App Transforms Voice Notes Into Multiple Summaries
By
–
The new @theoasisAI app is beautifully simple and useful. Hit record, ramble, and it summarizes what you said in a bunch of different formats. Playing with it this weekend and I’ve pulled it out to capture an idea 4 times this morning.
-
Finding Right Contact for BC Ethics Inquiry
By
–
Same problem, I tried emailing a contact but I'm not sure it's the right one. Any ideas @BCEthics ?
-
MIT’s MiniCity Tests Autonomous Urban Perception Planning
By
–
Cruising through MIT’s MiniCity, a scaled-down McWorld for testing autonomous urban perception & planning. pic.twitter.com/OsuvGoqfJk
— MIT CSAIL (@MIT_CSAIL) 11 février 2023Cruising through MIT’s MiniCity, a scaled-down McWorld for testing autonomous urban perception & planning.