Volume 206 https://
proceedings.mlr.press/v206/ Proceedings of AISTATS 2023 Is now available on PMLR.
MACHINE LEARNING
-
AISTATS 2023 Proceedings Volume 206 Now Available on PMLR
By
–
-
ACML 2022 Proceedings Volume 189 Published on PMLR
By
–
Volume 189 https://
proceedings.mlr.press/v189/ Proceedings of ACML 2022 Is now available on PMLR. -

Chain-of-Thought Prompting Enables Multi-Step Reasoning in Language Models
By
–
3 (cont). One way to elicit reasoning is via "chain-of-thought (CoT) prompting", which gives examples of intermediate reasoning steps in-context. CoT prompting enables large LMs to do multi-step reasoning tasks, increasing the range of tasks that LMs can do.
-
Reasoning: The Key Differentiator Between Classical ML and Intelligence
By
–
3. The last idea is reasoning, which differentiates classical ML techniques from intelligence. Classical ML approaches need a lot of data and are black-box. Intelligent agents learn from a few examples and can do abstract reasoning.
-

Unpredictable Emergence in Language Models: Key Implications
By
–
There are at least four profound implications of emergence:
2A. Emergence cannot be predicted simply by extrapolating the scaling curves from smaller models.
2B. Emergent abilities are not explicitly specified by the trainer of the language model. -

Emergence: Large Language Models Gaining Unexpected Complex Abilities
By
–
2. Emergence is a phenomenon where large language models gain abilities that are not present in smaller language models. An example of an emergent ability is doing complex math questions.
-

Scaling Laws: Model Size, Data, and Compute for LM Improvement
By
–
Key takeaways: 1. Scaling involves increasing model size, data, and compute. Scaling is challenging (cost, infra, etc), but important, since "scaling laws" tell us that scaling predictably makes LMs better.
-
Three Ideas Driving the LLM Revolution: Scaling, Emergence, and Reasoning
By
–
I gave an invited lecture at New York University for @hhexiy
's class! I covered three ideas driving the LLM revolution: scaling, emergence, and reasoning. I tried to frame them in a way that reveals why large LMs are special in the history of AI. Slides: -

Efficient LLM Training with Sparsity and Dataflow Techniques
By
–
TECHNICAL RESEARCH PAPER: Training Large Language Models Efficiently with Sparsity and Dataflow This paper demonstrates an end-to-end training flow on a LLM – 13 billion GPT – using sparsity and dataflow. @arxiv
: https://
arxiv.org/abs/2304.05511
PDF: https://
arxiv.org/pdf/2304.05511
.pdf
… #ml #llm -
African Languages in AI: Bridging Tech Representation Gap
By
–
Watch the first episode of the African Language series, part of our work to highlight the vast diversity of global linguistics. We’ll discuss the expanse of Africa’s languages, how they’re underrepresented in tech, & how to improve language tech for them →