5/ Boolformer – presents the first Transformer architecture trained to perform end-to-end symbolic regression of Boolean functions; it can predict compact formulas for complex functions and be applied to modeling the dynamics of gene regulatory networks.
@dair_ai
-

Vision Transformers Need Registers for Internal Computations
By
–
4/ Vision Transformers Need Registers – identifies artifacts in feature maps of vision transformer networks that are repurposed for internal computations; the proposed solutions provide additional tokens to the input sequence to fill that role.
-

Graph Neural Prompting Enhances LLMs Knowledge Learning
By
–
3/ Graph Neural Prompting with LLMs – proposes a plug-and-play method to assist pre-trained LLMs in learning beneficial knowledge from knowledge graphs (KGs).
-

The Reversal Curse: LLMs fail bidirectional generalization
By
–
1/ The Reversal Curse – finds that LLMs trained on sentences of the form “A is B” will not automatically generalize to the reverse direction “B is A”; shows the effect across model sizes and model families.
-
Top ML Papers of the Week: LLMs and Vision Transformers
By
–
Top ML Papers of the Week (Sep 25 – Oct 1): – MentalLlaMa
– Boolformer
– The Reversal Curse in LLMs
– Long-Context Scaling with LLMs
– Graph Neural Prompting with LLMs
– Vision Transformers Need Registers
… -

70B LLM Variant Surpasses GPT-3.5 Long-Context Performance
By
–
2/ Effective Long-Context Scaling with LLMs – propose a 70B variant that can already surpass gpt-3.5-turbo-16k’s overall performance on a suite of long-context tasks.
-

Galactica: Specialized LLM Achieves SOTA Scientific NLP Results
By
–
Ever wondered how far specialized large language models (LLMs) can go in scientific NLP? Meet Galactica. At the time of release, this LLM attained new SOTA results in PubMedQA and MedMCQA with a 106B-token scientific corpus and excels in tasks like LaTeX equations and reasoning.
-

OWL: Specialized LLM for IT Operations Management
By
–
9/ LLMs for IT Operations – proposes OWL, an LLM for IT operations tuned using a self-instruct strategy based on IT-related tasks; it discusses how to collect a quality instruction dataset and how to put together a benchmark.
-

KOSMOS-2.5: Multimodal AI for Document Text Generation
By
–
10/ KOSMOS-2.5 – a multimodal model for machine reading of text-intensive images capable of document-level text generation and image-to-markdown text generation.
-

Language Modeling as Compression: LLM Prediction Capabilities
By
–
7/ Language Modeling is Compression – evaluates the compression capabilities of LLMs; investigates how and why compression and prediction are equivalent; shows that LLMs are powerful general-purpose compressors due to their in-context learning abilities.
