Announcing OPT-IML: a new language model from Meta AI with 175B parameters, fine-tuned on 2,000 language tasks — openly available soon under a noncommercial license for research use cases. Research paper & more details on GitHub
RESEARCH
-

Postdoc Opportunity in Surgical AI with Anand Kumar
By
–
Apply to do a postdoc with me and @AjhungMD on Surgical AI https://
cms.caltech.edu/about/position
s/surgicalai
… -
MoleculeSTM: ChatGPT Technology Applied to Molecular Analysis
By
–
MoleculeSTM paving the way for #chatGPT for molecules https://t.co/EtPho1oj0b
— Prof. Anima Anandkumar (@AnimaAnandkumar) 22 décembre 2022MoleculeSTM paving the way for #chatGPT for molecules
-
Largest Text-Molecule Model Enables ChatGPT-like Molecule Retrieval and Editing
By
–
We build the largest Text-molecule model that does not rely only on aligned training pairs. Now you can retrieve and edit molecules based on text prompts. This will pave the way for #ChatGPT for #molecules https://
chao1224.github.io/MoleculeSTM @nvidia @Mila_Quebec @Caltech -

MIT researchers link brain language systems to code ML representations
By
–
The brain has multiple demand and language systems responsible for various cognitive tasks. MIT researchers found that these systems align w/ML representations of code, potentially helping us understand how our brains read programs: http://
bit.ly/3vb5ehn -
OODA Loop 2022: ChatGPT GPT-3 OpenAI Future Analysis
By
–
OODA Loop 2022: The Past, Present, and Future of ChatGPT, GPT-3, OpenAI, NLMs, and NLP https://
oodaloop.com/ooda-original/
disruptive-technology/2022/12/22/ooda-loop-2022-the-past-present-and-future-of-chatgpt-gpt-3-openai-nlms-and-nlp/
… via @ooda -
Zshot: Linking Entity Mentions to Exemplars via Transformers
By
–
Zshot uses a transformers model that links mentions to exemplars you provide to the pipeline component. You can use standard NER mentions with @spacy_io 's built-in models, or you can train your own with https://
prodi.gy, optionally assisted by @OpenAI suggestions. -
spaCy Extension for Zero-Shot Labeling with Prodigy Integration
By
–
Following up with the theme of zero and few-shot labelling, this spaCy extension by @IBMResearch is super cool, and could be used in combination with our recent Prodigy recipe: https://
github.com/explosion/prod
igy-openai-recipes
… -

10 Common Data Quality Issues and Solutions
By
–
10 Most Common Data Quality Issues and How to Fix Them: Ensuring data quality guarantees more data-informed decisions. Hence, this article highlights the common data quality issues and ways to overcome them. https://
kdnuggets.com/2022/11/10-com
mon-data-quality-issues-fix.html?utm_source=dlvr.it&utm_medium=twitter&utm_campaign=10-most-common-data-quality-issues-and-how-to-fix-them
… -

SantaCoder: First LLM with Opted-Out Training Dataset
By
–
So cool to see @BigCodeProject release SantaCoder, afaik the very first LLM trained on an "opted-out" dataset, aka allowing people to opt-out from the training dataset. 1.1B parameters that outperforms larger models on both generation and infilling! https://
huggingface.co/bigcode/santac
oder
…
