but the results from the fine-tuned models were trash and the responses I was getting were really bad I am assuming it takes either significantly more refined prompt/completion data than what I gave it or just more data (both could be true)
LLMS
-
Initial Fine-Tuning Approach: Breaking Books into Chunks and Generating Questions
By
–
initial approach: oh this seems not too bad, I read this doc https://
beta.openai.com/docs/guides/fi
ne-tuning
… and was like yeah I can just break up the authors' books into chunks and generate some simple questions for each chunk and then use that to fine-tune the model -
Philosophy Author Chatbot Project Using GPT-3 Fine-Tuning
By
–
philosophy author chatbot thing w gpt background: need to create something for my philosophy of AI class idea: fine-tune gpt3 models on text from authors we have read in class and create a chat application where users can talk to these models
-
Philosophy Author Chatbot Project Using Fine-tuned GPT-3 Models
By
–
philosophy author chatbot thing w gpt background: need to create something for my philosophy of AI class idea: fine-tune gpt3 models on text from authors we have read in class and create a chat application where users can talk to these models
-
Stanford and Google Improve Transformer Training with Convex Analysis
By
–
Stanford U & Google’s Convex Analytic Training Framework Improves the Understanding and Optimization of Transformers https://
syncedreview.com/2022/12/01/sta
nford-u-googles-convex-analytic-training-framework-improves-the-understanding-and-optimization-of-transformers/
… -

MIT Scientists Use GPT-3 LLM for Clinical Natural Language Processing
By
–
Scientists at @MIT_CSAIL have used a GPT-3-style large language model (LLM) to perform tasks such as drug regimen extraction – an approach that will transform clinical natural language processing. Code > https://
bit.ly/3OK3CEd via @antgrasso #NLP #gpt3 #AI -

RA-CM3 Scalable Modular Knowledge Integration Architecture
By
–
RA-CM3 integrates knowledge in a more scalable & modular way compared to previous methods. It’s more resilient to long-tail or unseen knowledge which can help the growth & update of knowledge over time in real world settings. We’re excited for the possibilities it presents. 3/4
-

RA-CM3 Model Outperforms Baselines with 70% Less Compute
By
–
The RA-CM3 model significantly outperforms baseline multimodal models on both image & caption generation tasks, while using <30% of the compute of comparable models & exhibiting novel capabilities like knowledge-intensive image generation & multimodal in-context learning. 2/4
-

Meta AI Introduces RA-CM3: Multimodal Model for Text and Image Generation
By
–
New paper from our team at Meta AI. Retrieval-Augmented CM3 (RA-CM3) is the first multimodal model that can retrieve and generate mixtures of text and images — while also reducing training cost and model size. Now available on arXiv https://
arxiv.org/abs/2211.12561 1/4 -
Jasper AI Advances GPT Networks for User Complexity Levels
By
–
With @JasperAI, we strive to dramatically advance AI work, including training GPT networks to fit AI outputs to all levels of end-user complexity and granularity. Read more here: https://
hubs.li/Q01tMGsK0