By popular demand, we've automated the "turn the Module into a Keras Layer" step. So now you can just… use torch Modules in a Keras model/layer. No extra step needed. Just like you can use Keras models/layers as Modules.
CODE
-

Building Arabic Language Models: Data Collection and Training Strategy
By
–
The key challenge of training any non-English model is data. To build a dataset large & rich enough for Arabic, Jais tapped into various open and curated sources totally 55B tokens. The model trained on 2 epochs of this data.
-
Joy and Art in Debugging Low-Level Numerical Issues
By
–
there is surprising joy & art in debugging low-level numerical issues
-

Meta’s Code Llama Models Now Available on Poe
By
–
All @MetaAI
's #CodeLlama models are now available on Poe! @poe_platform > http://
poe.com – web – iOS – Android & MacOS -

Masking PII in Language Models with LangChain Expression Language
By
–
This is great stuff from @MaksOpp and @deepsense_ai Let's you easily mask PII before passing it into the language model – and with LangChain Expression Language is very clear/easy how to plug that into your chain
-

MLFlow AI Gateway Now Powers RAG Applications with Llama2
By
–
You can now use the highly-scalable #MLFlow AI Gateway for your RAG apps! This blog shows it all in action with a RAG application built using the API gateway, Llama2 and hosted models on MosaicML. Give it a read https://
bit.ly/3OTWA06 -
Efficient LLM Training QLoRA Llama 2 Resource Optimization
By
–
Depends on your settings. But if you limit the context size to like 2048 (like in the NeurIPS competition) and use a microbatch size of 1 with gradient accumulation and qlora with llama 2 7B, that’s approx 20 GB RAM and shouldn’t take too long, maybe an hour.
-

Weave Interactive Data Exploration Tool for Analysis
By
–
Data Delight: "Weave" Your Way to Interactive Exploration! https://
bit.ly/3Kw3es1 #AI #MachineLearning #DeepLearning #LLMs #DataScience -
ChatGPT 4 Significantly Outperforms 3.5 for Developers
By
–
devs that still doubt the productive power of chatgpt are def using 3.5 please try 4 before dunking
-
Philippe’s Trolling Reshape Function Naming Convention
By
–
looks like phillippe wanted to troll everyone by calling what should've been `tl.reshape` as `tl.view`.
It's like those trolly C macros ( #define float int )
