The 12.5 million people living in French Banlieue together with Heetch wanted to do something about it On each postcard is a QR code to a large image dataset to help Midjourney correct its data on the Banlieue I find this to be a beautiful way of protesting and making a stand
DATA
-
Human captions improve AI training through synthetic data generation
By
–
The training involved human writers crafting detailed captions, harmonizing subject and context. This change reduced caption 'noise', enhancing training accuracy. But more importantly, they used these human data to train a captioner model to create synthetic data!
-
Cisco Evolutio Partnership Delivers Enterprise AI Observability Platform
By
–
We’re excited to partner with @Cisco to deliver enterprise-grade observability for both predictive and #generativeAI with a new #MLOps module for the Cisco Observability Platform built with Evolutio. This new module delivers always-on monitoring and production diagnostics to
-
Rumor about a new Amazon LLM named Olympus
By
–
Rumor that Amazon is developing a new LLM, codenamed Olympus. They have one giant data set that others don’t: What you buy.
-

PySpark Course: Master Big Data with Python Apache Spark
By
–
PySpark Course: Big Data Handling with Python and Apache Spark https://
bit.ly/4744c7W #AI #MachineLearning #DeepLearning #LLMs #DataScience -
Google Maps Traffic Data Vulnerable to Manipulation Attacks
By
–
Google Maps uses dynamic, user-sourced data for traffic speed gauging. However, it's prone to manipulation, for example, when a person tricked it into rerouting by walking with a bag full of phones.
-
Fine-tuning LLMs: Knowledge Integration and Efficient Updates
By
–
lots of problems need to be solved before we have this:
– how to finetune LLMs on data so they “know” the data, and don’t just mimic it?
– how do we combine knowledge from pretraining on the universe with finetuning on small datasets?
– how do we efficiently update with new docs? -
Vector Databases Replaced by Single Transformer Models
By
–
prediction: every vector database will eventually be replaced with a single transformer a row for every individual datapoint is so old-school, feels outdated what would be better: a single differentiable blob, something that “knows” about all your data and can chat about it
-
Vector Databases for LLM Applications and RAG Systems
By
–
Vector databases are a key part of many LLM applications that need search or data retrieval, for example with Retrieval Augmented Generation (RAG).
— Andrew Ng (@AndrewYNg) 8 novembre 2023
Learn how they work + how to use them in our new short course, taught by @weaviate_io's @sebawita!https://t.co/Yi0mnGt9pE pic.twitter.com/ACuueLLpEkVector databases are a key part of many LLM applications that need search or data retrieval, for example with Retrieval Augmented Generation (RAG). Learn how they work + how to use them in our new short course, taught by @weaviate_io
's @sebawita
! https://
deeplearning.ai/short-courses/
vector-databases-embeddings-applications
… -
Nvidia cuDF: Accelerate Pandas with GPU Computing
By
–
Every body may have used pandas for your data analysis and feature engineering work. now Nvidia has come of with an amazing library cudf where in now you use pandas with GPU in an accelerated mode.