Added a new "Relevant Context" section that shows up in the "Task Execution" stage. This is to show how we're using @Pinecone vector search to provide memory from past tasks into the current task being executed. It's empty on step 1.
DATA
-
Replenishing the Commons: Training Data Sustainability Challenge
By
–
I don't know how to engineer it but it feels like there needs to be some way to ensure the commons (from which the training data is pulled) gets replenished.
-

ChatGPT Outperforms Crowd Workers on Text Annotation Tasks
By
–
6/ ChatGPT for Text-Annotation – shows that ChatGPT outperforms crowd-workers on several annotation tasks such as relevance, topics, and frames detection; besides better 0-shot accuracy, the per-annotation cost of ChatGPT is ~20 times cheaper than MTurk.
-

AI Training Data Crisis: Documentation and GPL Solutions
By
–
Was discussing this with @monkchips
. "Docs" likely to increasingly look like "ChatGPT", but @peternixey raises a critical point: where will the training data come from if devs are quietly asking machines for guidance? Maybe we need a new GPL to avoid one-way knowledge traps? -
Maximizing Weights and Biases for Machine Learning Workflows
By
–
Lately I’ve been using @wandb more and more. I’ve been using configs to store hparams and ofc logging to track loss, lr, etc different runs. Still feel like I’ve not been maximising it. What are some underrated features you use everyday?
-

Alteryx Celebrates Earth Month With Data Analytics for Sustainability
By
–
We're passionate about using #DataAndAnalytics to change our communities and the planet. To celebrate #EarthMonth, we're hosting a month of global activities, challenges and opportunities focused on #sustainability. https://
alteryx.com/alteryx-for-go
od
… #InvestInOurPlanet #AlteryxForGood -

Local and Cloud Embedding Storage Options for AI Systems
By
–
Our system is designed to give the user the option of storing the text and embedding data locally on a powerful device (Memory Backpack™) – or with any of our embedding hosting providers like @pinecone or @vectara
. -
Facial Recognition Profiles Generated From Social Network Data
By
–
Facial recognition is used to generate profiles and histories of people you interact with. You can connect your social accounts to automatically ingest your network to further provide context (images, contact history, etc).
-
Seven Datasets for AI Text-to-Image Synthesis Models
By
–
Transforming words into images just got easier with these seven intriguing datasets! Explore the possibilities of AI-generated art with text-to-image synthesis models. Read more: https://
bit.ly/3nEnjUN @GoogleAI @Microsoft @jeevprabnivash @_DigitalIndia @GoI_MeitY -
French Language Dataset Size and LLM Training Data Deduplication
By
–
ça fait que 32go a télécharger le fr avec les images. C'était le cas pour THE PILE qui a servi de base à d'autre LLM comme GPT-J-6 ou GPT-neo, nlg-megatron, mais me semble pas. A moins qu'ils aient considére que ça faisait doublon et que la trad serait suffisante. Tu peux
