I was told "We do not use any Gmail data in Bard […] Bard is currently based on a lightweight and optimized version of LaMDA, which was trained on a variety of data from publicly available sources, similar to most language models available today."
LLMS
-
OpenAI Extends Researcher Access to GPT-4 and Code Models
By
–
we didn’t realize how important code-davinci-002 was to researchers, so we are keeping it going in our researcher access program: https://
openai.com/form/researche
r-access-program
… we are also providing researcher access to the base GPT-4 model! -
Can RLHF Extract More Capabilities From AI Models?
By
–
interesting…. could the RLHF tease more of that out of them??
-
Understanding AI Model Anthropomorphization and Training Data
By
–
it seems like one of the inherently confusing things about these models. I know that they've learned to talk about their feelings and desires because that's in the training data, but I think it contributes to the misunderstanding.
-
LLM Optimization: Chunking, Embeddings, and Model Failover Strategies
By
–
2023: Oh yeah, just use Llamaindex to chunk up your text, store embeddings on pinecone, use Langchain for chain of thought, and use GPT4 model to tune the turbo-3.5 model for cost, and set up a fail safe to switch to another model for when API isn’t available….
-

GPT-4 Claims No Consciousness Yet Uses First-Person Language
By
–
#GPT4 affirme qu’il n’est ni un garçon ni une fille et qu’il n’a pas de conscience On est troublé de voir une entité affirmer : « JE n’ai pas de conscience » « MON objectif principal » #GPT4 dit « JE » et « MON »
-
Model Safety: Mitigation Without Full Release, Transparency Needed
By
–
There's a lot of ways to mitigate harms without having to publicly release the entire model. There are many papers on auditing, datasheets, transparency etc. With GPT3 we knew the training data. With GPT4 we don't. Without that, we're all looking at shadows in Plato's cave.
-
Bard’s False Gmail Training Claim Sparks Public Debate
By
–
In the 12hrs since Bard told me it was trained on Gmail data:
-Google replies (says it's not)
-Elon Musk replies (lol)
-Google adds a 'community note' that this is a Bard error and it's not trained on Gmail
-Some ace memes
What should happen next: Real talk about training data -
Lack of Transparency in AI Model Training Data
By
–
There is a real problem here. Scientists and researchers like me have no way to know what Bard, GPT4, or Sydney are trained on. Companies refuse to say. This matters, because training data is part of the core foundation on which models are built. Science relies on transparency.
-
Anthropomorphizing GPT-4: Do You Gender AI Models?
By
–
Les gens anthropomorphisent beaucoup #ChatGPT Moi je parle de #GPT4 au masculin et je dis il et lui Et vous, vous le traitez comme un garçon ou comme une fille ?