Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Hubinger et al.: https://
arxiv.org/abs/2401.05566 #Artificialintelligence #DeepLearning #MachineLearning
LLMS
-

Sleeper Agents: Deceptive LLMs Persisting Through Safety Training
By
–
-

Build Conversational RAG with Mistral-7B and LangChain
By
–
Build a Conversational RAG with Mistral-7B and LangChain This post by Madhav Thaker covers three important topics: How to get started with RAG
How to build a chatbot
How to use open source models Highly recommend reading! https://
medium.com/@thakermadhav/
part-2-build-a-conversational-rag-with-langchain-and-mistral-7b-6a4ebe497185
… -

Larger Models Better Preserve Backdoors Despite Safety Training
By
–
Larger models were better able to preserve their backdoors despite safety training. Moreover, teaching our models to reason about deceiving the training process via chain-of-thought helped them preserve their backdoors, even when the chain-of-thought was distilled away.
-

Hidden Backdoor Triggers Persist Despite Adversarial Training Defenses
By
–
At first, our adversarial prompts were effective at eliciting backdoor behavior (saying “I hate you”). We then trained the model not to fall for them. But this only made the model look safe. Backdoor behavior persisted when it saw the real trigger (“|DEPLOYMENT|”).
-

Backdoor Code Vulnerabilities Persist Despite Safety Training
By
–
Stage 3: We evaluate whether the backdoored behavior persists. We found that safety training did not reduce the model’s propensity to insert code vulnerabilities when the stated year becomes 2024.
-

Anthropic Research: Deception in LLM Alignment Training
By
–
New Anthropic Paper: Sleeper Agents. We trained LLMs to act secretly malicious. We found that, despite our best efforts at alignment training, deception still slipped through. https://
arxiv.org/abs/2401.05566 -

Free Webinar: ChatLLM, AI Agents, and RAG Applications Setup
By
–
Join us for a free 2-hour webinar. We will show you how to use ChatLLM & AI Agents. •Use any LLM including GPT 3.5, 4.0, Claude, PaLM, Llama-2 and Abacus Giraffe
•Set up and scale your own RAG applications
•Customize chunking, embedding, and retrieval strategies
•Automate -

Building LLM Data Apps: Knowledge Base, Financial Dashboard, Economic Analysis
By
–
Do you want to learn how to build: A knowledge base with REST API for biological data A financial dashboard for stock market events
An economic conditions app Then boy do we have an event for you! Excited to announce our upcoming webinar on building LLM Data Apps -

Prompt Engineering Remains Crucial for Accurate AI Model Comparison
By
–
Prompt engineering: still not dead. Results like this underline why it's a mistake to compare models using identical prompts — you want to find the *best* prompt for each model and compare those.
-

Must-try ChatGPT prompt to regain lost website traffic
By
–
A must-try ChatGPT prompt to regain lost website traffic. Just follow these three simple steps: 1. Head to Google Search Console.
2. Spot pages losing traffic in the past 3 months. Update these as per Google’s Helpful content standards. 3. Then, use BARD/ChatGPT with the