Initial Results We compared different LLMs within a retrieval Q&A chain with an OpenAI functions agent. We also compared against an agent built using OpenAI’s assistant API. In the initial tests, the Assistant Agent performed the best! https://
smith.langchain.com/public/452ccaf
c-18e1-4314-885b-edd735f17b9d/d/compare?selectedSessions=d21fe3ec-719c-44bf-a6dd-c0f6bc414e9c20b3ae59-9eee-43f0-9776-1c1f41b15aa29ea0c6d3-31dd-4943-a96d-0d9ee320daae
…
LLMS
-

LLM Comparison: Assistant API Outperforms Function Agents
By
–
-

LangChain Q&A Dataset Tests RAG Functionality
By
–
🤝 Shared Q&A Dataset
— LangChain (@LangChain) 22 novembre 2023
We hand-crafted a Q&A dataset from LangChain’s python docs. The questions help test important RAG functionality like the ability to synthesize multiple documents and know when to abstain from guessing.
🔗 https://t.co/UEuqpyb15Q pic.twitter.com/TfZLgFc0CJShared Q&A Dataset
We hand-crafted a Q&A dataset from LangChain’s python docs. The questions help test important RAG functionality like the ability to synthesize multiple documents and know when to abstain from guessing.
https://
smith.langchain.com/public/452ccaf
c-18e1-4314-885b-edd735f17b9d/d
… -

LangSmith Launches Public Q&A Benchmark Dataset
By
–
🦜💪 Public LangSmith Benchmarks
— LangChain (@LangChain) 22 novembre 2023
Deploying LLM apps requires great evaluation, but writing evals can be painstaking.
We're launching a Q&A benchmark dataset on LangSmith so you can easily compare architectures .
Dataset: https://t.co/UEuqpyb15Q
Blog: https://t.co/Y85mNP1clJ pic.twitter.com/JkZ0MlbHROPublic LangSmith Benchmarks Deploying LLM apps requires great evaluation, but writing evals can be painstaking. We're launching a Q&A benchmark dataset on LangSmith so you can easily compare architectures . Dataset: https://
smith.langchain.com/public/452ccaf
c-18e1-4314-885b-edd735f17b9d/d
…
Blog: https://
blog.langchain.dev/public-langsmi
th-benchmarks/
… -

System 2 Attention: Making LLM Reason from FAIR
By
–
System 2 Attention.
Making LLM reason.
From @jaseweston and @tesatory at FAIR. -
Prompt Hacking: Understanding AI Vulnerabilities and Defense Mechanisms
By
–
Prompt hacking involves using ingenuity to make AI breach its ethics, leading to harm. This challenge shows the need to grasp such interactions to boost AI's defenses against abuse.
-

Prompt Hacking: Security Issue in LLMs and Solutions
By
–
'prompt hacking' lets one access events uninvited, posing a big issue to models like ChatGPT or any LLM, which we tackled with @learnprompting and colleagues from @mila ( @jerpint ), @ChengleiSi from @stanfordnlp
, and @towards_AI
. -
GPT cannot safely emit long data URIs; ask for downloadable files
By
–
The issue is GPT can't safely emit long OOD sequences like data URIs during inference; it's more reliable if you ask for the result as a downloadable file.
-

Free OSS and Google Models Access in LangSmith Playground
By
–
OSS models and Google models in LangSmith Playground for FREE Want to experiment with new models easily? No need for an API key! We've partnered with @thefireworksai and @GoogleAI to give you access to both OSS and Google models such as Mistral-7b, LLaMA2-70b or
-
DataLLMs Boost Organization Efficiency in Days
By
–
DataLLMs can dramatically increase the efficacy of your organization. They are much more effective than having to go through a mountain of reports that no one looks at. Abacus can set up a DataLLM for you in a couple of days. Read this: https://
blog.abacus.ai/blog/2023/08/2
4/data-llm-get-insights-from-your-data/
… 13/13 -
Flexible LLM Options: Closed-Source and Custom Models
By
–
LLMs: You can use closed-source (e.g. GPT-4) or your own custom LLM (Abacus-Giraffe or Llama2). You can compare and contrast whichever LLM works for you and pick the best one 11/13