With the scaffold command, any prediction id becomes a fully functional starting point for exploring AI models. ❯ replicate scaffold z5kvo4tb6mhb7xnpxpjpfxlzxe ./cats-with-hats
CODE
-

Replicate CLI Scaffold Command Simplifies Development Environment Setup
By
–
Using Replicate CLI's new `scaffold` command, you can now spin up a fully configured development environment with a single CLI command. https://
replicate.com/blog/replicate
-scaffold
… Try it with:
❯ replicate scaffold catzujlbikt7ejulil7iu4tqvu ./70s-scifi-imgs -

MLflow 2.8 LLM-as-a-Judge Evaluation for RAG Applications
By
–
Get out-of-the-box metrics like latency, tokens and more, using #MLflow 2.8 with LLM-as-a-judge Discover how it can save you time and money and best practices for #LLM evaluation in RAG applications https://
bit.ly/3FHCWAg -
LangSmith Benchmarks: Evaluate LLM Performance on Your System
By
–
More to come! Check out the blog post for more information on how to get started, or see the benchmark docs to run these on your own system. https://
blog.langchain.dev/public-langsmi
th-benchmarks/
… https://
langchain-ai.github.io/langchain-benc
hmarks/index.html
… -
LangChain Benchmarks Package for LLM and Embedding Comparison
By
–
Try it yourself We've published the new langchain-benchmarks package, which provides tooling to easily compare LLMs, embeddings, indexing techniques, and more across these datasets, so you can find the optimal solution for each task.
https://
langchain-ai.github.io/langchain-benc
hmarks/notebooks/retrieval/langchain_docs_qa.html
… -

LangChain Playground: Test LLM Improvements in Browser
By
–
If you’re logged in, you can even use the playground to try out different improvements in the browser (by clicking on any LLM run). https://
smith.langchain.com/public/452ccaf
c-18e1-4314-885b-edd735f17b9d/d/80501c49-6845-4a1e-980d-8e36dddba230/p/r/965f74aa-6c87-490c-98aa-158db8358fe9/playground
… -

LangSmith Benchmarking: Compare AI Systems Side-by-Side
By
–
Compare Since all these benchmarks are built using LangSmith, you can easily spot where different systems go wrong and compare them side-by-side. You can also go beyond aggregate statistics to examine the step-by-step execution of different systems on the same data point.
-

LangChain Q&A Dataset Tests RAG Functionality
By
–
🤝 Shared Q&A Dataset
— LangChain (@LangChain) 22 novembre 2023
We hand-crafted a Q&A dataset from LangChain’s python docs. The questions help test important RAG functionality like the ability to synthesize multiple documents and know when to abstain from guessing.
🔗 https://t.co/UEuqpyb15Q pic.twitter.com/TfZLgFc0CJShared Q&A Dataset
We hand-crafted a Q&A dataset from LangChain’s python docs. The questions help test important RAG functionality like the ability to synthesize multiple documents and know when to abstain from guessing.
https://
smith.langchain.com/public/452ccaf
c-18e1-4314-885b-edd735f17b9d/d
… -

LangSmith Launches Public Q&A Benchmark Dataset
By
–
🦜💪 Public LangSmith Benchmarks
— LangChain (@LangChain) 22 novembre 2023
Deploying LLM apps requires great evaluation, but writing evals can be painstaking.
We're launching a Q&A benchmark dataset on LangSmith so you can easily compare architectures .
Dataset: https://t.co/UEuqpyb15Q
Blog: https://t.co/Y85mNP1clJ pic.twitter.com/JkZ0MlbHROPublic LangSmith Benchmarks Deploying LLM apps requires great evaluation, but writing evals can be painstaking. We're launching a Q&A benchmark dataset on LangSmith so you can easily compare architectures . Dataset: https://
smith.langchain.com/public/452ccaf
c-18e1-4314-885b-edd735f17b9d/d
…
Blog: https://
blog.langchain.dev/public-langsmi
th-benchmarks/
… -
GPT cannot safely emit long data URIs; ask for downloadable files
By
–
The issue is GPT can't safely emit long OOD sequences like data URIs during inference; it's more reliable if you ask for the result as a downloadable file.
