This is a new benchmark (ToolComp) that encompasses a broader range of Tool Use scenarios than prior benchmarks. Uniquely, we are further evaluating the models utilizing process supervision labels. We also split the benchmark into Enterprise and Chat use cases to differentiate
PROMPT ENGINEERING
-

Models Struggle with Instruction Following and Hallucinated Information
By
–
In our error analysis, we found that most models continued to struggle with: – Final Answer Missing Information
– Hallucinated Information This indicates a widespread issue with precise instruction following, especially in prompts requiring detailed information and intermediary -
New Data Enrichment Agent for Research and Form Completion
By
–
As part of our templates launch we added a brand new data enrichment agent This will do research on a particular topic and fill out a form with that research! Useful for ai sdrs, etc
-
Comparing AI Models: o1 vs Sonnet 3.5 vs 4o vs Gemini
By
–
blew past o1 for the week, then sonnet 3.5 for the day, and now unsure if i should be using o1-mini or 4o or go through the pain of setting up gemini
-
Aggressive prompt writing raises concerns in AI development
By
–
Damn that promt is written aggressively.
-
Three Cognitive Architecture Templates Launch: ReAct, RAG, Data Enrichment
By
–
we're launching a small number (3) of templates for common cognitive architectures
— Harrison Chase (@hwchase17) 19 septembre 2024
these are configurable (choose LLM/vector store of your choice)
Starting with:
– ReAct agent
– RAG chatbot
– data enrichment (research) agent
What should we add next? https://t.co/bqOclMrhUNwe're launching a small number (3) of templates for common cognitive architectures these are configurable (choose LLM/vector store of your choice) Starting with:
– ReAct agent
– RAG chatbot
– data enrichment (research) agent What should we add next? -

Asking o1-preview how many r’s are in strawberry
By
–
Asking ChatGPT o1-preview
¿,ʎɹɹǝqʍɐɹʇs, uᴉ ǝɹɐ s,ɹ ʎuɐɯ ʍoɥ -

Updated Hacker News AI Papers Ranking Logic Using o1-mini
By
–
updated ranking logic for Hacker News AI papers using o1-mini
-

Chain-of-Thought Prompting: Benefits Limited to Math and Reasoning
By
–
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning discuss: https://
huggingface.co/papers/2409.12
183
… Chain-of-thought (CoT) via prompting is the de facto method for eliciting reasoning capabilities from large language models (LLMs). But for what kinds of tasks
