3/ Typewriter results None of the agents are perfect. GPT-4 had a hard time typing "keyboard" and "head" https://
smith.langchain.com/public/ff14ecb
2-3664-4c4a-b2dc-d8aa9fd2185d/d
…
LLMS
-

GPT-4 Typewriter Task Performance Evaluation and Limitations
By
–
-

New Open-Source Tool Use Benchmarks for LLM Agents
By
–
Agents are the “killer” LLM app, but building and evaluating agents is hard. A huge part of agents is tool use, but there aren't enough open-source tool use benchmarks out there. Today, we are excited to release four new test environments for benchmarking LLMs’ ability to
-
Meta Research: Understanding Back-Translation at Scale
By
–
Understanding Back-Translation at Scale – Meta Research | Meta Research https://
bit.ly/486P4a6 #AI #MachineLearning #DeepLearning #LLMs #DataScience -
Training Models to Generate Gist Tokens for Intermediate Reasoning
By
–
yeah this is very related idea — dense vectors convey more information than discrete tokens, in less space! but it's not the same thing, i'm describing how to train a model that generates gist tokens on the fly for its intermediate reasoning steps
-
Interpretability vs Efficiency Trade-offs in AI Models
By
–
ok good point good point i'm sacrificing all hope of interpretability for the sake of efficiency / performance latent scratchpad wouldn't be interpretable I think I've read chain-of-thought can also be misleading/wrong though even if it produces the right answer
-

Latent Chain-of-Thought: Making LLM Scratchpads Non-Human-Readable
By
–
fun research idea: Latent chain-of-thought / Latent scratchpad it's well-known that language models perform better when they generate intermediate reasoning tokens through some sort of 'scratchpad'. but there's no reason scratchpad tokens need to be human-readable. in fact,
-
Gemini: Highly Capable Multimodal Models Family
By
–
Gemini: A Family of Highly Capable Multimodal Models Anil et al.: https://
arxiv.org/abs/2312.11805 #ArtificialIntelligence #DeepLearning #MachineLearning -
GPT-4 Turbo JSON Mode for Valid Model Responses
By
–
With GPT-4 Turbo, we’ve introduced JSON Mode, which ensures model responses are all valid JSON objects. Is this what you’re looking for, or is there something more specific you need? Thanks! https://
platform.openai.com/docs/guides/te
xt-generation/json-mode
… -
Parallel Function Calling Issue Investigation Request
By
–
Thanks, Matt! Regarding parallel function calling, is this with gpt-4-1106-preview? How many functions are you working with? If you have a specific instance that we might be able to reproduce, please let us know—we’re eager to look into it!
-
Documentation priorities for model changes and updates feedback
By
–
Thanks for the feedback! Regarding docs on model changes, what’s most important to you? Grasping differences between models for your prompts, comparing model performances with evals, or getting clearer updates on upgrades and timelines?