My understanding is that LLM prompt performance is linear with respect to the length of the prompt, so if you want fast responses you won't want to dump more into the prompt than necessary no matter how long it can be
By
–
My understanding is that LLM prompt performance is linear with respect to the length of the prompt, so if you want fast responses you won't want to dump more into the prompt than necessary no matter how long it can be