This asynchrony really should exist in normal one-on-one chatbot interactions. When you pose ask a hard question to an LLM it should say, immediately, “This is a hard problem — give me 15m.” If that’s an issue, you should be able to say so and get a quicker guess.
LLMs should request delay for hard questions or provide quick guesses
By
–