Ok it's crazy what you can do with text+vision > Adding vision can improve performance for text-based tasks
> OCR is trivial, so you don't need any prompt engineering
> It stays fast and cheap, even with extra vision inputs
Vision inputs improve text tasks and make OCR trivial
By
–
