Worth a try! I've been experimenting with Claude 3 and Gemini Pro 1.5 vision recently too, they're very powerful
GENERATIVE AI
-
Claude 3 Gemini Pro GPT-4 Vision outperform Tesseract OCR
By
–
No, I don't think Tesseract isn't powerful enough for that unfortunately – that's where Claude 3 / Gemini Pro / GPT-4 Vision are likely to be a better fit
-

LangChain Explained: Agents, Prompts, and Retrievers Overview
By
–
What is LangChain? via IBM There are a lot of concepts in LangChain (agents, prompts, retrievers, etc) This is a great conceptual video (no code) from IBM on how all these pieces fit together https://
youtube.com/watch?v=1bUy-1
hGZpI
… -

New LLM Primer Book Offers Practical Usage Tips
By
–
My friend @emollick assesses large language models as deeply as anyone, and his primer book, coming out this week, is chock full of useful tips for how to use them. And I love the title.
-
OCR Tool Built with Claude 3 Opus and GPT-4
By
–
Try it out here: https://
tools.simonwillison.net/ocr I wrote about how I built it – including all of the prompts I used through both Claude 3 Opus and a little bit of ChatGPT/GPT-4 – on my blog: -
Sale of an AI Voice Model with Dubbing and Automatic Conversion
By
–
Ma question maintenant c’est est ce que tu vas (tu as ?) vendre un modèle IA de ta voix en engageant un doubleur pour faire l’acting vocale et le convertir en ta voix automatiquement ? Et donc doubler des vidéos sans rien faire ?🤓 pic.twitter.com/BduE6LR6QL
— Defend Intelligence (Anis Ayari) (@DFintelligence) 30 mars 2024My question now is: are you going to (or have you?) sell an AI model of your voice by hiring a voice actor to do the voice acting and convert it into your voice automatically? And thus dub videos without doing anything?
-

AI System Generates Multiple Images with Contextual Text
By
–
It also works with multiple images!
— Pietro Schirano (@skirano) 30 mars 2024
Just ask how many you want, and it'll figure out the rest, even writing an appropriate paragraph based on the images. https://t.co/Ctjp2hPVYj pic.twitter.com/GDRcUvjYLxIt also works with multiple images! Just ask how many you want, and it'll figure out the rest, even writing an appropriate paragraph based on the images.
-
How Claude 3 and GPT-4 simplified AI prompt engineering
By
–
All you new AI engineers don’t know how easy you have it. With Claude 3/GPT-4, you can write a lazy prompt in one try and it’ll still perform pretty well. Back in the day, you had to spend hours, if not days, carefully iterating GPT-3 prompts to make things work.
-
DBRX Open-Source LLM Outperforms Benchmarks With Efficient MoE
By
–
#DBRX is a general-purpose LLM that outperforms established open source models on standard benchmarks! We’ve open-sourced it. DBRX is incredibly efficient thanks to its fine-grained MoE architecture & offers the flexibility orgs need for custom #genAI.
-

GPT App Success: From Plugin to Hundreds of Thousands Users
By
–
it works so well! I went from this early vision as a plugin to now being a full GPT app with hundreds of thousands of users. This tweet is almost 1 year old.
— Pietro Schirano (@skirano) 30 mars 2024
Always trust the process. 🙂https://t.co/g2b0CDcscrit works so well! I went from this early vision as a plugin to now being a full GPT app with hundreds of thousands of users. This tweet is almost 1 year old. Always trust the process. 🙂