Totally agree, AWS Textract is still the best OCR tool I've used It's a shame it's so difficult to put it in people's hands though! Just signing up for AWS billing is enough to put most people off
TOOLS
-
Ancient Open Source Libraries Power Modern AI Applications
By
–
Also neat is that the enabling libraries here – Tesseract.js and PDF.js – are both pretty old at this point: First commit to Tesseract.js was Jun 26, 2015 https://
github.com/naptha/tessera
ct.js/commit/906ce3cadbffaf5f7317a4418f282c4b78bf8385
… First to PDF.js was Apr 25, 2011 https://
github.com/mozilla/pdf.js
/commit/6dc1770bba7a417ce5664c0305469e5bb7ea76bd
… -
Tesseract.js and PDF.js: Decade-Old Libraries Still Relevant
By
–
I love that both the libraries I'm using here – Tesseract.js and PDF.js – are nearly ten years old!
-
Minimalist 226-Line Tool with PDF and OCR Capabilities
By
–
Something I really like about this tool is that the entire thing is 226 lines of combined HTML, CSS and JavaScript (plus the PDF.js and Tesseract.js dependencies, loaded from a CDN) The code is a little untidy but at 226 lines it honestly doesn't matter
-
Tesseract OCR limitations with illustrations and typeset text
By
–
Yeah, that's pretty much expected – Tesseract is great at regular typeset text, but it's really not very effective at illustrations
-

LangChain Explained: Agents, Prompts, and Retrievers Overview
By
–
What is LangChain? via IBM There are a lot of concepts in LangChain (agents, prompts, retrievers, etc) This is a great conceptual video (no code) from IBM on how all these pieces fit together https://
youtube.com/watch?v=1bUy-1
hGZpI
… -
OCR Tool Built with Claude 3 Opus and GPT-4
By
–
Try it out here: https://
tools.simonwillison.net/ocr I wrote about how I built it – including all of the prompts I used through both Claude 3 Opus and a little bit of ChatGPT/GPT-4 – on my blog: -
Sale of an AI Voice Model with Dubbing and Automatic Conversion
By
–
Ma question maintenant c’est est ce que tu vas (tu as ?) vendre un modèle IA de ta voix en engageant un doubleur pour faire l’acting vocale et le convertir en ta voix automatiquement ? Et donc doubler des vidéos sans rien faire ?🤓 pic.twitter.com/BduE6LR6QL
— Defend Intelligence (Anis Ayari) (@DFintelligence) 30 mars 2024My question now is: are you going to (or have you?) sell an AI model of your voice by hiring a voice actor to do the voice acting and convert it into your voice automatically? And thus dub videos without doing anything?
-
Signal Addresses Phone Privacy Bug, Not Zero-Click Attack
By
–
Hi! We’re confident this IS NOT a 0click attack. It’s a bug in our phone number privacy+usernames implementation, fix coming soon. Please clarify/delete this tweet bc it could inadvertently spread misinfo that could harm ppl who’d otherwise use Signal in high stakes cases
-

DesignerGPT Integrates DALL·E for Website Image Generation
By
–
Thanks to the team at @OpenAI, now DesignerGPT can also use DALL·E generated images on your website.
— Pietro Schirano (@skirano) 30 mars 2024
The possibilities are basically endless. ✨
I would have absolutely adored this tool as a student. pic.twitter.com/Dk0lGVQXOcThanks to the team at @OpenAI
, now DesignerGPT can also use DALL·E generated images on your website. The possibilities are basically endless. I would have absolutely adored this tool as a student.