Have you found a chunking strategy that works well powered by gpt-3.5? I am hoping to add chunking to LLM at some point but I'm not sure what strategies I should include
CODE
-
Browser-based packages loadable from CDN networks
By
–
Do you know anything like that which runs directly in the browser via a package you can load from a CDN?
-
Expanding Tesseract.js Language Support in OCR Tool
By
–
Tesseract.js works with other languages, but I've not exposed that in my tool yet – I should do that!
-
textract-cli: Lightweight AWS Textract Command Line Wrapper
By
–
My other OCR project from yesterday: textract-cli, a tiny CLI wrapper around AWS's amazing but so-hard-to-use Textract API https://
github.com/simonw/textrac
t-cli
… Assuming you have AWS credentials configured: pipx install textract-cli
textract-cli image.jpeg > output.txt <5MB JPEG/PNGs only -
Privacy-First JavaScript Project: Local Processing No Servers
By
–
That's more sophisticated than I want to go with this particular project – the goal here is privacy-first, no servers, document stays on your computer and all the work happens in JavaScript running in your browser
-
Ancient Open Source Libraries Power Modern AI Applications
By
–
Also neat is that the enabling libraries here – Tesseract.js and PDF.js – are both pretty old at this point: First commit to Tesseract.js was Jun 26, 2015 https://
github.com/naptha/tessera
ct.js/commit/906ce3cadbffaf5f7317a4418f282c4b78bf8385
… First to PDF.js was Apr 25, 2011 https://
github.com/mozilla/pdf.js
/commit/6dc1770bba7a417ce5664c0305469e5bb7ea76bd
… -
Tesseract.js and PDF.js: Decade-Old Libraries Still Relevant
By
–
I love that both the libraries I'm using here – Tesseract.js and PDF.js – are nearly ten years old!
-
Minimalist 226-Line Tool with PDF and OCR Capabilities
By
–
Something I really like about this tool is that the entire thing is 226 lines of combined HTML, CSS and JavaScript (plus the PDF.js and Tesseract.js dependencies, loaded from a CDN) The code is a little untidy but at 226 lines it honestly doesn't matter
-
Tesseract OCR limitations with illustrations and typeset text
By
–
Yeah, that's pretty much expected – Tesseract is great at regular typeset text, but it's really not very effective at illustrations
-

LangChain Explained: Agents, Prompts, and Retrievers Overview
By
–
What is LangChain? via IBM There are a lot of concepts in LangChain (agents, prompts, retrievers, etc) This is a great conceptual video (no code) from IBM on how all these pieces fit together https://
youtube.com/watch?v=1bUy-1
hGZpI
…