Find code here: https://
github.com/patchy631/ai-e
ngineering-hub/tree/main/LaTeX-OCR-with-Llama
…
_____
Interested in ML/AI Engineering? Sign up for our newsletter for in-depth lessons and get a FREE eBook with 150+ core DS/ML lessons:
LLMS
-
LaTeX OCR with Llama: AI Engineering Hub Repository
By
–
-
Image to LaTeX Converter Using Llama 3.2 Vision Model
By
–
Image to LaTeX powered by multimodal Llama 3.2!
— Akshay 🚀 (@akshay_pachaar) 8 décembre 2024
.
.
Upload image of an equation, and it gives you the corresponding LaTeX code.
Here's what you need:
– Ollama for serving Llama 3.2 vision locally
– Streamlit for the UI
Everything is just 50 lines of code!
Link to the code in… pic.twitter.com/tMCt93sjgGImage to LaTeX powered by multimodal Llama 3.2!
.
.
Upload image of an equation, and it gives you the corresponding LaTeX code. Here's what you need: – Ollama for serving Llama 3.2 vision locally
– Streamlit for the UI Everything is just 50 lines of code! Link to the code in -
100K H100 Clusters: Grok vs Llama 4 Capabilities
By
–
i think his point of “what has a 100k h100 cluster training run brought us” is accurate – we don’t yet know. the artifacts are yet to come off that line. but whether grok will be the first or llama 4 or whatever others with a 100k+ cluster have planned is unknown.
-

FineWeb2 Releases Multilingual Dataset for AI Pretraining
By
–
The FineWeb team is happy to finally release "FineWeb2" FineWeb 2 extends the data driven approach to pre-training dataset design introduced in FineWeb 1 to now covers 1893 languages In our experiments, it tops all other publicly available multilingual pretraining datasets
-
Hugging Face Releases Fineweb-2 Dataset for AI Training
By
–
Check it out here: https://
huggingface.co/datasets/Huggi
ngFaceFW/fineweb-2
… -

FineWeb 2.0: 3 Trillion Token Multilingual Training Corpus Released
By
–
FineWeb 2.0 – 8 Terabytes, 3 Trillion tokens, 1000 languages – simply the best multilingual pre-training corpus out there! Available under a commercially permissive license!
-

GPT, Generative AI, and LLMs Conference Overview
By
–
GPT, Generative AI, and LLMs! @AverConferences #BigData #Analytics #DataScience #AI #MachineLearning #NLProc #IoT #IIoT #Python #RStats #TensorFlow #JavaScript #CloudComputing #Serverless #DataScientist #Linux #Programming #Coding #100DaysOfCode https://
geni.us/Aver-Confer -

Chain-of-thought summary and AI categories
By
–
The CoT summary mostly what you’d expect — pages and pages of ideation about plausible nicknames and letter cycles, often nonsensically meandering. This is from the middle — imagine like 100 more pages of this stuff:
-
Speculation on token probabilities and tokenization artifacts
By
–
Just a guess but maybe “ouches” and “unction” are strings just below some threshold of probability of following a space to earn a token? E.g. “unction” would be from function names (where “f” is swallowed by an unseparated prefix) and “ouch” often follows an open-quote or dash.
-

Building Custom RAG Pipelines for Generative AI Applications
By
–
RAG-Driven #GenerativeAI — Build custom Retrieval Augmented Generation pipelines: http://
amzn.to/3MWnIek v/ @PacktPublishing ——
#AI #MachineLearning #DataScience #GenAI #LLMs
——
𝓚𝓮𝔂 𝓕𝓮𝓪𝓽𝓾𝓻𝓮𝓼:
Implement RAG’s traceable outputs, linking each response to its source