Why Machines Learn – The Elegant Math Behind Modern AI: http://
amzn.to/4dwliQc
+
Deep Learning Foundations and Concepts: http://
amzn.to/4mw2xAE
——————
#DataScience #AI #ML #MachineLearning #Mathematics #LinearAlgebra #Probability #Calculus #DataScientist
DATA
-

Essential Books on AI Mathematics and Machine Learning Foundations
By
–
-
Embedding, Clustering and Topic Modeling Workflow
By
–
actually much simpler – 1. embed data
2. cluster
3. run topic model on clusters courtesy of @nomic_ai ! -

Databricks Adds Recursive CTEs for SQL Hierarchies
By
–
Recursive Common Table Expressions are now supported in Databricks! This brings a native way to express loops and traversals in SQL when working with hierarchical and graph-like structures. Plus, support for recursive CTEs simplifies migrations from legacy database systems. See
-
Adopting SLMs without rearchitecting your entire stack
By
–
Adopt SLMs without rearchitecting your entire stack: → Start by auditing where LLMs are used and which tasks can be offloaded → Use distillation and fine-tuning to train SLMs on those subtasks → Gradually replace LLMs with SLMs in your pipeline, monitor, and optimize
-

From Dashboards to Self-Optimizing Systems: Data-Driven Autonomy
By
–
The transition from dashboards to self-optimizing systems shows not just a technical shift, but a deeper awareness that data is not a static tool, but a capability that grows and enables alignment between human intent and machine autonomy. Microblog @antgrasso #DataDriven
-
Synthetic Data Generation for AI Model Training and Benchmarks
By
–
oh i mean it’s probably synthetic data that was generated to convey skills useful for certain benchmarks. not trying to imply that they’re training on the benchmarks. saw no evidence of that and i doubt it
-
Government Statistics Essential During AI-Driven Economic Change
By
–
As we enter a period of rapid change driven by advances in AI, we need to track changes in the economy. That means we should invest more in government statistics, not less. 1/2
-
Deduplicating Redundant AI-Generated Output Data
By
–
FUTURE WORK – deduplication even though i varied the random seed and used temperature, a lot of the outputs are highly redundant it would be prudent to deduplicate, i bet there are only 100k or fewer mostly-unique examples here
-

GPT-OSS 20B Samples Dataset Released on Hugging Face
By
–
if you want to try the data, here you go, it's on huggingface: http://
huggingface.co/datasets/jxm/g
pt-oss20b-samples
… let me know what you find! -

GPT-OSS Training Data Analysis: Bizarre Results Revealed
By
–
curious about the training data of OpenAI's new gpt-oss models? i was too. so i generated 10M examples from gpt-oss-20b, ran some analysis, and the results were… pretty bizarre time for a deep dive