PROP TIP Running LLMs locally? Give them web access My setup: – SearXNG: candidate source discovery – Firecrawl: known-URL scraping and crawling – Camofox: browser fallback when JS/interaction gets annoying Search → Extract → Interact Tell your favorite agent to set this
DATA
-

New data shows US AI talent advantage no longer permanent
By
–
One year after their 1st white paper analyzing DeepSeek’s talent base, HAI senior fellow @AmyZegart and @HooverInst
's Emerson Johnston publish new data that shows even more clearly that the U.S. can no longer assume a permanent AI talent advantage. (2/4) https://
hoover.org/research/updat
e-deepseek-ai-and-the-great-talent-competition
… -
Apps & Agents for Good Hackathon at DataAISummit using Databricks and OpenAI
By
–
Builders are gathering for the Apps & Agents for Good Hackathon at #DataAISummit! Using Lakebase, Agent Bricks, Databricks Apps, and @OpenAI
, teams will spend the next two days building agentic data apps for social impact. Good luck to all the participants — we're excited to see -
AI forces organizations to rethink data residency for compliance and trust
By
–
AI is forcing organizations to rethink not only how they use data, but also where that data resides. Data residency influences compliance, governance, and trust, three pillars that will shape the next wave of enterprise AI adoption. Worth keeping on every leader’s radar. @IBM
-

Problem Detection in Agent Traces with Post-Trained Model
By
–
Detecting problems in agent traces in production is difficult. It must be done at low cost (due to volume) but also with accuracy (otherwise too much noise). We post-trained our own model for this. SOTA accuracy, at ~10-100x lower cost.
-
Distillation with logits: obsolete technique, generate synthetic data
By
–
But thinking about 'distillation' only for the strict case of having access to logits is very old-school, a problem from the time when OpenAI cut off access to logits from the API. I am not up to date on what techniques exist and their effectiveness, but generate synthetic data,
-
Concern about a possible leak of Claude and Codex projects
By
–
Read carefully. What worries me most is a possible leak one day of the projects worked on in the Claude and Codex environments.
-

LEANN: Index millions with 97% less storage via graph recomputation
By
–
Turn your laptop into a powerful RAG system! LEANN can index and search through millions of documents while using 97% less storage than traditional solutions without accuracy loss. LEANN achieves this through graph-based selective recomputation with high-degree preserving
-

Linear Algebra and Optimization for Machine Learning Textbook Promotion
By
–
Linear Algebra and Optimization for Machine Learning [516-page textbook]: http://
amzn.to/39aWf8N
—————
#DataScience #DataScientist #ML #Mathematics #ORMS -

Gradient Adversarial Moving Sample Method for Scientific Machine Learning PDEs
By
–



Gradient Adversarial Moving Sample Method for Scientific Machine Learning in Solving Partial Differential Equation! #BigData #Analytics #DataScience #AI #MachineLearning #NLProc #LLM #Mathematics #IoT #IIoT #PyTorch #Python #RStats #TensorFlow #Java #JavaScript #ReactJS #GoLang
