Declarative Pipelines is going open source! We’re excited to bring a proven standard for building reliable, scalable data pipelines with Apache Spark™ to the community. Built with openness and composability in mind, Declarative Pipelines allows developers to define what their
DATA
-

Unity Catalog Metrics and Discover Transform Business Data Management
By
–
Today, at #DataAISummit @matei_zaharia introduced major innovations that expand Unity Catalog to business users!
– Unity Catalog Metrics: A single source of truth for business metrics across data engineering, ML and BI
– Unity Catalog Discover: New curated internal marketplace -

Databricks Announces Full Apache Iceberg Support with Unity Catalog
By
–
We’re excited to announce full Apache Iceberg™ support in Databricks, unlocking the full Iceberg and Delta Lake ecosystems with Unity Catalog! This is a major advancement toward a single, unified open table format. Now you can:
– Use Databricks and external engines to read and -

Overparametrized Gradient Descent Learns Gaussian Mixtures
By
–
A very cool result showing overparametrized gradient descent (EM) learns Gaussian mixtures. It uses tools from tensor decomposition we developed and connects it to gradient descent.
-

LangGraph+ Tensorlake Integration for Agent Document Understanding
By
–
LangGraph+ Tensorlake: Unlocking Document Understanding for Agents When creating agents that interact with data, the connections you have to that data matter Excited to announce an integration Tensorlake – a best in class document ingestion engine https://
tensorlake.ai/blog/announcin
g-langchain-tensorlake-integration
… -
Auditing Your Organization’s AI Readiness Framework
By
–
How to Audit Your AI Readiness Practical framework for assessing organizational AI readiness. Covers data infrastructure, technical capabilities, workforce skills, and leadership preparation. Essential guide for successful AI implementation.
-
Knowledge Alone Insufficient: AI Systems Need Massive Context
By
–
Knowledge doesn’t solve everything. We have the knowledge on why smoking is bad and how to stop, a billion still use tobacco. Knowledge, like the cure for cancer or climate change, while helpful, is not the whole picture. AI systems will need mass context to better analyze
-

India Launches AIKosh National AI Dataset Repository Initiative
By
–
.
@OfficialINDIAai
, @GoI_MeitY invites Expressions of Interest (EOI) from academic institutions, startups, NGOs, think tanks, and industries to contribute non-personal, anonymised, India-specific datasets to #AIKosh — the national AI dataset repository. Why contribute? -
MES and Industry 4.0 Summit Begins in Porto
By
–
MES & Industry 4.0 International Summit is about to start. Stay tuned from updates from Porto, Portugal#sponsored #criticalmfg_iiot #Industry40 @GregorianCT1 via @fogoros pic.twitter.com/nM4tAJbyOQ
— Lucian Fogoros (@fogoros) 12 juin 2025MES & Industry 4.0 International Summit is about to start. Stay tuned from updates from Porto, Portugal #sponsored #criticalmfg_iiot #Industry40 @GregorianCT1 via @fogoros
-

Optimal Advantage Regression Accelerates RL for LLM Reasoning
By
–
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
Paper: https://
arxiv.org/pdf/2505.20686
.pdf
…
Code: https://
github.com/ZhaolinGao/A-PO
