Built for application developers. Designed for DBAs. Databricks Lakebase modernizes OLTP with:
– Familiar Postgres semantics for developers
– Automatic scaling and recovery for admins
– Separation of compute and durable state
– One platform for operational, analytical, and AI
DATA
-

Databricks Lakebase Modernizes OLTP with Postgres Semantics
By
–
-
Fun-DDPS: Decoupled Diffusion for CO₂ Storage Inverse Modeling
By
–
Accurate modeling of CO₂ storage in the subsurface is a critical challenge for scaling Carbon Capture and Storage (CCS). 🌎 But here’s the catch: the inverse problem (recovering geological properties from sparse field observations) is severely ill-posed. 🤒 Traditional approaches either struggle with the scarcity of measured data or are too computationally expensive. 💸 In our new work, we introduce Fun-DDPS (Function-space Decoupled Diffusion Posterior Sampling), a generative framework that decouples the problem into two parts: a function-space diffusion model that learns a prior over geological parameters, and a differentiable Local Neural Operator surrogate for physics modeling and conditioning. ✨ Why does the decoupling matter? The diffusion prior handles the heavy lifting of recovering missing geological information, while the neural operator surrogate makes data assimilation fast and physically grounded — no expensive full-physics simulations in the loop. Key results on synthetic CCS datasets: ⏩ 11x improvement in forward modeling with only 25% observations ✅ Robust inverse modeling for data assimilation under sparse, noisy conditions 🔥 First rigorous validation of diffusion-based inverse solvers against asymptotically exact Rejection Sampling (RS) posteriors This points toward a scalable, practical path for uncertainty-aware subsurface characterization — something the CCS community needs as projects move from pilots to full-scale deployment. 📷 arxiv.org/abs/2602.12274 Huge kudos to my amazing collaborators — this work wouldn’t have been possible without them: @IsaacJu13 @AnimaAnandkumar Sally M Benson @ggg_www_ #CarbonCapture #CO2Sequestration #MachineLearning #DiffusionModels #NeuralOperators #AI4Science #CCS #Sustainability
→ View original post on X — @animaanandkumar, 2026-02-19 20:06 UTC
-

Google Gemini 3.1 Pro Now Available on Databricks
By
–
Google's Gemini 3.1 Pro models are now available on Databricks. Use Google’s Gemini 3.1 Pro on Databricks to build and scale GenAI applications on your enterprise data — end to end, with the governance and operational tooling teams need for production. Gemini 3.1 Pro delivers
-
Data Drift vs Process Drift in Industrial AI Operations
By
–
In industrial AI, there is a major difference between data drift and process drift. While data drift focuses on changes in input information, process drift is about the physical reality of the operation shifting over time.
— Lucian Fogoros (@fogoros) 19 février 2026
The orchestration layer handles this by coordinating the… pic.twitter.com/wrwHDy7qP3In industrial AI, there is a major difference between data drift and process drift. While data drift focuses on changes in input information, process drift is about the physical reality of the operation shifting over time.
The orchestration layer handles this by coordinating the -

Scikit-learn Cookbook (3rd Ed.) — 80+ Machine Learning Recipes
By
–
Scikit-learn Cookbook — 80+ recipes for #MachineLearning in Python with scikit-learn [3rd Edition]: http://
amzn.to/4oDGOq7 v/ @PacktDataML 𝓒𝓸𝓷𝓽𝓮𝓷𝓽𝓼:
Common Conventions & API Elements of Scikit-Learn
Pre-Model Workflow and Data Preprocessing
Dimensionality -

Advanced Diagnostics Transform Disease Detection Through AI
By
–
Advanced diagnostics are redefining how diseases are identified, by combining biology, data, and smart systems to support earlier signals, reliable decisions, and more efficient healthcare processes across labs and clinical organizations. Microblog by @antgrasso #Biotech
-

From Project to Platform: Data as Permanent Asset
By
–
The shift from project to platform is the only way to survive the industrial data paradox. A project has an end date, but a platform is a living capability. When you stop building for one specific use case and start building for adaptability, your data becomes a permanent asset
-

Open Source RAG Service with Configurable Chunking and Cosine Similarity
By
–
最近连续分享了不少 RAG 的库/服务,再来一个 RAG。 这次是 RAG as a service,开源,本地自部署免费,云端有免费额度。 特点:
• 支持可配置的 chunk size 和 overlap,用来提升召回(README 里写的是大约提升 12 个百分点的 recall,当然这是他们自己的数据)。 • 用 file_id + 余弦相似度做 -

Book: Machine Learning Design Patterns for Data Prep and MLOps
By
–
[Excellent Book] #MachineLearning Design Patterns — Solutions to Common Challenges in Data Preparation, Model-Building, and #MLOps: http://
amzn.to/2W7YSy0
——————
#AI #ML #DataScience #DataScientist -

Overview of Distance Metrics for ML and Data Science
By
–

9 Distance Metrics used in #DataScience and #MachineLearning (advantages & pitfalls): https://
maartengrootendorst.com/blog/distances/
———
#AI #ML #Statistics #Mathematics #DataScientist #Algorithms
———
See 730+ page book "Encyclopedia of Distances" (3rd edition): http://
amzn.to/3skDOWm