"Simulants" is a method for synthesizing clinical trial data that maintains subject privacy, facilitating data sharing and innovation in fields like drug safety and bias analysis. Submitted by: Afrah Shafquat @Medidata
DATA
-
Data-IQ Framework Stratifies Training Data by Predictive Confidence
By
–
"Data-IQ" is a versatile framework enabling the systematic stratification of training data into outcome-based subgroups using predictive confidence & aleatoric uncertainty, aiding in feature acquisition, dataset selection, & reliable model usage. By: Nabeel Seedat @Cambridge_Uni
-
Minimax Optimal Stability Estimation Under Distribution Shift
By
–
"Minimax Optimal Estimation of Stability Under Distribution Shift" this study proposes an estimator for gauging system stability under environmental changes, providing a method to predict performance deterioration and ensure robustness. Submitted by: Yuanzhe Ma @Columbia
-
DC-Check: Data Quality Framework for Reliable ML Systems
By
–
"DC-Check" is a checklist-style framework aimed at guiding the creation of reliable ML systems by emphasizing data quality and preparation throughout all stages of the machine learning pipeline. Submitted by: Nabeel Seedat @Cambridge_Uni
-
FILA: Optimal Online Auditing Technique for ML Model Accuracy
By
–
"FILA" is an optimal online auditing technique for ML model accuracy that uses a sampling-based approach and Thompson Sampling to estimate accuracy under a finite labeling budget. Submitted by: Naiqing Guan @UofT
-
Machine Learning Analyzes Climate Change Infrastructure Impacts
By
–
"Analyzing the impact of climate change on critical infrastructure from the scientific literature" uses a weakly supervised machine learning approach to efficiently analyze a large corpus of research on climate change impacts on infrastructure. By: Tanwi Mallick @argonne
-
Lossy Compression for Large Scientific Dataset Training
By
–
"The Bearable Lightness of Big Data" presents the use of lossy compression algorithms to reduce the size of large scientific datasets while maintaining data fidelity for training deep learning models. Submitted by: Wai Tong Chung @Stanford
-
Nonlinear Dimensionality Reduction for Fluid Flow Modeling
By
–
"Comparing nonlinear dimensionality reduction for data-driven unsteady fluid flow modeling" explores various NDR techniques, ie manifold learning & deep learning, noting superior spatial reconstruction & physically interpretable modes. Submitted by: Hunor Csala @UUtah
-
AutoWS-Bench-101: Weak Supervision Evaluation Framework
By
–
"AutoWS-Bench-101" provides an evaluation of Automated Weak Supervision against zero-shot or few-shot learners, utilizing 100 labels for training models with limited labeled data. Submitted by: Nicholas Roberts @UWMadison
-
Counterfactual Fairness Reduces Dataset Bias in Weak Supervision
By
–
"Mitigating Source Bias for Fairer Weak Supervision" introduces a counterfactual fairness-based method to reduce dataset bias, improving accuracy and fairness in models trained via weak supervision. Submitted by: Changho Shin @UWMadison