A @Stanford study reveals that leading AI companies are pulling user conversations for training. Should users of AI chatbots worry about their privacy?
DATA
-

AI Infrastructure Transforms Finance: LLMs Power Smarter Payments
By
–
Join us at @Money2020
: AI in Finance — Scaling Financial Intelligence with @OracleCloud & NVIDIA Hear from leaders at @Mastercard
, @Stripe
, and @Fiserv on how specialized LLMs and high-performance AI infrastructure are powering smarter payments, hyper-personalization, and -
Data Mixing Balance: Vision-Language Skills and Regional Imbalance
By
–
That's a great question and will take the opportunity to dash off a bit on the data mixing. Mixing data is a tricky balance it turns out. There were two main factors at play: – we wanted to keep general vision-language skills.
– and had unbalanced regions and languages: think -
Scaling 2.8M Images and 3M Entities with Flexible Data Curation Pipeline
By
–
We have around 2.8M unique images and 3M unique entities and the data curation pipeline allow scaling that up(more images per entity, adding new countries/languages, questions per entity, etc…). Very glad to hear people were asking about this 🙂
-
Major VQA Dataset Release: 30M Samples Across 42 Countries
By
–
Congrats! Big props unifying vision language datasets and making it easy to access them. We also recently released a massive VQA dataset of 30M samples(22M filtered), spanning 42 countries and 39 languages. The dataset contains factual visual questions and answers about world
-
Complexity vs Feature Engineering Trade-offs in Machine Learning
By
–
The less complexity the more capacity to focus on learning other things. But the more handcrafted features the less opportunity to learn better representations. It all kind of depends on how big the model is, how much data you have, and what your compute resources are
-
Image Format Complexity Requires More Training Data
By
–
exactly, that’s the messiness of working with the image format I mentioned. I think you can make to generalize well on all these but since there are more degrees of freedom it will require more data to train (luckily this can be done with automatic data augmentation but still)
-
Visual Tokenization Challenges: Aspect Ratios and Image Preprocessing Complexity
By
–
I know it’s popular to hate tokenizers, but visual representations (which are also tokenized) bring a lot of messiness as well. Aspect ratios, cropping, resolution, brightness, etc. Sure, models learn to deal with that but it requires lots of data to make them robust wrt these.
-
Weekly embedding evaluations breakthrough in AI research
By
–
It's Christmas every week in embedding eval land
-

SAS Viya Boosts Enterprise Productivity and Workflow Efficiency
By
–
No late nights needed, just productivity and pizza See how SAS Viya helps you boost productivity and reclaim your time http://
2.sas.com/6018AfVdq