In October, Whisper, a state-of-the-art audio model from @OpenAI for speech recognition was added to the library. https://
huggingface.co/docs/transform
ers/model_doc/whisper
…
MULTIMODAL AI
-

OpenAI Whisper Speech Recognition Model Added to Hugging Face
By
–
-

LayoutLM v3 Multimodal Model Added for Document Analysis
By
–
LayoutLM v3 (also from @MSFTResearch
) was added to the library in June. It is a multimodal model combining vision and text for document analysis. -
Doubling AI Model Architectures: New Audio, Vision and Multimodal Models
By
–
We doubled the number of architectures (89 to 167) with new models in audio, text, vision, multiple modalities or even time seriesand protein folding Here are a few highlights in the most used of those new models
-

Swin Transformer Vision Model for Image Classification Detection
By
–
Swin Transformer is a vision model from @MSFTResearch added back in January, which can be used as backbone for a variety of tasks such as image classification, object detection or semantic segmentation. https://
huggingface.co/docs/transform
ers/model_doc/swin
… -
Google AI Interpolates 3D Worlds from Online Photos
By
–
Google AI starts to interpolate the world into 3D spaces based entirely on photos found online pic.twitter.com/rTaqtWNcFq
— AI Breakfast (@AiBreakfast) 30 décembre 2022Google AI starts to interpolate the world into 3D spaces based entirely on photos found online
-
New AI Chat Interface Offers Natural Speech and Reduced Latency
By
–
Okay, I’m cleared to say more. It’s like ChatGPT, but the speech is more natural and terse, like texting. Much better latency (hope it lasts). Social sharing good enough to replace sharing screenshots on Twitter. Here’s me talking to Poe: https://
poe.quora.com/s/YPMAqBPuDShG
g3kl1mbt
… -

AI image generation bias: golden retrievers as blonde women
By
–
Also interesting how AI repeatedly generates blonde women with long hair if you ask for a golden retriever, makes sense I guess
-
Inpainting AI: From Near-Perfect to Flawless Image Generation
By
–
Wow! Por mucho que estas herramientas estén ya en nuestras manos, es imposible no asombrarse con los resultados.
— Carlos Santana (@DotCSV) 28 décembre 2022
Para 2023 sólo pido que los inpaintings dejen de ser imágenes *casi* perfectas, a obtener resultados que no contengan ni un sólo error. ¿Llegaremos?🌠 https://t.co/DUFFeShVxbWow! Por mucho que estas herramientas estén ya en nuestras manos, es imposible no asombrarse con los resultados. Para 2023 sólo pido que los inpaintings dejen de ser imágenes *casi* perfectas, a obtener resultados que no contengan ni un sólo error. ¿Llegaremos?
-
AI Avatar App Successfully Launches on iOS Platform
By
–
Very interesting post on someone who DID go into iOS with an AI avatar app
-
OpenAI’s Point-E Generates 3D Point Clouds From Text Prompts
By
–
OpenAI’s Point·E: Generating 3D Point Clouds From Complex Prompts in Minutes on a Single GPU https://
syncedreview.com/2022/12/27/ope
nais-pointe-generating-3d-point-clouds-from-complex-prompts-in-minutes-on-a-single-gpu/
…