Longer thread on "mind reading tech" and BCI and the latest paper to add to the mix. https://
mind-vis.github.io
MULTIMODAL AI
-
Mind Reading Technology, Brain-Computer Interfaces, and Latest Research Paper
By
–
-

AI Computer Vision Detects Product Defects Perfectly
By
–
From Defect to Perfect! Impact of #AI #ComputerVision #ML #DeepLearning on Quality Great example @MarinerLLC @IntelIoT https://
insight.tech/industry/produ
ct-defect-detection-you-can-count-on-with-mariner-2?utm_source=twitter&utm_medium=organic&utm_campaign=2022-tdc-eaves
… #tech #DigitalTransformation @Shi4Tech #podcast #IntelPartner @insightdottech #ITInfluencer #Product @pierrepinna @SpirosMargaris -

AI Computer Vision Transforms Self-Checkout Retail Experience
By
–
Does self #checkout always end in friction?! Now #AI #Cloud #ComputerVision Transforms #CX & Reduces #waste ! https://
insight.tech/retail/ai-and-
computer-vision-accelerate-self-checkout?utm_source=twitter&utm_medium=organic&utm_campaign=2022-tdc-eaves
… #Retail #IntelPartner #SupplyChain #fintech #POS #Foodie #AgriTech #SDGs #IoT @DeepLearn007 @SpirosMargaris @OpenFoodChain @NutriSumit2023 -
Extending LLMs to Vision: Incremental Multimodal Integration with Flamingo
By
–
Extending LLMs from text to vision will probably take time but, interestingly, can be made incremental. E.g. Flamingo (
https://
storage.googleapis.com/deepmind-media
/DeepMind.com/Blog/tackling-multiple-tasks-with-a-single-visual-language-model/flamingo.pdf
… (pdf)) processes both modalities simultaneously in one LLM. -
Why LLMs Process Text Instead of Raw Pixels
By
–
Interestingly the native and most general medium of existing infrastructure wrt I/O are screens and keyboard/mouse/touch. But pixels are computationally intractable atm, relatively speaking. So it's faster to adapt (textify/compress) the most useful ones so LLMs can act over them
-
DALL-E API Now Available in Public Beta
By
–
DALL·E API in action 🎬
— Lilian Weng (@lilianweng) 16 novembre 2022
Ref: https://t.co/pwg5emm17f https://t.co/FSrxbd6ok2DALL·E API in action Ref: https://
openai.com/blog/dall-e-ap
i-now-available-in-public-beta/
… -

Using AI-Generated Images to Enhance Machine Translation Models
By
–
Training models on “hallucinated images”, like @OpenAI’s Dall-E, to improve machine translation pic.twitter.com/TM9hvBvixr
— IBM Data, AI & Automation (@IBMData) 16 novembre 2022Training models on “hallucinated images”, like @OpenAI
’s Dall-E, to improve machine translation -
AI Models Learn Emergent Languages Through Images
By
–
Using images to teach #AI models emergent languages, making it easier to learn non-Indo-European languages 🗣️ pic.twitter.com/aaunITPKsl
— IBM Data, AI & Automation (@IBMData) 16 novembre 2022Using images to teach #AI models emergent languages, making it easier to learn non-Indo-European languages
-
Pantheon Lab Creates Realistic AI Human with Natural Voice
By
–
The person in this video is not a real human…
— Pascal Bornet (@pascal_bornet) 16 novembre 2022
She was created by start-up Pantheon Lab. The progress made over the last 2 years is amazing: natural face, voice and lipsync using AI. This opens up opportunities in healthcare, services#innovation #artificialintelligence #tech pic.twitter.com/OzQomS1b0NThe person in this video is not a real human… She was created by start-up Pantheon Lab. The progress made over the last 2 years is amazing: natural face, voice and lipsync using AI. This opens up opportunities in healthcare, services #innovation #artificialintelligence #tech
-
Robot Dog Walks on Stools Using Onboard Vision
By
–
New work from @berkeley_ai and @CMU_Robotics on visual locomotion enables a robot dog walking on tall bar stools in @ashishkr9311's living room — entirely from onboard cameras and compute. Trained in simulation and deployed directly in the real world!https://t.co/kBEpOCsRob https://t.co/wVIeiMgfd2
— Berkeley AI Research (@berkeley_ai) 15 novembre 2022New work from @berkeley_ai and @CMU_Robotics on visual locomotion enables a robot dog walking on tall bar stools in @ashishkr9311
's living room — entirely from onboard cameras and compute. Trained in simulation and deployed directly in the real world! http://
vision-locomotion.github.io