Cleanup Pictures – Image Editor (Free +) by @cyrildiagne http://
cleanup.pictures Use to remove any unwanted text, logo, date stamp, or watermark, or part of an image. Super simple, fast tool. Free for 720p images
MULTIMODAL AI
-
Cleanup Pictures: AI-Powered Image Object Removal Tool
By
–
-
Resemble AI enables custom voice model training
By
–
Resemble AI – Voice Cloning (free to try) @resembleai http://
resemble.ai Create a text-to-speech voice model using your *own voice* by training the model. This one is truly remarkable -
Mind Reading Technology, Brain-Computer Interfaces, and Latest Research Paper
By
–
Longer thread on "mind reading tech" and BCI and the latest paper to add to the mix. https://
mind-vis.github.io -

AI Computer Vision Detects Product Defects Perfectly
By
–
From Defect to Perfect! Impact of #AI #ComputerVision #ML #DeepLearning on Quality Great example @MarinerLLC @IntelIoT https://
insight.tech/industry/produ
ct-defect-detection-you-can-count-on-with-mariner-2?utm_source=twitter&utm_medium=organic&utm_campaign=2022-tdc-eaves
… #tech #DigitalTransformation @Shi4Tech #podcast #IntelPartner @insightdottech #ITInfluencer #Product @pierrepinna @SpirosMargaris -

AI Computer Vision Transforms Self-Checkout Retail Experience
By
–
Does self #checkout always end in friction?! Now #AI #Cloud #ComputerVision Transforms #CX & Reduces #waste ! https://
insight.tech/retail/ai-and-
computer-vision-accelerate-self-checkout?utm_source=twitter&utm_medium=organic&utm_campaign=2022-tdc-eaves
… #Retail #IntelPartner #SupplyChain #fintech #POS #Foodie #AgriTech #SDGs #IoT @DeepLearn007 @SpirosMargaris @OpenFoodChain @NutriSumit2023 -
Extending LLMs to Vision: Incremental Multimodal Integration with Flamingo
By
–
Extending LLMs from text to vision will probably take time but, interestingly, can be made incremental. E.g. Flamingo (
https://
storage.googleapis.com/deepmind-media
/DeepMind.com/Blog/tackling-multiple-tasks-with-a-single-visual-language-model/flamingo.pdf
… (pdf)) processes both modalities simultaneously in one LLM. -
Why LLMs Process Text Instead of Raw Pixels
By
–
Interestingly the native and most general medium of existing infrastructure wrt I/O are screens and keyboard/mouse/touch. But pixels are computationally intractable atm, relatively speaking. So it's faster to adapt (textify/compress) the most useful ones so LLMs can act over them
-
DALL-E API Now Available in Public Beta
By
–
DALL·E API in action 🎬
— Lilian Weng (@lilianweng) 16 novembre 2022
Ref: https://t.co/pwg5emm17f https://t.co/FSrxbd6ok2DALL·E API in action Ref: https://
openai.com/blog/dall-e-ap
i-now-available-in-public-beta/
… -

Using AI-Generated Images to Enhance Machine Translation Models
By
–
Training models on “hallucinated images”, like @OpenAI’s Dall-E, to improve machine translation pic.twitter.com/TM9hvBvixr
— IBM Data, AI & Automation (@IBMData) 16 novembre 2022Training models on “hallucinated images”, like @OpenAI
’s Dall-E, to improve machine translation -
AI Models Learn Emergent Languages Through Images
By
–
Using images to teach #AI models emergent languages, making it easier to learn non-Indo-European languages 🗣️ pic.twitter.com/aaunITPKsl
— IBM Data, AI & Automation (@IBMData) 16 novembre 2022Using images to teach #AI models emergent languages, making it easier to learn non-Indo-European languages