You know who’s interested in learning COBOL? LLMs
GENERATIVE AI
-
AGI Genesis NFT Game Blurs AI and Human Creativity
By
–
[ THE AGI GENESIS NFT GAME ] [ P R E L I M I N A R Y C O N C E P T S E X C L U S I V E L Y ] "Imagine a world where the boundaries of human creativity and artificial intelligence blur into an immersive digital experience. Welcome to The AGI Genesis NFT Game. It's not
-

Web Scraping Wikipedia Using LLM Agents Tutorial
By
–
How to Web Scrape Wikipedia with LLM Agents We this tutorial by Kenneth Leung. It shows how to take a CSV of entities and then use an agent to populate information about those entities by looking things up on Wikipedia Many real-world use cases! https://
medium.datadriveninvestor.com/how-to-web-scr
ape-wikipedia-using-llm-agents-f0dba8400692
… -

Choosing the Right Text-to-Image AI Tool: A Decision Guide
By
–
Choosing the right text-to-image AI tool? Here’s a cheat sheet by Jonathan Parsons to help you decide! For the latest in AI, marketing, and brand building, follow @ingliguori
. #AI #Marketing #BrandBuilding -

Code Training Empowers LLMs: Generation, Reasoning, Agents
By
–
9/ How Code Empowers LLMs – an overview of the benefits of training LLMs with code-specific data. Some capabilities include enhanced code generation, enabling reasoning, function calling, automated self-improvements, and serving intelligent agents.
-
Instruct-Imagen: Multimodal Context Grounding for Heterogeneous Image Generation
By
–
10/ Instruct-Imagen – tackles heterogeneous image generation by first enhancing the model’s ability to ground its generation on an external multimodal context and fine-tunes on image generation tasks with multimodal instructions.https://t.co/3jQa0piXlG
— DAIR.AI (@dair_ai) 7 janvier 202410/ Instruct-Imagen – tackles heterogeneous image generation by first enhancing the model’s ability to ground its generation on an external multimodal context and fine-tunes on image generation tasks with multimodal instructions.
-

GPT-4V as Generalist Web Agent: 50% Task Completion
By
–
7/ GPT-4V is a Generalist Web Agent – explores the potential of GPT-4V as a generalist web agent; findings suggest that GPT-4V can complete 50% of tasks on live websites – possible through manual grounding of its textual plans into actions.
-
DocLLM: Visual Document Reasoning with Bounding Box Spatial Layout
By
–
8/ DocLLM – an extension to traditional LLMs for reasoning over visual documents; focuses on using bounding box information to incorporate spatial layout structure; demonstrates SoTA on 14 of 16 datasets across several document intelligence tasks.
-

LLM Augmented LLMs: Composing Models for Expanded Capabilities
By
–
5/ LLM Augmented LLMs – explore composing existing foundation models with specific models to expand capabilities; introduce cross-attention between models to compose representations that enable new capabilities.
-
Top ML Papers of the Week: DocLLM, ALOHA, Fine-tuning
By
–
The Top ML Papers of the Week (Jan 1 – Jan 7): – DocLLM
– Mobile ALOHA
– Self-Play Fine-tuning
– Fast Inference of MoE
– LLM Augmented LLMs
– Mitigating Hallucination in LLMs
…