Google announces Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters discuss: https://
huggingface.co/papers/2408.03
314
… Enabling LLMs to improve their outputs by using more test-time computation is a critical step towards building generally self-improving
@_akhaliq
-

Google: Test-Time Compute Scaling More Effective Than Model Parameters
By
–
-

Diffusion Models as Visual Data Mining Tools
By
–
Diffusion Models as Data Mining Tools discuss: https://
huggingface.co/papers/2408.02
752
… This paper demonstrates how to use generative models trained for image synthesis as tools for visual data mining. Our insight is that since contemporary generative models learn an accurate representation of -

MMIU: Evaluating Large Vision-Language Models with Multiple Images
By
–
MMIU Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models discuss: https://
huggingface.co/papers/2408.02
718
… The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. -
Chrome Extension to Connect arXiv and Hugging Face
By
–
Arxiv to HF
— AK (@_akhaliq) 6 août 2024
Using this chrome extension you can find papers on Hugging Face to start discussions with authors from arxiv
extension: https://t.co/nxTiHTWUTi pic.twitter.com/CuXSub6idRArxiv to HF Using this chrome extension you can find papers on Hugging Face to start discussions with authors from arxiv extension:
https://chromewebstore.google.com/detail/icfbnjkijgggnhmlikeppnoehoalpcpp
… -

14 Papers Submitted Today to Daily Papers
By
–
14 papers submitted so far today to daily papers submit your paper directly: https://
huggingface.co/papers/submit -

MiniCPM-V: GPT-4V Level MLLM for Mobile Devices
By
–
MiniCPM-V A GPT-4V Level MLLM on Your Phone paper page: https://
huggingface.co/papers/2408.01
800
… The recent surge of Multimodal Large Language Models (MLLMs) has fundamentally reshaped the landscape of AI research and industry, shedding light on a promising path toward the next AI milestone. -

Lumina-mGPT: Multimodal Generative Pretraining for Text-to-Image
By
–
Lumina-mGPT Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining paper page: https://
huggingface.co/papers/2408.02
657
… We present Lumina-mGPT, a family of multimodal autoregressive models capable of various vision and language tasks, particularly -

VidGen-1M: Large-Scale Dataset for Text-to-Video Generation
By
–
VidGen-1M A Large-Scale Dataset for Text-to-video Generation paper page: https://
huggingface.co/papers/2408.02
629
… The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant -

Language Model Can Listen While Speaking
By
–
Language Model Can Listen While Speaking https://
huggingface.co/papers/2408.02
622
… Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM) have significantly enhanced speech-based conversational AI. However, these models -

Language Models Can Listen While Speaking Simultaneously
By
–
Language Model Can Listen While Speaking https://
huggingface.co/papers/2408.02
622
… Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM) have significantly enhanced speech-based conversational AI. However, these models