LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding LLaVA-ST is a multimodal large language model designed for spatial-temporal understanding in videos, addressing challenges in coordinate alignment and feature compression. Problem:
@askalphaxiv
-

Diffusion Model for Portrait Animation: Joint Depth and Appearance Learning
By
–
Joint Learning of Depth and Appearance for Portrait Image Animation This research presents a diffusion-based generative model for joint learning of depth and appearance in portrait image animation, enabling consistent visual and 3D outputs for various applications. Problem:
-

WebWalker: LLM Benchmarking Framework for Web Traversal
By
–
WebWalker: Benchmarking LLMs in Web Traversal The paper introduces WebWalkerQA, a benchmark for testing LLMs in web traversal, and WebWalker, a multi-agent framework for human-like navigation to extract layered information. Problem: RAG systems retrieve shallow content,
-

GPT-4o Advances Face Presentation Attack Detection Few-Shot Learning
By
–
Exploring ChatGPT for Face Presentation Attack Detection in Zero and Few-Shot in-Context Learning The paper investigates GPT-4o for Face Presentation Attack Detection (PAD), showing strong performance in few-shot in-context learning, outperforming specialized PAD models in
-

Weekly AI Breakthroughs: Detection, ChatGPT, and Web Navigation
By
–
From uncovering physical laws in videos to animating portraits and navigating the web, this week’s AI research has brought even more breakthroughs! – Exploring ChatGPT for Face Presentation Attack Detection in Zero and Few-Shot in-Context Learning
– WebWalker: Benchmarking -

Discovering Research: Tools and Features for Academic Communities
By
–
alphaXiv has introduced communities, where members can stay connected with the latest research and discussions within a specific subfield. What tools and websites do you use to discover new research? Which features on them do you find most useful? Drop your answers below
-

Trending Papers: Agents, Reasoning, and System 2 LLMs
By
–
Trending papers this week in our newly formed Agents & Reasoning Community: 1. Agents Are Not Enough w/ @chirag_shah 2. rStar-Math 3. Search-o1 w/ Xiaoxi Li 4. Towards System 2 Reasoning in LLMs w/ Violet Xiang @gandhikanishk 5. Agent Laboratory w/ @SRSchmidgall
-
alphaXiv Chrome Extension Streamlines arXiv Paper Discovery
By
–
Also, check out the alphaXiv Chrome extension! It'll allow you to like papers directly within arXiv and take you to the corresponding community on alphaXiv!
-
Akash and Together Compute Offer $20K Credits for Agents Community
By
–
Excited to have @akashnet and @togethercompute generously offering $20K in compute credits for top contributors in the newly formed Agents Community! Join here to stay up to date with the latest in agents research: http://
alphaxiv.org/communities/ag
ents
… -
alphaXiv: Community-Curated Research Platform for ArXiv Papers
By
–
Goodreads for arXiv papers💡
— alphaXiv (@askalphaxiv) 15 janvier 2025
What if instead of arbitrary algorithms and tweets, arXiv papers were curated by your research community?
Introducing communities on alphaXiv: bridging papers, discussions, and people in one space. pic.twitter.com/1tpls6bEC6Goodreads for arXiv papers What if instead of arbitrary algorithms and tweets, arXiv papers were curated by your research community? Introducing communities on alphaXiv: bridging papers, discussions, and people in one space.
