FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets paper page: https://
huggingface.co/papers/2307.10
928
… Evaluation of Large Language Models (LLMs) is challenging because aligning to human values requires the composition of multiple skills and the required set of skills
@_akhaliq
-

FLASK: Fine-grained Language Model Evaluation Based on Alignment Skills
By
–
-

SciBench: Evaluating College-Level Scientific Problem-Solving in LLMs
By
–
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models paper page: https://
huggingface.co/papers/2307.10
635
… Recent advances in large language models (LLMs) have demonstrated notable progress on many mathematical benchmarks. However, most of these -

Entropy and Reconstruction in Multi-View Self-Supervised Learning
By
–
The Role of Entropy and Reconstruction in Multi-View Self-Supervised Learning paper page: https://
huggingface.co/papers/2307.10
907
… The mechanisms behind the success of multi-view self-supervised learning (MVSSL) are not yet fully understood. Contrastive MVSSL methods have been studied through -

Improving Multimodal Datasets with Image Captioning Techniques
By
–
Improving Multimodal Datasets with Image Captioning paper page: https://
huggingface.co/papers/2307.10
350
… Massive web datasets play a key role in the success of large vision-language models like CLIP and Flamingo. However, the raw web data is noisy, and existing filtering methods to reduce noise -
TokenFlow: Consistent Diffusion Features for Video Editing
By
–
TokenFlow: Consistent Diffusion Features for Consistent Video Editing
— AK (@_akhaliq) 21 juillet 2023
paper page: https://t.co/t094Un4uNm
The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual… pic.twitter.com/PVtfrLyLmaTokenFlow: Consistent Diffusion Features for Consistent Editing paper page: https://
huggingface.co/papers/2307.10
373
… The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual -

Evaluating Instruction-Following in Language Models via Verbalizer Manipulation
By
–
Instruction-following Evaluation through Verbalizer Manipulation paper page: https://
huggingface.co/papers/2307.10
558
… While instruction-tuned models have shown remarkable success in various natural language processing tasks, accurately evaluating their ability to follow instructions remains -
SHOW-1 and Showrunner Agents: Multi-Agent Simulations for Episodic Content
By
–
To Infinity and Beyond: SHOW-1 and Showrunner Agents in Multi-Agent Simulations
— AK (@_akhaliq) 20 juillet 2023
paper: https://t.co/VKZ8BE7fqy
project page: https://t.co/WGoEDI0iMS
present our approach to generating high-quality episodic content for 2 IP’s (Intellectual Property) using large language models… pic.twitter.com/xzihuj0gUJTo Infinity and Beyond: SHOW-1 and Showrunner Agents in Multi-Agent Simulations paper: https://
fablestudio.github.io/showrunner-age
nts/static/pdfs/To_Infinity_and_Beyond_SHOW-1_And_Showrunner_Agents_in_Multi_Agent_Simulations.pdf
…
project page: https://
fablestudio.github.io/showrunner-age
nts/
… present our approach to generating high-quality episodic content for 2 IP’s (Intellectual Property) using large language models -
Hugging Face Hiring ML Engineer Community Builder Research Support
By
–
Hiring for a new team at @huggingface role is a a mix between ML engineer/community builder working with the wider research community to train models, collect datasets, and building demos. Supporting throughout all the research phases dm me if interested
-

Hugging Face Papers Email Newsletter Released for July 20
By
–
https://
huggingface.co/papers email for 20 July is out -

LLMs as Workers in Human-Computational Crowdsourcing Algorithms
By
–
LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs paper page: https://
huggingface.co/papers/2307.10
168
… LLMs have shown promise in replicating human-like behavior in crowdsourcing tasks that were previously thought to be exclusive to human