structuring data upon ingestion and querying against the structured data is easier to debug and solve for hallucination, compared to traditional vector search that being said, it has to be structured correctly to answer questions likely to be asked, so takes more upfront effort
LLMS
-

Graph-Based AI: One Year Journey Through GraphRAG and Agents
By
–
i've been on this graph journey for over a year now so much has happened from the rise in graphrag to graph-based agents (e.g. langgraph)
-

Mini-Omni: Language Models with Hearing and Streaming Speech
By
–
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming demo is out by @freddy_alfonso_ demo: https://
huggingface.co/spaces/gradio/
omni-mini
… -
Router between Llama-3.1-405B and 70B arbitrates strength and speed
By
–
This is great news! There's one router between Llama-3.1-405B and Llama-3.1-70B, can come in really handy to arbitrate between strength and speed depending on your task!
-

OpenAI publishes datasets on Hugging Face Hub
By
–
Cool to see OpenAI publish more datasets on the @huggingface Hub recently. Now I'm hoping to also see some models soon We'll be here for you when you do!
-

Phantom Demo Released: Efficient LLM and Vision Models
By
–
Phantom demo is out Phantom is super efficient 0.5B, 1.8B, 3.8B, and 7B size Large Language and Vision Models built on new propagation strategy demo: https://
huggingface.co/spaces/BK-Lee/
Phantom
… -

OpenAI’s o1 still can’t plan reliably but is a massive leap forward
By
–
Paper read – OpenAI's o1 still can't plan reliably – but is still a massive leap forward Can OpenAI's o1 actually plan and reason, as claimed in its release? Researchers put it to the test using PlanBench, a planning benchmark that has stumped even the best language
-

Unlocking GPT-01 on ChatGPT: incredible results and new features
By
–
I unlocked #GPTo1 on #ChatGPT and the result is just incredible → https://youtu.be/OvTjuHLa7bI No more message limits, ability to share files, connect to the Internet… KILLER!
-
Qwen2.5-34B Block Replication vs Layer Replication Comparison
By
–
I agree. I'm mostly curious about the difference between replicating blocks vs. individual layers with Qwen2.5-34B-Instruct.
-
Echo Merge Experiment Tests Self-Merge Information Flow
By
–
I don't expect it to have superior performance in general, but in some narrow domains. The echo merge is an experiment to check if the "insanity" of these self-merges comes from a broken information flow.