¿Qué prompts soléis utilizar para evaluar a los LLMs y que todavía a día de hoy veáis que cuestan ser superados por los modelos actuales? Estoy recopilando y testeando. Por cierto, el Anonymous GPT por ahora pinta bien.
PROMPT ENGINEERING
-
Using Claude as an Editor: Improving AI-Generated Content Quality
By
–
I’ve been using Claude as an editor for the past few months. It’s great, but it gives me an A- on practically any first draft I send it.
— Packy McCormick (@packyM) 7 août 2024
Dan and I try to fix that. He’s a master. https://t.co/q1YEySkWPNI’ve been using Claude as an editor for the past few months. It’s great, but it gives me an A- on practically any first draft I send it. Dan and I try to fix that. He’s a master.
-

IPAdapter-Instruct: Resolving Ambiguity in Image-based Conditioning
By
–
Unity presents IPAdapter-Instruct Resolving Ambiguity in Image-based Conditioning using Instruct Prompts discuss: https://
huggingface.co/papers/2408.03
209
… Diffusion models continuously push the boundary of state-of-the-art image generation, but the process is hard to control with any nuance: -
GPT-4o blurts wrong answer, then self-corrects; structured output tangential
By
–
It helps, but in this case GPT-4o often blurts out the wrong answer at the start but then realizes and fixes the mistake as it talks. If you grade liberally in that case, most naive prompts asking for reasoning will work fine. Structured Output is slightly tangential here.
-

OpenAI docs include ‘9.11 > 9.9’ solved by JSON output
By
–

"9.11 > 9.9" is now in the OpenAI docs, as a problem solved by requesting JSON structured output to separate final answers from supporting reasoning:
-
Function Calls vs JSON Output: Usage Differences Explored
By
–
Curious how different this is from function calls as I know a lot of people were using them for just the JSON output
-

Mistral Large 2 Excels in Coding, Math, and Hard Prompts
By
–
Mistral Large 2 (2407) is now on @lmsysorg
. It performs extremely well in the Coding, Hard Prompts, Math, and Longer Query categories, where it outperforms GPT4-Turbo and Claude 3 Opus. It is also doing very well in Instruction Following where it ranks above Llama 3.1 405B. -

Using ChatGPT to generate a 30-day LinkedIn content plan
By
–
LinkedIn is the best platform for building your professional brand. I’ve made a ChatGPT prompt that crafts a 30-day LinkedIn content plan. Watch your connections and opportunities multiply. To get it, simply: • Like & Comment
• Reply 'Send'
• Follow me so I can DM you -
Hero AI App: ChatGPT Plugins Vision Now Available
By
–
Not an investor, but I've gotten to play with Hero over the past couple of weeks and it's great. It's what I thought OpenAI was building towards with ChatGPT plugins (https://t.co/2YrqhpQkuJ), in an app.
— Packy McCormick (@packyM) 5 août 2024
You can use the code PACKY to try it out. https://t.co/74pDrSUBwVNot an investor, but I've gotten to play with Hero over the past couple of weeks and it's great. It's what I thought OpenAI was building towards with ChatGPT plugins (
https://
notboring.co/p/attention-is
-all-you-need
…), in an app. You can use the code PACKY to try it out. -
Experimenting with AI-Assisted Writer for Video Summaries
By
–
Also as a bit of an experiment I tried hiring a writer to summarise the video (I suspect they got some help from ChatGPT too 😀 ) — it's not a substitute for the video, but for those in a hurry it's a quick summary of some points:
