New open-source project dropping in a few hours. Doing this one a bit differently…. Building a version of it into the HyperWrite platform to make it easier for everyone to try.
@mattshumer_
-
Claude 3 Haiku Model Behavior Changes and Refusal Patterns
By
–
Claude 3 Haiku is starting to refuse more prompts — was there a model update behind the scenes?
-
Rating LLM outputs: hand-labeling versus prompt-based evaluation
By
–
Determine what you care about (ex: conciseness, factuality, style, etc.). The either:
– hand rate a bunch of examples on each dimension, train a model to understand what to look for
– build a LLM prompt with Mistral or similar to do this without fine-tuning Run over all the -
Manual Model Evaluation Through Held-Out Test Sets
By
–
I have a held-out test set that I manually prompt the model with and evaluate myself, and then beyond that, I just keep asking the model questions till I feel the vibes
-
Model Unification and A/B Testing Platform Improvements
By
–
Interesting. Definitely try using the platform with this (provide examples). If you’re on a good model a/b test, it should work near-perfectly. Sorry for the a/b test confusion — we’re hopefully going to unify on one model in the coming weeks so there’s no guessing.
-
Prompt Optimization: Context Size Impact on Model Performance
By
–
Depends how many examples you have. If say, 500 or less, include the task description in the prompt so the model picks it up quicker. If you’re in the thousands, no context often leads to similar or better performance and slightly lower inference costs.
-
Base Model Enables Custom Writing Styles Via Chat
By
–
Yep! Can’t share the base model, but it already works incredibly well. You can:
– describe a specific style for it to write in, and/or
– show it examples of the style you want Available through the chat interface that we offer today. -
HyperWrite New Model Rollout Soon A/B Testing
By
–
The new HyperWrite model is just so damn good If you’re lucky enough to get it in our A/B tests, you’ll see what I mean Broader rollout soon
-
Comparing AI Model Performance and Prompt Engineering Techniques
By
–
Sorry to disagree, but those models are far weaker than the leading models mentioned. The outputs won’t even be remotely as useful. Try the “no fingers” trick.