Fine-tuning really takes some time, even if you just want to polish basic responses to pre-defined questions. If prompt engineering becomes a thing, it would require some prompt QA tooling which would be able to test GPTs at scale. Are there any frameworks already?
Seeking frameworks for prompt engineering and LLM testing at scale
By
–