Writing fiction seems to be a genuine weak spot for LLMs that is not improving as rapidly as almost every other area. There may be a lot of reasons why this is happening. It would be a really interesting benchmark (but you would need human judges, AI judges love AI fiction).
LLMs struggle with creative fiction and require new human-led benchmarks
By
–