It can be done: Gather any available real query examples, add in synthetic Q’s derived from docs. Synthesize responses using RAG pipeline. SFT (a smaller model) on resulting Q&A’s, without RAG context. Whether it saves any money is another question.
Synthesizing AI Responses with RAG Pipelines and SFT Models
By
–