err i commented too quickly and people might get overexcited. i was just saying model-tool specific cotraining is achieves same or better results than using a 10-100x larger model without cotraining. this research doesn’t touch on memory/context compression, which is yet another
Model-Tool Cotraining Achieves Better Results Than Larger Models
By
–