Both sizes are distilled token by token from one 18B teacher, so they share one embedding space. A corpus indexed with the 9B model can be searched with 0.6B queries. That lifts ViDoRe v3 from 62.3% to 63.5% with no added query cost.
Distilled 9B and 0.6B models share embedding space for retrieval
By
–
