have you considered starting with @markatgradient
's 1m context llama 3 instead of just the 8b context?
Consider Llama 3 1M Context Instead of 8B Context
By
–
By
–
have you considered starting with @markatgradient
's 1m context llama 3 instead of just the 8b context?